Two-way communication system, method, and program
The two-way communication system adjusts speech rates based on emotion analysis, normal speech comparison, or time constraints to synchronize speaking speeds, improving communication efficiency and meeting completion.
Patent Information
- Application Number
- JP2025010138
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2025-01-23
- Publication Date
- 2025-08-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing two-way audio communication systems fail to adjust speaking speeds appropriately, leading to issues where users may speak too quickly or too slowly, affecting communication efficiency.
A two-way communication system that adjusts the speech rate of audio collected by one terminal and transmits it to another, using emotion analysis, comparison with normal speech rates, or remaining time to synchronize speaking speeds, with automatic answering devices generating responses at adjusted rates.
The system effectively influences speakers to match their speech rate with others, enhancing communication efficiency and meeting completion within time constraints.
Smart Images

Figure 2025115971000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a two-way communication system, method, and program. [Background technology]
[0002] Conventionally, there is known a technology for two-way audio communication between multiple terminals (for example, video conferencing, etc.). For example, Patent Document 1 describes that a neck-worn terminal equipped with a microphone and a speaker wirelessly communicates with another neck-worn terminal. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6786139 Summary of the Invention [Problem to be solved by the invention]
[0004] However, there are cases where the speaking speed of the users of each device is not appropriate. For example, you may want to prompt a person who is speaking too quickly to slow down, or a person who is speaking too slowly to speed up.
[0005] The present disclosure aims to make a speaker speak slower or faster in two-way audio communication between multiple terminals. [Means for solving the problem]
[0006] A two-way communication system according to a first aspect of the present disclosure includes: A two-way communication system for two-way audio communication between a plurality of terminals, changing the speaking rate of the speech collected by at least one terminal or generated by an automatic answering machine; The speech whose speech rate has been changed is transmitted to a terminal other than the terminal from which the speech was collected.
[0007] According to a first aspect of the present disclosure, a person who hears speech whose speech rate has been changed is induced to speak slower without realizing it, or is induced to speak faster without realizing it, due to the speech rate of the speech.
[0008] A second aspect of the present disclosure is a two-way communication system according to the first aspect, The speech rate of the voice collected by at least one of the terminals or the voice generated by the automatic answering device is changed in response to an instruction from a user of any of the terminals or based on a judgment by the two-way communication system.
[0009] According to the second aspect of the present disclosure, any user can instruct a change in speech speed, or the speech speed can be changed automatically.
[0010] A third aspect of the present disclosure is a two-way communication system according to the second aspect, Based on the result of analyzing the emotion of the user of the terminal where the voice is collected, the speech speed of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is changed.
[0011] According to the third aspect of the present disclosure, it is possible to change the speech rate according to the emotion of the terminal user.
[0012] A fourth aspect of the present disclosure is a two-way communication system according to the third aspect, When the emotion of the user of the terminal where the voice is collected is impatience, the speech rate of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is reduced.
[0013] According to the fourth aspect of the present disclosure, when a person is in a hurry and hears a voice with a slowed-down speaking rate, the person is lured by the slowed-down speaking rate and starts speaking slower without realizing it.
[0014] A fifth aspect of the present disclosure is a two-way communication system according to the second aspect, Based on the difference between the normal speaking speed of the user of the terminal where the voice is collected and the speaking speed of the voice collected by the terminal where the voice is collected, the speaking speed of the voice collected by a terminal other than the terminal of the user or the voice generated by the automatic answering device is changed.
[0015] According to the fifth aspect of the present disclosure, it is possible to change the speech speed in accordance with the difference from the normal speech speed.
[0016] A sixth aspect of the present disclosure is a two-way communication system according to the fifth aspect, When the speech speed of the voice collected at the terminal collecting the voice is faster than the normal speech speed of the user of the terminal collecting the voice, the speech speed of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is reduced.
[0017] According to the sixth aspect of the present disclosure, when a person who speaks faster than normal hears a voice with a slowed-down speaking rate, that person is lured into speaking slower without realizing it by the slowed-down speaking rate of the voice.
[0018] A seventh aspect of the present disclosure is a two-way communication system according to the fifth aspect, When the speech speed of the voice collected by the terminal collecting the voice is slower than the normal speech speed of the user of the terminal collecting the voice, the speech speed of the voice collected by a terminal other than the terminal of the user or the voice generated by the automatic answering device is increased.
[0019] According to the seventh aspect of the present disclosure, when a person who speaks slower than normal hears a voice with an increased speaking speed, the person is lured by the increased speaking speed of the voice and begins to speak faster without realizing it.
[0020] An eighth aspect of the present disclosure is a two-way communication system according to the second aspect, The speech rate of the voice collected by the at least one terminal or the voice generated by the automatic answering device is changed based on the remaining time of the conference or work by the two-way voice communication.
[0021] According to the eighth aspect of the present disclosure, the speech rate can be changed according to the remaining time of the meeting or work.
[0022] A ninth aspect of the present disclosure is a two-way communication system according to the eighth aspect, When the remaining time of the conference or work by the two-way voice communication falls below a threshold, the speech rate of the voice collected by all the terminals and the voice generated by the automatic answering device is increased.
[0023] According to a ninth aspect of the present disclosure, by listening to the voice with the increased speaking speed, everyone is lured into speaking faster without realizing it, allowing the meeting or work to be completed within the remaining time.
[0024] A method according to a tenth aspect of the present disclosure, comprising: A method executed by a two-way communication system for two-way audio communication between a plurality of terminals, comprising: changing the speech rate of the voice collected by at least one terminal or the voice generated by the automatic answering device; transmitting the speech whose speech rate has been changed to a terminal other than the terminal from which the speech was collected; Includes.
[0025] A program according to an eleventh aspect of the present disclosure includes: On the computer, a step of changing the speech rate of the voice collected by at least one terminal among a plurality of terminals that perform two-way voice communication or the voice generated by the automatic answering device; a step of transmitting the speech whose speech rate has been changed to a terminal other than the terminal from which the speech is collected; Execute the following. [Brief explanation of the drawings]
[0026] [Figure 1] FIG. 1 is a diagram illustrating an overall configuration according to an embodiment of the present disclosure. [Figure 2] 1 is an example of a wearable terminal according to an embodiment of the present disclosure. [Figure 3]FIG. 1 is a diagram illustrating a hardware configuration of a wearable terminal according to an embodiment of the present disclosure. [Figure 4] FIG. 2 is a diagram illustrating a hardware configuration of a two-way communication management device according to an embodiment of the present disclosure. [Figure 5] 1 is a functional block diagram of a two-way communication system according to an embodiment of the present disclosure. [Figure 6] FIG. 10 is a sequence diagram of a speech speed conversion process based on emotion analysis according to an embodiment of the present disclosure (in the case of a remote supporter terminal). [Figure 7] FIG. 10 is a sequence diagram (in the case of a remote supporter terminal) of a speech speed conversion process based on a comparison with a normal speech speed according to an embodiment of the present disclosure. [Figure 8] FIG. 10 is a sequence diagram of a speech speed conversion process based on the remaining time of a meeting or task according to an embodiment of the present disclosure (in the case of a remote supporter terminal). [Figure 9] FIG. 10 is a sequence diagram of a speech speed conversion process based on emotion analysis according to an embodiment of the present disclosure (in the case of an automatic answering device). [Figure 10] FIG. 10 is a sequence diagram of a speech speed conversion process based on a comparison with a normal speech speed according to an embodiment of the present disclosure (in the case of an automatic answering device). [Figure 11] FIG. 10 is a sequence diagram of a speech speed conversion process based on the remaining time of a meeting or task according to an embodiment of the present disclosure (in the case of an automatic answering device). [Figure 12] 1 is a diagram illustrating a speech section and a non-speech section according to an embodiment of the present disclosure. FIG. [Figure 13] FIG. 10 is a diagram illustrating a method for lowering the speech rate according to an embodiment of the present disclosure. [Figure 14] FIG. 10 is a diagram illustrating a method for increasing the speech rate according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0028] <Terminology> In this specification, a "terminal" refers to a terminal for bidirectionally communicating at least sound between multiple terminals. For example, the terminal may be a wearable terminal equipped with a microphone and a speaker. The terminal may also be a personal computer, smartphone, tablet, etc. equipped with a microphone and a speaker, or a personal computer, smartphone, tablet, etc. connected to a microphone and a speaker. In this specification, "speech rate" refers to the speed of speech, and is, for example, the value obtained by dividing the number of moras in speech by the duration of the speech (for example, the number of moras per second, the number of moras per minute, etc.).
[0029] <Summary> In a two-way communication system according to an embodiment of the present disclosure, the speech rate of speech collected by at least one terminal is changed, and the speech with the changed speech rate is transmitted to a terminal other than the terminal from which the speech was collected. (Note that this includes cases where a device managing two-way communication (two-way communication management device 20) converts the speech rate of the speech, where the terminal from which the speech was collected converts the speech rate of the speech, and where a terminal other than the terminal from which the speech was collected converts the speech rate of the speech.) When a speaker hears speech with a slowed speech rate, he or she is persuaded to speak more slowly, or when a speaker hears speech with an increased speech rate, he or she is persuaded to speak more quickly. Note that speakers are known to synchronize with the speech rate of the other party (see, for example, "Komatsu Takanori and Morikawa Koji, 'Speech Rate Entrainment in Interactive Communication between Humans and Artificial Objects,' Research Report of the Information Processing Society of Japan, 2004-ICS-137(10), October 29, 2004"). The speech rate of the voice generated by the automatic answering device may be changed.
[0030] <Overall composition> Fig. 1 is a diagram showing the overall configuration according to one embodiment of the present disclosure. The two-way communication system 1 includes wearable terminals 10A, 10B, and 10C (hereinafter collectively referred to as wearable terminals 10), a two-way communication management device 20, and at least one of a remote supporter terminal 30 and an automatic response device 40. Note that Fig. 1 illustrates a case where there are three wearable terminals 10 and one remote supporter terminal 30, but the number of terminals is not limited to this. Each of these will be described below.
[0031] <<Wearable devices>> The wearable terminal 10 is a terminal (computer) equipped with a microphone function and a speaker function, used by a field worker (it is assumed that field worker 11A wears wearable terminal 10A, field worker 11B wears wearable terminal 10B, and field worker 11C wears wearable terminal 10C. Hereinafter, field workers 11A, 11B, and 11C are collectively referred to as field workers 11). For example, the wearable terminal 10 is a neck-worn terminal. The wearable terminal 10 can send and receive data to and from the two-way communication management device 20 via any network.
[0032] <<Two-way communication management device>> The two-way communication management device 20 is a device that manages at least audio two-way communication between multiple terminals (wearable terminals 10A, 10B, and 10C and remote supporter terminal 30 in the example of FIG. 1). The two-way communication management device 20 is composed of one or more computers.
[0033] <<Remote supporter terminal>> The remote supporter terminal 30 is a terminal used by the remote supporter 31. For example, the remote supporter 31 is a person who supports the work of the on-site worker 11 at a remote location away from the site where the on-site worker 11 is working. For example, the remote supporter terminal 30 is a personal computer. The remote supporter terminal 30 can send and receive data with the two-way communication management device 20 via any network.
[0034] <<Automatic answering machine>> The automatic answering device 40 uses artificial intelligence (AI) or the like to recognize the voice of the field worker 11 wearing the wearable device 10, and generates a voice that indicates a response (e.g., an answer to the question) to the content of the voice (e.g., a question). The automatic answering device 40 is composed of one or more computers. The automatic answering device 40 and the two-way communication management device 20 may be implemented in a single device.
[0035] <Wearable device configuration> FIG. 2 shows an example of a wearable terminal 10 according to an embodiment of the present disclosure. For example, the wearable terminal 10 is a terminal shaped like a neck strap as shown in FIG. 2. For example, the wearable terminal 10 includes a sound input unit (microphone) 105, a sound output unit (speaker) 106, and an earphone jack 100 as shown in FIG. 2. Furthermore, the wearable terminal 10 may include an operation unit 104 and an imaging unit (camera) 107 as shown in FIG. Each unit will be described in detail with reference to FIG. 3.
[0036] <Hardware configuration> An example of the hardware configuration of the wearable terminal 10 will be described below with reference to FIG. 3, and an example of the hardware configuration of the two-way communication management device 20 will be described with reference to FIG.
[0037] 3 is a diagram illustrating a hardware configuration of a wearable terminal 10 according to an embodiment of the present disclosure. The wearable terminal 10 can include a control unit (processor) 101, a storage unit (memory) 102, a communication unit 103, an operation unit 104, a sound input unit (microphone) 105, a sound output unit (speaker) 106, an imaging unit (camera) 107, and various sensors 108. Each of these will be described below.
[0038] The control unit (processor) 101 is a processor that controls the wearable terminal 10. For example, the control unit (processor) 101 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0039] The storage unit (memory) 102 is a memory that stores any data.
[0040] The communication unit 103 connects to any network and communicates with other computers (such as the two-way communication management device 20).
[0041] The operation unit 104 is a button or the like that allows the field worker 11 to input instructions to the wearable terminal 10.
[0042] The sound input unit (microphone) 105 collects sound.
[0043] The sound output unit (speaker) 106 outputs sound (for example, sound collected by another terminal and acquired via the two-way communication management device 20).
[0044] An imaging unit (camera) 107 captures still images and moving images.
[0045] The various sensors 108 are one or more arbitrary sensors such as an acceleration sensor, a GPS (Global Positioning System), and the like.
[0046] 4 is a diagram showing the hardware configuration of the two-way communication management device 20 according to an embodiment of the present disclosure. The same applies to the remote supporter terminal 30 and the automatic response device 40. The two-way communication management device 20 can include a control unit (processor) 201, a storage unit (memory) 202, and a communication unit 203. Each of these will be described below.
[0047] The control unit (processor) 201 is a processor that controls the two-way communication management device 20. For example, the control unit (processor) 201 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0048] The storage unit (memory) 202 is a memory that stores any data.
[0049] The communication unit 203 connects to any network and communicates with other computers (such as the wearable terminal 10 and the remote supporter terminal 30).
[0050] FIG. 5 is a functional block diagram of a two-way communication system 1 according to an embodiment of the present disclosure.
[0051] For example, the wearable terminal 10 includes a sound receiving unit 111, a sound collecting unit 112, and a sound transmitting unit 113. For example, the wearable terminal 10 functions as the sound receiving unit 111, the sound collecting unit 112, and the sound transmitting unit 113 by executing a program.
[0052] For example, the two-way communication management device 20 includes a voice receiving unit 211, a determination unit 212, a speech speed conversion unit 213, and a voice transmitting unit 214. For example, the two-way communication management device 20 functions as the voice receiving unit 211, the determination unit 212, the speech speed conversion unit 213, and the voice transmitting unit 214 by executing a program.
[0053] For example, the remote supporter terminal 30 includes a voice receiving unit 311, a voice collecting unit 312, and a voice transmitting unit 313. For example, the remote supporter terminal 30 functions as the voice receiving unit 311, the voice collecting unit 312, and the voice transmitting unit 313 by executing a program.
[0054] For example, the automatic answering device 40 includes a voice receiving unit 411, a voice recognition unit 412, a voice generation unit 413, and a voice transmission unit 414. For example, the automatic answering device 40 functions as the voice receiving unit 411, the voice recognition unit 412, the voice generation unit 413, and the voice transmission unit 414 by executing a program.
[0055] Note that Figure 5 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0056] [Wearable devices] The audio receiving unit 111 receives audio from the two-way communication management device 20 and reproduces it through a speaker or the like.
[0057] The sound collection unit 112 acquires the sound (that is, the sound collected by the microphone) uttered by the person wearing the wearable device 10 (that is, the field worker 11).
[0058] The voice transmitting unit 113 transmits the voice to the two-way communication management device 20.
[0059] [Two-way communication management device] The voice receiving unit 211 receives voice from the wearable terminal 10, the remote supporter terminal 30, and the automatic response device 40.
[0060] The decision unit 212 decides whether or not to change the speech rate of the voice.
[0061] The speech speed conversion unit 213 changes the speech speed of the voice (specifically, decreases or increases the speech speed of the voice).
[0062] As described above, the wearable terminal 10 and the remote supporter terminal 30 may convert the speech speed of the voice (that is, the wearable terminal 10 and the remote supporter terminal 30 may include the determination unit 212 and the speech speed conversion unit 213).
[0063] The voice transmitting unit 214 transmits the voice to the wearable terminal 10, the remote supporter terminal 30, and the automatic response device 40.
[0064] [Remote supporter terminal] The audio receiving unit 311 receives audio from the two-way communication management device 20 and reproduces it through a speaker or the like.
[0065] The voice collection unit 312 acquires the voice uttered by the user of the remote supporter terminal 30 (that is, the remote supporter 31) (that is, the voice collected by the microphone).
[0066] The voice transmitting unit 313 transmits the voice to the two-way communication management device 20 .
[0067] [Automatic answering machine] The voice receiving unit 411 receives voice from the two-way communication management device 20 .
[0068] The voice recognition unit 412 performs voice recognition of the voice received by the voice receiving unit 411. For example, the voice recognition unit 412 uses AI or the like to perform voice recognition of the voice received by the voice receiving unit 411 (for example, the voice of the field worker 11 wearing the wearable terminal 10 asking a question) and converts the voice into text.
[0069] The voice generation unit 413 generates a voice indicating a response to the result of voice recognition by the voice recognition unit 412. For example, the voice generation unit 413 uses AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to text (e.g., a question) that is the result of voice recognition, and generates a voice of the text indicating the response (i.e., performs voice synthesis).
[0070] The voice transmitting unit 414 transmits the voice generated by the voice generating unit 413 to the two-way communication management device 20.
[0071] Here, an example of changing the speech rate of the voice will be described. For example, the speech rate of the voice can be changed in response to an instruction from a user of one of the terminals (for example, the remote supporter terminal 30) or based on the judgment of the two-way communication management device 20. As examples of changing the speech rate of the voice based on the judgment of the two-way communication management device 20, "speech rate conversion based on the user's emotions," "speech rate conversion based on a comparison with the normal speech rate," and "speech rate conversion based on the remaining time of the meeting or work" will be described.
[0072] [Speech speed conversion based on user emotions] For example, the speech rate can be changed based on the result of analyzing the emotion of the person wearing the wearable terminal 10 (i.e., the field worker 11). For example, the determination unit 212 of the two-way communication management device 20 analyzes the emotion of the person (i.e., the field worker 11) who has uttered the voice (i.e., the voice collected by the microphone) uttered by the person wearing the wearable terminal 10, and when the emotion is a specific emotion (e.g., impatience), determines that the speech rate of the voice of the other party in the conversation (i.e., the remote supporter 31) or the speech rate of the voice generated by the automatic response device 40 should be changed (e.g., lower the speech rate).
[0073] [Speech speed conversion based on comparison with normal speaking speed] For example, the speech rate can be changed based on the difference between the speech rate of the voice collected by the wearable terminal 10 and the normal speech rate of the person wearing the wearable terminal 10 (i.e., the field worker 11). For example, when the speech rate of the person wearing the wearable terminal 10 (i.e., the field worker 11) is faster than normal, the determination unit 212 of the two-way communication management device 20 determines that the speech rate of the voice of the other party in the conversation (i.e., the remote supporter 31) or the speech rate of the voice generated by the automatic answering device 40 should be changed (e.g., to lower the speech rate). Also, for example, when the speech rate of the person wearing the wearable terminal 10 (i.e., the field worker 11) is slower than normal, the determination unit 212 of the two-way communication management device 20 determines that the speech rate of the voice of the other party in the conversation (i.e., the remote supporter 31) or the speech rate of the voice generated by the automatic answering device 40 should be changed (e.g., to increase the speech rate).
[0074] [Adjust speech rate based on remaining meeting or task time] For example, the speech rate can be changed based on the remaining time of a conference or task via two-way voice communication. For example, when the determination unit 212 of the two-way communication management device 20 detects that the remaining time of the conference or task has fallen below a threshold, it determines that the speech rate of the voices of all terminals and the speech rate of the voice generated by the automatic answering device 40 should be changed (for example, the speech rate should be increased).
[0075] <Processing method> A method for speech speed conversion processing will be described below with reference to Figures 6 to 8. Note that speech speed conversion may be repeated to control the speed of a speaker's speech (for example, to gradually slow down or gradually speed up).
[0076] Fig. 6 is a sequence diagram (in the case of a remote supporter terminal) of speech speed conversion processing based on emotion analysis according to an embodiment of the present disclosure. Fig. 6 illustrates an example in which speech speed conversion is performed on speech collected by the remote supporter terminal 30 based on the result of analyzing the emotion of the user of the wearable terminal 10 (i.e., the field worker 11). However, speech speed conversion may also be performed on speech collected by a terminal other than an arbitrary terminal based on the result of analyzing the emotion of the user of the terminal.
[0077] Note that Figure 6 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0078] In step 101 (S101), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (that is, the field worker 11).
[0079] In step 102 (S102), the wearable terminal 10 transmits the voice collected in S101 to the two-way communication management device 20.
[0080] In step 103 (S103), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S102.
[0081] In step 104 (S104), the two-way communication management device 20 transmits the voice received in S103 to the remote supporter terminal 30 (that is, transmits it to the remote supporter 31 without speech speed conversion).
[0082] In step 105 (S105), the remote supporter terminal 30 receives the voice transmitted by the two-way communication management device 20 in S104.
[0083] In step 106 (S106), the remote supporter terminal 30 outputs the audio received in S105 (for example, plays it on a speaker).
[0084] In step 107 (S107), the two-way communication management device 20 analyzes the emotion of the person who made the voice (i.e., the field worker 11) based on the voice received in S103. For example, the two-way communication management device 20 can analyze the emotion based on the frequency, volume, quality, and speaking speed of the voice. Note that the two-way communication management device 20 may also analyze the emotion of the person wearing the wearable terminal 10 based on something other than the voice (for example, based on biometric information). The two-way communication management device 20 saves the result of the emotion analysis.
[0085] S107 may be executed simultaneously with S104 to S106 or before S104 to S106.
[0086] In step 108 (S108), the remote supporter terminal 30 collects the voice uttered by the user of the remote supporter terminal 30 (that is, the remote supporter 31).
[0087] In step 109 (S109), the remote supporter terminal 30 transmits the voice collected in S108 to the two-way communication management device 20.
[0088] In step 110 (S110), the two-way communication management device 20 receives the voice transmitted by the remote supporter terminal 30 in S109.
[0089] In step 111 (S111), the two-way communication management device 20 converts the speech rate of the voice received in S110. Specifically, the two-way communication management device 20 reads the result of the emotion analysis in S107 and converts the speech rate of the voice based on the result. For example, if the emotion is a specific emotion, the two-way communication management device 20 changes the speech rate of the voice (for example, if the emotion is impatience (for example, panic, losing one's usual composure, being flustered, being confused, etc.), the speech rate is reduced).
[0090] In step 112 (S112), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed (for example, whose speech rate has been reduced) in S111.
[0091] In step 113 (S113), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S112.
[0092] In step 114 (S114), the wearable device 10 outputs (for example, plays on a speaker) the speech received in S113 with the speech rate changed (for example, the speech rate slowed down). Therefore, for example, the field worker 11 who has been speaking quickly due to impatience will be influenced by the speech of the remote supporter 31 whose speech rate has been slowed down and will also start speaking slowly.
[0093] 7 is a sequence diagram (in the case of a remote supporter terminal) of speech speed conversion processing based on a comparison with a normal speech speed according to an embodiment of the present disclosure. In FIG. 7, an example is described in which speech speed conversion is performed on speech collected by the remote supporter terminal 30 based on the result of comparison with the normal speech speed of the user of the wearable terminal 10 (i.e., the field worker 11). However, speech speed conversion may also be performed on speech collected by a terminal other than a given terminal based on the result of comparison with the normal speech speed of the user of the given terminal.
[0094] Note that Figure 7 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0095] In step 200 (S200), the two-way communication management device 20 stores information about the normal speaking speed of the person wearing the wearable terminal 10 (that is, the field worker 11).
[0096] In step 201 (S201), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (that is, the field worker 11).
[0097] In step 202 (S202), the wearable terminal 10 transmits the voice collected in S201 to the two-way communication management device 20.
[0098] In step 203 (S203), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S202.
[0099] In step 204 (S204), the two-way communication management device 20 transmits the voice received in S203 to the remote supporter terminal 30 (that is, transmits it to the remote supporter 31 without speech speed conversion).
[0100] In step 205 (S205), the remote supporter terminal 30 receives the voice transmitted by the two-way communication management device 20 in S204.
[0101] In step 206 (S206), the remote supporter terminal 30 outputs the sound received in S205 (for example, plays it on a speaker).
[0102] In step 207 (S207), the two-way communication management device 20 compares the speech rate of the voice received in S203 with the normal speech rate (the normal speech rate saved in S200) of the person who made the voice (i.e., the field worker 11). The two-way communication management device 20 saves the result of the comparison with the normal speech rate.
[0103] S207 may be executed simultaneously with S204 to S206 or before S204 to S206.
[0104] In step 208 (S208), the remote supporter terminal 30 collects the voice uttered by the user of the remote supporter terminal 30 (that is, the remote supporter 31).
[0105] In step 209 (S209), the remote supporter terminal 30 transmits the voice collected in S208 to the two-way communication management device 20.
[0106] In step 210 (S210), the two-way communication management device 20 receives the voice transmitted by the remote supporter terminal 30 in S209.
[0107] In step 211 (S211), the two-way communication management device 20 converts the speech rate of the voice received in S210. Specifically, the two-way communication management device 20 reads the result of the comparison with the normal speech rate in S207, and converts the speech rate of the voice based on the result. For example, if the speech rate is faster than normal, the two-way communication management device 20 reduces the speech rate of the voice. Also, for example, if the speech rate is slower than normal, the two-way communication management device 20 increases the speech rate of the voice.
[0108] In step 212 (S212), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed in S211.
[0109] In step 213 (S213), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S212.
[0110] In step 214 (S214), the wearable device 10 outputs the voice with the changed speaking speed received in S213 (for example, by playing it through a speaker). Therefore, for example, the on-site worker 11 who has been speaking faster than usual will be influenced by the voice of the remote supporter 31 whose speaking speed has been slowed down and will also start speaking slower. Also, for example, the on-site worker 11 who has been speaking slower than usual will be influenced by the voice of the remote supporter 31 whose speaking speed has been increased and will also start speaking faster.
[0111] FIG. 8 is a sequence diagram of a speech speed conversion process based on the remaining time of a meeting or task according to an embodiment of the present disclosure (in the case of a remote supporter terminal).
[0112] Note that Figure 8 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0113] In step 301 (S301), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (that is, the field worker 11).
[0114] In step 302 (S302), the wearable terminal 10 transmits the voice collected in S301 to the two-way communication management device 20.
[0115] In step 303 (S303), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S302.
[0116] In step 304 (S304), the two-way communication management device 20 detects the remaining time of the conference or work by two-way voice communication. If speech speed conversion is necessary (for example, if the remaining time is equal to or less than the threshold), the process proceeds to step 305. If speech speed conversion is not necessary (for example, if the remaining time is not equal to or less than the threshold), the process proceeds to step 306.
[0117] In step 305 (S305), the two-way communication management device 20 changes the speech rate of the voice received in S303 (for example, increases the speech rate).
[0118] In step 306 (S306), the two-way communication management device 20 transmits to the remote supporter terminal 30 the voice whose speech rate has been changed (for example, increased) in S305 (if speech rate conversion is required), or the voice received in S303 (if speech rate conversion is not required).
[0119] In step 307 (S307), the remote supporter terminal 30 receives the voice transmitted by the two-way communication management device 20 in S306.
[0120] In step 308 (S308), the remote supporter terminal 30 outputs (for example, plays on a speaker) the voice with the changed speech rate (for example, increased speech rate) received in S307. As a result, the remote supporter 31, influenced by the voice of the field worker 11 with the increased speech rate, also starts to speak faster (as a result, the meeting or work can be completed within the remaining time).
[0121] In step 309 (S309), the remote supporter terminal 30 collects the voice uttered by the person using the remote supporter terminal 30 (that is, the field worker 11).
[0122] In step 310 (S310), the remote supporter terminal 30 transmits the voice collected in S309 to the two-way communication management device 20.
[0123] In step 311 (S311), the two-way communication management device 20 receives the voice transmitted by the remote supporter terminal 30 in S310.
[0124] In step 312 (S312), the two-way communication management device 20 detects the remaining time of the conference or work by two-way voice communication. If speech speed conversion is necessary (for example, if the remaining time is equal to or less than the threshold), the process proceeds to step 313, and if speech speed conversion is not necessary (for example, if the remaining time is not equal to or less than the threshold), the process proceeds to step 314.
[0125] In step 313 (S313), the two-way communication management device 20 changes the speech rate of the voice received in S311 (for example, increases the speech rate).
[0126] In step 314 (S314), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed (for example, increased) in S313 (if speech rate conversion is required), or the voice received in S311 (if speech rate conversion is not required).
[0127] In step 315 (S315), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S314.
[0128] In step 316 (S316), the wearable device 10 outputs (for example, plays through a speaker) the speech received in S315 with the speech rate changed (for example, increased). As a result, the on-site worker 11, influenced by the speech of the remote supporter 31 with the increased speech rate, also starts speaking faster (as a result, the meeting or work can be completed within the remaining time).
[0129] A method for converting the speech speed of a voice generated by an automatic answering device will be described below with reference to Figures 9 to 11. Note that the speech speed of the speaker may be controlled (for example, by gradually slowing down or gradually speeding up) by repeating the speech speed conversion.
[0130] For example, speech speed conversion is used when the automatic answering device 40 generates an answer to a question from the field worker 11 and causes the wearable device 10 of the field worker 11 to output a voice indicating the answer. Also, speech speed conversion is used when the automatic answering device 40 analyzes a real-time video of a work operation captured by the wearable device 10 of the field worker 11 and causes the wearable device 10 of the field worker 11 to output work instructions based on the results of the analysis by voice (provide voice guidance). Also, speech speed conversion is used when the automatic answering device 40 causes the wearable device 10 of the field worker 11 to output work procedures described in a work procedure manual by voice (provide voice guidance).
[0131] 9 is a sequence diagram (in the case of an automatic answering device) of a speech rate conversion process based on emotion analysis according to an embodiment of the present disclosure. In FIG. 9, an example is described in which the speech rate of the voice generated by the automatic answering device 40 is converted based on the result of analyzing the emotion of the user of the wearable device 10 (i.e., the field worker 11).
[0132] Note that Figure 9 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0133] In step 1001 (S1001), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (that is, the field worker 11).
[0134] In step 1002 (S1002), the wearable terminal 10 transmits the voice collected in S1001 to the two-way communication management device 20.
[0135] In step 1003 (S1003), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S1002.
[0136] In step 1004 (S1004), the two-way communication management device 20 transmits the voice received in S1003 to the automatic answering device 40 (that is, transmits the voice to the automatic answering device 40 without speech speed conversion).
[0137] In step 1005 (S1005), the automatic answering device 40 receives the voice transmitted by the two-way communication management device 20 in S1004.
[0138] In step 1006 (S1006), the automatic answering device 40 performs speech recognition on the voice received in S1005. For example, the automatic answering device 40 uses AI or the like to perform speech recognition on the voice received in S1005 (for example, the voice of the field worker 11 wearing the wearable terminal 10 asking a question) and converts the voice into text.
[0139] In step 1007 (S1007), the two-way communication management device 20 analyzes the emotion of the person who made the voice (i.e., the field worker 11) based on the voice received in S1003. For example, the two-way communication management device 20 can analyze the emotion based on the frequency, volume, quality, and speaking speed of the voice. Note that the two-way communication management device 20 may also analyze the emotion of the person wearing the wearable terminal 10 based on something other than the voice (for example, based on biometric information). The two-way communication management device 20 saves the result of the emotion analysis.
[0140] Note that S1007 may be executed simultaneously with S1004 to S1006 or before S1004 to S1006.
[0141] In step 1008 (S1008), the automatic answering device 40 generates a voice indicating a response to the result of the voice recognition in S1006. For example, the automatic answering device 40 uses an AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to the text (e.g., a question) that is the result of the voice recognition, and generates a voice of the text indicating the response (i.e., performs voice synthesis).
[0142] In step 1009 (S1009), the automatic answering device 40 transmits the voice generated in S1008 to the two-way communication management device 20.
[0143] In step 1010 (S1010), the two-way communication management device 20 receives the voice transmitted by the automatic answering device 40 in S1009.
[0144] In step 1011 (S1011), the two-way communication management device 20 converts the speech rate of the voice received in S1010. Specifically, the two-way communication management device 20 reads the result of the emotion analysis in S1007 and converts the speech rate of the voice based on the result. For example, if the emotion is a specific emotion, the two-way communication management device 20 changes the speech rate of the voice (for example, if the emotion is impatience (for example, panic, losing one's usual composure, being flustered, being confused, etc.), the speech rate is reduced).
[0145] In step 1012 (S1012), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed (for example, whose speech rate has been reduced) in S1011.
[0146] In step 1013 (S1013), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S1012.
[0147] In step 1014 (S1014), the wearable device 10 outputs (for example, plays through a speaker) the voice with the speech rate changed (for example, slowed down) received in S1013. Therefore, for example, the field worker 11 who has been speaking quickly due to impatience will be persuaded by the voice of the automatic answering device 40 with the speech rate slowed down and will also start speaking slowly.
[0148] 10 is a sequence diagram (in the case of an automatic answering device) of a speech speed conversion process based on a comparison with a normal speech speed according to an embodiment of the present disclosure. In FIG. 10, an example of converting the speech speed of the voice generated by the automatic answering device 40 based on the result of a comparison with the normal speech speed of the user of the wearable device 10 (i.e., the field worker 11) is described.
[0149] Note that Figure 10 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0150] In step 2000 (S2000), the two-way communication management device 20 stores information about the normal speaking speed of the person wearing the wearable terminal 10 (that is, the field worker 11).
[0151] In step 2001 (S2001), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (that is, the field worker 11).
[0152] In step 2002 (S2002), the wearable terminal 10 transmits the voice collected in S2001 to the two-way communication management device 20.
[0153] In step 2003 (S2003), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S2002.
[0154] In step 2004 (S2004), the two-way communication management device 20 transmits the voice received in S2003 to the automatic answering device 40 (that is, transmits the voice to the automatic answering device 40 without speech speed conversion).
[0155] In step 2005 (S2005), the automatic answering device 40 receives the voice transmitted by the two-way communication management device 20 in S2004.
[0156] In step 2006 (S2006), the automatic answering device 40 performs speech recognition on the voice received in S2005. For example, the automatic answering device 40 uses AI or the like to perform speech recognition on the voice received in S2005 (for example, the voice of the field worker 11 wearing the wearable terminal 10 asking a question) and converts the voice into text.
[0157] In step 2007 (S2007), the two-way communication management device 20 compares the speech rate of the voice received in S2003 with the normal speech rate (the normal speech rate saved in S2000) of the person who made the voice (i.e., the field worker 11). The two-way communication management device 20 saves the result of the comparison with the normal speech rate.
[0158] Note that S2007 may be executed simultaneously with S2004 to S2006 or before S2004 to S2006.
[0159] In step 2008 (S2008), the automatic answering device 40 generates a voice indicating a response to the result of the voice recognition in S2006. For example, the automatic answering device 40 uses an AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to the text (e.g., a question) that is the result of the voice recognition, and generates a voice of the text indicating the response (i.e., performs voice synthesis).
[0160] In step 2009 (S2009), the automatic answering device 40 transmits the voice generated in S2008 to the two-way communication management device 20.
[0161] In step 2010 (S2010), the two-way communication management device 20 receives the voice transmitted by the automatic answering device 40 in S2009.
[0162] In step 2011 (S2011), the two-way communication management device 20 converts the speech rate of the voice received in S2010. Specifically, the two-way communication management device 20 reads the result of the comparison with the normal speech rate in S2007, and converts the speech rate of the voice based on the result. For example, if the speech rate is faster than normal, the two-way communication management device 20 reduces the speech rate of the voice. Also, for example, if the speech rate is slower than normal, the two-way communication management device 20 increases the speech rate of the voice.
[0163] In step 2012 (S2012), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed in S2011.
[0164] In step 2013 (S2013), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S2012.
[0165] In step 2014 (S2014), the wearable device 10 outputs the voice with the changed speaking speed received in S2013 (for example, by playing it through a speaker). Therefore, for example, a field worker 11 who has been speaking faster than usual may be persuaded by the voice of the automatic answering device 40 with the slowed speaking speed to also speak more slowly. Also, for example, a field worker 11 who has been speaking more slowly than usual may be persuaded by the voice of the automatic answering device 40 with the increased speaking speed to also speak more quickly.
[0166] FIG. 11 is a sequence diagram of a speech speed conversion process based on the remaining time of a meeting or task according to an embodiment of the present disclosure (in the case of an automatic answering device).
[0167] Note that Figure 11 describes a case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0168] In step 3001 (S3001), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (that is, the field worker 11).
[0169] In step 3002 (S3002), the wearable terminal 10 transmits the voice collected in S3001 to the two-way communication management device 20.
[0170] In step 3003 (S3003), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S3002.
[0171] In step 3004 (S3004), the two-way communication management device 20 transmits the voice received in S3003 to the automatic answering device 40.
[0172] In step 3005 (S3005), the automatic answering device 40 receives the voice transmitted by the two-way communication management device 20 in S3004.
[0173] In step 3006 (S3006), the automatic answering device 40 performs speech recognition on the voice received in S3005. For example, the automatic answering device 40 uses AI or the like to perform speech recognition on the voice received in S3005 (for example, the voice of the field worker 11 wearing the wearable terminal 10 asking a question) and converts the voice into text.
[0174] In step 3007 (S3007), the automatic answering device 40 generates a voice indicating a response to the result of the voice recognition in S3006. For example, the automatic answering device 40 uses an AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to the text (e.g., a question) that is the result of the voice recognition, and generates a voice of the text indicating the response (i.e., performs voice synthesis).
[0175] In step 3008 (S3008), the automatic answering device 40 transmits the voice generated in S3007 to the two-way communication management device 20.
[0176] In step 3009 (S3009), the two-way communication management device 20 receives the voice transmitted by the automatic answering device 40 in S3008.
[0177] In step 3010 (S3010), the two-way communication management device 20 detects the remaining time of the conference or work by two-way voice communication. If speech speed conversion is necessary (for example, if the remaining time is equal to or less than the threshold), the process proceeds to step 3011. If speech speed conversion is not necessary (for example, if the remaining time is not equal to or less than the threshold), the process proceeds to step 3012.
[0178] In step 3011 (S3011), the two-way communication management device 20 changes the speech rate of the voice received in S3009 (for example, increases the speech rate).
[0179] In step 3012 (S3012), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed (for example, increased) in S3011 (if speech rate conversion is required), or the voice received in S309 (if speech rate conversion is not required).
[0180] In step 3013 (S3013), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S3012.
[0181] In step 3014 (S3014), the wearable device 10 outputs (for example, plays through a speaker) the voice with the changed speech rate (for example, increased speech rate) received in S3013. As a result, the field worker 11, attracted by the voice of the automatic answering device 40 with the increased speech rate, also starts to speak faster (as a result, the meeting or work can be completed within the remaining time).
[0182] Below, the speech section and the non-speech section will be described with reference to FIG. 12, a method for lowering the speech rate will be described with reference to FIG. 13, and a method for increasing the speech rate will be described with reference to FIG.
[0183] 12 to 14, the shaded areas of "own speech" indicate that the user of the own terminal is speaking (i.e., the sound is being collected by the own terminal) during the "elapsed time (milliseconds)." The speech includes a section where speech (voice) is detected (a speech section, also called a voice section) and a section where speech (voice) is not detected (a non-speech section, also called a non-voice section).
[0184] The shaded areas of "Communication & Processing" in Figures 12 to 14 indicate that communication and processing are taking place from one's own terminal to a terminal other than one's own terminal (the other terminal) during the "Elapsed Time (milliseconds)" (for example, the two-way communication management device 20 is performing the processing).
[0185] The shaded areas of "Playback by other party" in FIGS. 12 to 14 indicate that the voice of the user of the user's own terminal is being played back at the other party's terminal during the "Elapsed time (milliseconds)".
[0186] Figure 12 shows a case where both the speech section and the non-speech section of "your own utterance" are sent to the other party's terminal without extracting the speech section from "your own utterance" (in Figure 12, both the speech section "aaa" and the non-speech section "aa" are sent to the other party's terminal and played back on the other party's terminal).
[0187] In one embodiment of the present disclosure, a speech section may be detected using VAD (Voice Activity Detection), and only the speech section may be sent to the other terminal. Hereinafter, with reference to Fig. 13, a case where the speech rate of only the speech section is lowered and sent to the other terminal (i.e., the non-speech section is not sent to the other terminal) will be described, and with reference to Fig. 14, a case where the speech rate of only the speech section is increased and sent to the other terminal (i.e., the non-speech section is not sent to the other terminal) will be described.
[0188] 13 is a diagram illustrating a method for slowing down the speech rate according to an embodiment of the present disclosure. As shown in FIG. 13, when a speech section ends and a non-speech section begins, the speech rate of the speech section is slowed down and sent to the other party's terminal for playback (i.e., the non-speech section is not sent to the other party's terminal). The process of slowing down the speech rate of the speech section may be performed by the two-way communication management device 20 in "Communication & Processing" in FIG. 13, by the user's own terminal, or by the other party's terminal.
[0189] 14 is a diagram illustrating a method for increasing the speech rate according to an embodiment of the present disclosure. As shown in FIG. 14, when a speech section ends and a non-speech section begins, the speech rate of the speech section is increased and sent to the other party's terminal for playback (i.e., the non-speech section is not sent to the other party's terminal). The process of increasing the speech rate of the speech section may be performed by the two-way communication management device 20 in "Communication & Processing" in FIG. 14, by the user's own terminal, or by the other party's terminal.
[0190] However, if the speech rate is increased too much, the speech of two speech periods will be played back consecutively without any non-speech period, making it difficult for the listener to hear. Therefore, a buffer (a buffer of a predetermined minimum length) may be provided between the speech of the two speech periods (i.e., a non-speech period may be provided).
[0191] Although the embodiments have been described above, it will be understood that various changes in form and details can be made without departing from the spirit and scope of the claims. [Explanation of symbols]
[0192] 1. Two-way communication system 10A Wearable Device 10B Wearable Device 10C Wearable Device 11A Field worker 11B Field worker 11C Field worker 20 Two-way communication management device 30 Remote supporter terminal 31 Remote Supporter 40 Auto Answering Machine 101 Control unit (processor) 102 Memory 103 Communications Department 104 Operation section 105 Sound input unit (microphone) 106 Sound output unit (speaker) 107 Imaging unit (camera) 108 Various sensors 100 earphone jack 111 Audio receiving unit 112 Audio pickup unit 113 Audio transmission unit 201 Control unit (processor) 202 Memory 203 Communications Department 211 Audio receiving unit 212 Judgment Department 213 Speech Speed Conversion Unit 214 Audio transmission unit 311 Audio receiving unit 312 Audio pickup unit 313 Audio transmission unit 411 Audio receiving unit 412 Speech Recognition Unit 413 Speech Generation Unit 414 Audio transmission unit
Claims
1. A two-way communication system for two-way audio communication between a plurality of terminals, changing the speech rate of the voice collected by at least one terminal or the voice generated by the automatic answering device; The speech whose speech rate has been changed is transmitted to a terminal other than the terminal from which the speech is collected. Two-way communication system.
2. The two-way communication system according to claim 1, wherein the speech rate of the voice collected by at least one of the terminals or the voice generated by the automatic answering device is changed in response to an instruction from a user of any of the terminals or based on a judgment of the two-way communication system.
3. 3. The two-way communication system according to claim 2, wherein the speech rate of the voice collected at a terminal other than the user's terminal or the voice generated by the automatic answering device is changed based on the results of analyzing the emotions of the user of the terminal at which the voice is collected.
4. 4. The two-way communication system according to claim 3, wherein when the emotion of the user of the terminal from which the voice is collected is impatience, the speaking speed of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is reduced.
5. 3. The two-way communication system according to claim 2, wherein the speech speed of the speech collected at a terminal other than the user's terminal or the speech generated by the automatic answering device is changed based on the difference between the normal speech speed of the user of the terminal from which the speech is collected and the speech speed of the speech collected at the terminal from which the speech is collected.
6. 6. The two-way communication system according to claim 5, wherein, when the speech speed of the speech collected at the terminal collecting the speech is faster than the normal speech speed of the user of the terminal collecting the speech, the speech speed of the speech collected at a terminal other than the terminal of the user or the speech generated by the automatic answering device is lowered.
7. 6. The two-way communication system according to claim 5, wherein, when the speech speed of the speech collected at the terminal collecting the speech is slower than the normal speech speed of the user of the terminal collecting the speech, the speech speed of the speech collected at a terminal other than the terminal of the user or the speech generated by the automatic answering device is increased.
8. 3. The two-way communication system according to claim 2, wherein the speech rate of the voice collected by the at least one terminal or the voice generated by the automatic answering machine is changed based on the remaining time of the conference or work performed by the two-way voice communication.
9. The two-way communication system according to claim 8, wherein when the remaining time of the conference or work via the two-way voice communication falls below a threshold, the speech rate of the voice collected by all terminals and the voice generated by the automatic answering device is increased.
10. A method executed by a two-way communication system for two-way audio communication between a plurality of terminals, comprising: changing the speech rate of the speech collected by at least one terminal or generated by an automatic answering machine; transmitting the speech whose speech rate has been changed to a terminal other than the terminal from which the speech was collected; A method comprising:
11. On the computer, A step of changing the speech rate of a voice collected by at least one terminal among a plurality of terminals that perform two-way voice communication or a voice generated by an automatic answering device; a step of transmitting the speech whose speech rate has been changed to a terminal other than the terminal from which the speech is collected; A program to execute.
Citation Information
Patent Citations
Voice interactive device, input voice optimizing method in the device and input voice optimizing processing program in the device
JP2003150194A
Apparatus and method for speech processing, recording medium, and program
JP2004258290A
Voice communication apparatus and voice communication system
JP2008032933A
Automatic voice recognition / voice conversion system
JP2014095753A
Aviation control voice communication device and voice processing method
JP2014228691A