Information presentation device, information presentation method, and information presentation program
The information presentation device addresses the challenge of controlling speaking speed and sound pressure in online communication by using pseudo-heartbeat sounds and bubble noise to adjust these elements, resulting in improved speech clarity for listeners.
Patent Information
- Application Number
- JP2023576314
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-01-26
AI Technical Summary
In online voice communication, users face challenges in controlling paralinguistic elements such as speaking speed and sound pressure, leading to difficulties in ensuring that their speech is easily conveyed to listeners.
An information presentation device that acquires speaking speed and sound pressure at intervals, determines the need for improvement, and outputs pseudo-heartbeat sounds or bubble noise to involuntarily adjust these elements.
Enables users to make speech that is easier for listeners to understand by involuntarily and imperceptibly controlling speaking speed and sound pressure.
Smart Images

Figure 0007694719000001 
Figure 0007694719000002 
Figure 0007694719000003
Abstract
Description
Technical Field
[0001] This invention relates to an information presentation device, an information presentation method, and an information presentation program.
Background Art
[0002] In recent years, due to the influence of COVID-19 and the like, online voice communication such as web conferencing is becoming mainstream.
[0003] For example, in a face-to-face meeting, a user who is a speaker can receive feedback on whether their speech is being correctly conveyed to the listener while grasping the listener's expression and situation. Then, based on this feedback, the user speaks while controlling paralinguistic elements (e.g., speech rate, sound pressure (volume), etc.) during speaking. However, in online voice communication, since users have few opportunities to obtain feedback on their speech, it tends to be more difficult to determine whether they are making their speech easy to convey to the listener compared to the face-to-face format.
[0004] In Non-Patent Document 1, it is proposed to quantitatively evaluate the speech rate before a meeting and convey to the speaker whether it is an appropriate speech rate, aiming to improve the speech rate.
[0005] In Non-Patent Document 2, it is proposed to measure the sound pressure and measure the improvement of the sound pressure by displaying feedback on the screen according to the sound pressure.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] In Non-Patent Document 1, there is a problem that even if the speaking speed of oneself is fed back and it is indicated that the speaking speed is inappropriate, the improvement of the speaking speed is limited. Further, in Non-Patent Document 2, there is a problem that even if the shortage of the speaking volume (sound pressure) is fed back acoustically or visually in real time during a presentation, the effect of increasing the sound pressure is limited.
[0008] In addition, when the user is in a tense state, the speaking speed may become fast or the user may not be able to speak at an appropriate volume. In such a case, even if the user tries to control the paralanguage elements consciously, there is a problem that the user makes a speech that is difficult for the listener to understand.
[0009] This invention has been made paying attention to the above circumstances, and its object is to provide a technology for involuntarily and non-perceptually controlling the paralanguage elements of a user in an online meeting or the like so that the user can make a speech that is easy for the listener to understand.
Means for Solving the Problems
[0010] In order to solve the above problems, one aspect of this invention is an information presentation device, which includes an acquisition unit that acquires the speaking speed and sound pressure of speech at a predetermined interval, a determination unit that determines whether it is necessary to improve at least one of the speaking speed or the sound pressure, a pseudo-heartbeat sound acquisition unit that acquires a pseudo-heartbeat sound when it is determined that it is necessary to improve the speaking speed, a bubble noise acquisition unit that acquires bubble noise when it is determined that it is necessary to improve the sound pressure, and a presentation content control unit that outputs presentation information including at least one of the pseudo-heartbeat sound or the bubble noise.
Effects of the Invention
[0011] According to one aspect of the present invention, in an online meeting or the like, it becomes possible to involuntarily and unconsciously control the paralinguistic elements of a user and enable the user to make utterances that are easily conveyed to the listener.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the following, the same or similar elements as those already described will be given the same or similar reference numerals, and redundant descriptions will be basically omitted.
[0014] [Embodiment] (Configuration) FIG. 1 is a block diagram showing an example of the hardware configuration of an information presentation device 1 according to an embodiment. The information presentation device 1 may be, for example, a user terminal used by a user participating in an online meeting. Here, the user terminal may be any computer that can be generally used by a user, such as a PC (Personal Computer), a smartphone, a tablet terminal, a wearable terminal, or the like. Further, the information presentation device 1 may be a server to which the user terminal is connected via a network. Also, the server may be any computer that can be used as a server.
[0015] The information presentation device 1 includes a control unit 10, a program storage unit 20, a data storage unit 30, a communication interface 40, and an input / output interface 50. The control unit 10, the program storage unit 20, the data storage unit 30, the communication interface 40, and the input / output interface 50 are communicably connected to each other via a bus.
[0016] The information presentation device 1 is realized by a computer such as a PC (Personal Computer). The information presentation device 1 includes a control unit 10, a program storage unit 20, a data storage unit 30, a communication interface 40, and an input / output interface 50. The control unit 10, the program storage unit 20, the data storage unit 30, the communication interface 40, and the input / output interface 50 are communicably connected to each other via a bus. Further, the information presentation device 1 is connected to a device or server used by other users via a network 6.
[0017] The control unit 10 controls the information presentation device 1. The control unit 10 includes a hardware processor such as a central processing unit (CPU: Central Processing Unit).
[0018] The program storage unit 20 is configured by, for example, combining a non-volatile memory such as an SSD (Solid State Drive) that can be written to and read from at any time as a storage medium and a non-volatile memory such as a ROM (Read Only Memory). In addition to middleware such as an OS (Operating System), it stores application programs necessary for executing various control processes according to one embodiment. Hereinafter, the OS and each application program are collectively referred to as a program.
[0019] The data storage unit 30 may be, for example, a combination of a non-volatile memory such as an SSD that can be written to and read from at any time as a storage medium and a volatile memory such as a RAM (Random Access Memory).
[0020] The communication interface 40 includes one or more wired or wireless communication modules. For example, the communication interface 40 includes a communication module for wired or wireless connection with a device or server used by other users via the network 6. For example, the communication interface 40 may include a wireless communication module capable of wireless connection with a Wi-Fi access point or the like. That is, the communication interface 40 may be a general communication interface as long as it can communicate with a device or server used by other users under the control of the control unit 10 and transmit and receive various information.
[0021] The input / output interface 50 is connected to an input device 51, an output device 52, a voice input device 53, a voice output device 54, and the like. The input / output interface 50 is an interface that enables transmission and reception of information between the input device 51, the output device 52, the voice input device 53, and the voice output device 54. The input / output interface 50 may include a wired or wireless communication interface. For example, the information presentation device 1 and at least one of the input device 51, the output device 52, the voice input device 53, and the voice output device 54 are wirelessly connected using a short-range wireless technology or the like, and information may be transmitted and received using the short-range wireless technology.
[0022] The input device 51 includes, for example, a keyboard, a pointing device, or the like for an owner (e.g., a user) of the information presentation device 1 to input an instruction to the information presentation device 1. The input device 51 may also include a reader for reading data to be stored in the data storage unit 30 from a memory medium such as a USB memory, or a disk device for reading such data from a disk medium.
[0023] The output device 52 includes a display for displaying output data to be presented to the owner from the information presentation device 1, a printer for printing the same, and the like.
[0024] The voice input device 53 may be, for example, a microphone. That is, the voice input device 53 is arranged near the user (speaker) and collects the user's speech and converts it into an electrical signal.
[0025] The voice output device 54 may be an ear device, a speaker, or the like. The voice output device 54 is used to output voice information spoken by another user or voice information stored in the data storage unit 30 as voice. Here, the voice output device 54 may be a sound conduction type ear device such as earphones or headphones, or may be a bone conduction type ear device.
[0026] Furthermore, the voice input device 53 and the voice output device 54 may be a device in which the voice input device 53 and the voice output device 54 are integrated, such as a headset or a speakerphone.
[0027] FIG. 2 is a block diagram showing the software configuration of the information presentation device 1 in the embodiment in association with the hardware configuration shown in FIG. 1. The data storage unit 30 includes a state storage unit 301, a threshold storage unit 302, and a presentation content storage unit 303.
[0028] The state storage unit 301 is used to store voice information and the like corresponding to the voice spoken by the speaker acquired by the information acquisition unit 101 of the control unit 10 described later.
[0029] The threshold storage unit 302 is used to store the threshold used when the user state determination unit 102 of the control unit 10 determines whether it is necessary to improve the user's speech.
[0030] The presentation content storage unit 303 stores the pseudo heart sound and the bubble noise used by the presentation content control unit 103 of the control unit 10 to improve the user's speech.
[0031] The control unit 10 includes an information acquisition unit 101, a user state determination unit 102, and a presentation content control unit 103.
[0032] The information acquisition unit 101 receives voice information from the voice input device 53 and stores the received voice information in the state storage unit 301 at predetermined intervals. Further, the information acquisition unit 101 can calculate the speech rate based on the voice information stored in the state storage unit 301 and measure the sound pressure. That is, the information acquisition unit 101 can acquire the speech rate and sound pressure of the utterance at predetermined intervals. Furthermore, the information acquisition unit 101 stores the calculated speech rate and the measured sound pressure in the state storage unit 301.
[0033] The user state determination unit 102 acquires the speech rate and sound pressure stored in the state storage unit 301. Then, the user state determination unit 102 calculates the average speech rate and the average sound pressure. Further, the user state determination unit 102 acquires the speech rate threshold value and the sound pressure threshold value stored in the threshold value storage unit 302. Then, the user state determination unit 102 compares the average speech rate with the speech rate threshold value and the average sound pressure with the sound pressure threshold value, respectively, and determines whether it is necessary to improve at least one of the speech rate or the sound pressure.
[0034] When it is determined that it is necessary to improve at least one of the speech rate or the sound pressure, the presentation content control unit 103 outputs presentation information for improving the user's utterance. For example, when it is determined that it is necessary to improve the speech rate, the presentation content control unit 103 acquires the pseudo-heartbeat sound stored in the presentation content storage unit 303 and outputs the presentation information including the pseudo-heartbeat sound to the voice output device 54. Alternatively, when it is determined that it is necessary to improve the sound pressure, the presentation content control unit 103 acquires the bubble noise and outputs the presentation information including the bubble noise to the voice output device 54.
[0035] Here, it is known that when presenting a pseudo-heartbeat sound with a heart rate different from the user's own heartbeat to the user, the user's own heartbeat approaches the pseudo-heartbeat sound (for example, see Nakamura et al., A Control System for Biological Information Using False Information Feedback, Information Processing Society of Japan Interaction 2012 (2021), etc.). Therefore, in order to relieve the user's tension and improve the speech rate, the presentation content control unit 103 presents the pseudo-heartbeat sound to the user as presentation information. Also, it is known that hearing is the most effective when presenting feedback information to the user (for example, see Miyata et al., Biofeedback Therapy, New Physiological Psychology Volume 2 (1997), etc.). Therefore, the presentation content control unit 103 may output the presentation information to the audio output device 54.
[0036] Furthermore, it is known that in a noisy environment, the sound pressure and fundamental frequency of the voice increase in order to make it easier to convey one's own voice to the other party compared to a silent environment (for example, see Lane, H. L, The Lombard Sign and the Role of Hearing in Speech, etc.). Therefore, in order to improve the sound pressure of the user's speech, the presentation content control unit 103 presents bubble noise to the user as presentation information.
[0037] (Operation) Figure 3 is a flowchart showing an example of the information presentation operation of the information presentation device 1. The operation of this flowchart is realized by the control unit 10 of the information presentation device 1 reading and executing the program stored in the program storage unit 20.
[0038] The operation starts when the user (speaker) of the information presentation device 1 participates in something where the user needs to speak online, such as an online meeting. In the following description, for simplicity, an online meeting is assumed for the explanation, but it is of course not limited to an online meeting, etc. For example, if it is possible to record, analyze the user's speech, and provide feedback to the user, it can also be a face-to-face meeting, etc.
[0039] The information acquisition unit 101 of the control unit 10 acquires voice information and stores the acquired voice information in the state memory unit 301 (step ST101). The voice input device 53 converts the voice uttered by the user into voice information and outputs the converted voice information to the information acquisition unit 101. Then, the information acquisition unit 101 stores the speech rate and sound pressure acquired from the voice information in the state memory unit 301.
[0040] FIG. 4 is a flowchart showing an example for explaining the operation of step ST101 in more detail. The information acquisition unit 101 stores the voice information segmented at predetermined intervals in the state memory unit 301 (step ST201). The information acquisition unit 101 segments the voice information received from the voice input device 53 at predetermined intervals T, assigns a recording ID to each of the segmented voice information, and stores it in the state memory unit 301. Here, the predetermined interval T may be an arbitrary interval, for example, 1 minute (60000 ms).
[0041] The information acquisition unit 101 calculates the speech rate based on the voice information stored in the state memory unit 301 (step ST202). The information acquisition unit 101 acquires the voice information during a predetermined interval T stored in the state memory unit 301. Then, the information acquisition unit 101 calculates the number of characters N word contained in the acquired voice information. That is, the calculated number of characters N word is the number of characters contained during the predetermined interval T. Then, the information acquisition unit 101 calculates the speech rate S (words / minute) from N word / T.
[0042] The information acquisition unit 101 measures the sound pressure based on the voice information stored in the state memory unit 301 (step ST203). The information acquisition unit 101 measures the sound pressure P (dB) from the voice information during the predetermined interval T acquired in step ST202. Here, the sound pressure P may be the average sound pressure during the interval T. Since the sound pressure P may be measured by a general method, detailed description here is omitted.
[0043] The information acquisition unit 101 stores the calculated speech rate S and the measured sound pressure P in the state storage unit 301 (step ST204). For example, the information acquisition unit 101 stores the speech rate and the sound pressure in the state storage unit 301 in association with a recording ID. As described above, the information acquisition unit 101 acquires the speech rate and the sound pressure of speech at a predetermined interval.
[0044] FIG. 5 is a diagram showing an example of the speech rate and the sound pressure stored in the state storage unit 301. As shown in FIG. 5, a recording ID is assigned to each piece of voice information for each interval T, and a table in which the speech rate S (words / minute) and the sound pressure (dB) for each recording ID are recorded is stored in the state storage unit 301. For example, during the period when the recording ID is 1, it is shown that the user is speaking at a speech rate S of 300 (words / minute) and a sound pressure of 50 (dB). Also, for example, the speech rate S and the sound pressure P obtained from new voice information may be stored in a new row at the bottom of the table in FIG. 5.
[0045] Referring to FIG. 3, the user state determination unit 102 of the control unit 10 determines whether it is necessary to present improvement information to the user (step ST102). The user state determination unit 102 determines whether to present presentation information for improving the user's speech based on the speech rate S and the sound pressure P stored in the state storage unit 301.
[0046] FIG. 6 is a flowchart showing an example for explaining the operation of step ST102 in more detail. The user state determination unit 102 acquires the speech rate S and the sound pressure P stored in the state storage unit 301 (step ST301). For example, the user state determination unit 102 may acquire the speech rate S and the sound pressure P for M rows stored in the state storage unit 301. Here, M may be a plurality of arbitrary numbers defined in advance.
[0047] The user state determination unit 102 calculates an average speech rate S ave and an average sound pressure P ave based on the acquired speech rate S and sound pressure P (step ST302). The user state determination unit 102 calculates the M acquired speech rates S ifrom the average speech rate S ave is calculated. Here, i is an arbitrary integer from 1 to M, indicating the i-th speech rate among the M acquired ones. For example, the user state determination unit 102 calculates the average speech rate S ave based on the following formula. S ave =ΣS i / M (i = 1, ···, M) Similarly, the user state determination unit 102 calculates the average sound pressure P i from the M acquired sound pressures P ave . For example, the user state determination unit 102 calculates the average sound pressure P ave based on the following formula. P ave = 10log 10 (Σ10 Pi / 10 ) - 10log 10 N (i = 1, ···, M) The user state determination unit 102 acquires the speech rate threshold S th and the sound pressure threshold P th stored in the threshold storage unit 302 (step ST303).
[0048] FIG. 7 is a diagram showing an example of the speech rate threshold S th and the sound pressure threshold P th stored in the threshold storage unit 302. As shown in FIG. 7, the threshold storage unit 302 stores a speech rate threshold S th indicating that the speech rate is inappropriate, and a sound pressure threshold P th indicating that the sound pressure is inappropriate. Note that in FIG. 7, for simplicity, only one speech rate threshold and one sound pressure threshold are shown, but of course, the threshold storage unit 302 may store a plurality of speech rate thresholds and a plurality of sound pressure thresholds. The user state determination unit 102 acquires the speech rate threshold and the sound pressure threshold from the threshold storage unit 302.
[0049] The user state determination unit 102 determines whether the average speech rate S ave is greater than the speech rate threshold S th (step ST304). The user state determination unit 102 determines whether the average speech rate Save and the speech speed threshold S th are compared, and the average speech speed S ave is greater than the speech speed threshold S th is determined. If the average speech speed S ave is greater than the speech speed threshold S th is determined, the process proceeds to step ST305, and if the average speech speed S ave is determined to be less than or equal to the speech speed threshold S th the process proceeds to step ST306.
[0050] The user state determination unit 102 outputs a speech speed improvement signal to the presentation content control unit 103 (step ST305). In step ST304, if the average speech speed S ave is determined to be greater than the speech speed threshold S th it means that the user is speaking faster than a speed understandable by the listener. Therefore, the user state determination unit 102 outputs a speech speed improvement signal indicating that it is necessary to improve the user's speech speed to the presentation content control unit 103.
[0051] The user state determination unit 102 determines whether the average sound pressure P ave is less than the sound pressure threshold P th (step ST306). The user state determination unit 102 compares the average sound pressure P ave with the sound pressure threshold P th to determine whether the average sound pressure P ave is less than the sound pressure threshold P th . If the average sound pressure P ave is determined to be less than the sound pressure threshold P th the process proceeds to step ST307, and if the average sound pressure P ave is determined to be less than or equal to the sound pressure threshold P th the process ends.
[0052] The user state determination unit 102 outputs a sound pressure improvement signal to the presentation content control unit 103 (step ST307). In step ST306, if the average sound pressure P ave is less than the sound pressure threshold P thIf it is determined to be smaller, it means that the user is speaking at a volume that is difficult for the listener to hear. Therefore, the user state determination unit 102 outputs a sound pressure improvement signal indicating that it is necessary to improve the sound pressure of the user to the presentation content control unit 103.
[0053] Here, when performing the process of step ST305 or step ST307, it corresponds to the case where it is determined that presentation is necessary in step ST102 shown in FIG. 3. On the other hand, when it is determined in step ST306 that the average sound pressure P ave is greater than the sound pressure threshold P th it corresponds to the case where it is determined that presentation is unnecessary in step ST102 shown in FIG. 3.
[0054] In the example of FIG. 6, in order to prioritize the improvement of the speech rate, first, it is determined whether the average speech rate S ave is greater than the speech rate threshold S th However, the improvement of the sound pressure may be prioritized. For example, when the user's speech content cannot be heard because the sound pressure of the utterance is too low, the control unit 10 may control to prioritize the improvement of the sound pressure. In this case, it can be implemented by swapping step ST304 and step ST306.
[0055] Also, the user state determination unit 102 may determine that it is necessary to improve both the speech rate and the sound pressure. In this case, after determining in step ST304 that it is necessary to improve the speech rate, step ST306 may be changed to be implemented. And when it is determined that it is necessary to improve both the speech rate and the sound pressure, the user state determination unit 102 may output an improvement signal indicating that it is necessary to improve both the speech rate and the sound pressure to the presentation content control unit 103.
[0056] Furthermore, the threshold storage unit 302 may store the speech rate lower limit threshold S th_min And in step ST303, the user state determination unit 102 may also acquire the speech rate lower limit threshold S th_min And in step ST304, the user state determination unit 102 may determine the average speech rate S aveis less than the lower limit of the speaking speed threshold S th_min It may be determined whether it is smaller. And the average speaking speed S ave is less than the lower limit of the speaking speed threshold S th_min If it is determined that it is smaller, in step ST305, the user state determination unit 102 may output a speaking speed improvement signal to the presentation content control unit 103. For example, the user's speaking speed is slow, and it is conceivable that the listener will get bored or irritated. Therefore, the user state determination unit 102 may output a speaking speed improvement signal to the presentation content control unit 103.
[0057] Referring to FIG. 3, the presentation content control unit 103 of the control unit 10 acquires presentation information (step ST103). The presentation content control unit 103 acquires the presentation information stored in the presentation content storage unit 303 according to the content of the signal received from the user state determination unit 102.
[0058] FIG. 8 is a diagram showing an example of the presentation information stored in the presentation content storage unit 303. As shown in FIG. 8, the presentation content storage unit 303 stores the user state, sound source type, audio file, designated playback time, etc. The user state indicates, for example, a state that the user should improve. For example, it indicates whether the user should improve the speaking speed (speaking speed improvement) or the sound pressure (sound pressure improvement). The sound source type stores a pseudo-heartbeat sound in the case of speaking speed improvement, and stores a bubble noise in the case of sound pressure improvement. Also, the designated playback time shows different examples for each file in the example of FIG. 8, but they may be the same.
[0059] For example, when receiving a speaking speed improvement signal from the user state determination unit 102, the presentation content control unit 103 may randomly select one from the files in which the user state is speaking speed improvement. For example, when receiving a speaking speed improvement signal, the presentation content storage unit 303 randomly selects either a pseudo-heartbeat sound that is heartbeat 1 or a pseudo-heartbeat sound that is heartbeat 2. Also, for example, only one file corresponding to the speaking speed improvement may be stored in the presentation content storage unit 303. For example, when receiving a speaking speed improvement signal, the presentation content control unit 103 will select the file.
[0060] Similarly, for example, when the user state determination unit 102 receives a sound pressure improvement signal, the presentation content control unit 103 may randomly select one file (pseudo heart sound) from the files where the user state is sound pressure improvement. For example, when receiving a sound pressure improvement signal, the presentation content storage unit 303 randomly selects either bubble noise 1 or bubble noise 2. Also, when receiving an improvement signal indicating that it is necessary to improve both the speech rate and the sound pressure, the presentation content control unit 103 may acquire files of both pseudo heart sounds and bubble noise. Here, for example, the presentation content storage unit 303 may store only one file corresponding to sound pressure improvement. For example, when receiving a sound pressure improvement signal, the presentation content control unit 103 will select the file.
[0061] The presentation content control unit 103 outputs the acquired presentation information through the input / output interface 50 (step ST104).
[0062] FIG. 9 is a flowchart showing an example for explaining the operation of step ST104 in more detail.
[0063] The presentation content control unit 103 acquires the number of audio output devices 54 connected to the information presentation device 1 through the input / output interface 50 (step ST401).
[0064] The presentation content control unit 103 determines whether the number of the acquired audio output devices 54 is plural (step ST402). For example, when the user is wearing a sound conduction type ear device and a bone conduction type ear device, two audio output devices 54 will be connected to the information presentation device 1. In such a case, the presentation content control unit 103 determines that the user is wearing a plurality of audio output devices 54. On the other hand, when the user is wearing either a sound conduction type ear device or a bone conduction type ear device, one audio output device 54 will be connected to the information presentation device 1. In such a case, the presentation content control unit 103 determines that the user is wearing one audio output device 54.
[0065] The presentation content control unit 103 transmits presentation information to a voice output device 54 that is not outputting voice information of other users (step ST403). For example, the presentation content control unit 103 receives voice information of other users through the communication interface 40 and the network 6. Then, the presentation content control unit 103 transmits the received voice information to one of the plurality of voice output devices 54 (the first voice output device), and outputs the presentation information to one of the plurality of voice output devices 54 that is not transmitting voice information (the second voice output device). For example, the presentation content control unit 103 transmits voice information of other users to a bone conduction type ear device, and transmits presentation information to a sound conduction type ear device. Here, of course, the presentation content control unit 103 may transmit voice information of other users to a bone conduction type ear device and transmit presentation information to a sound conduction type ear device. Also, when the presentation information includes both a pseudo heart sound and bubble noise, the presentation content control unit 103 may transmit the pseudo heart sound and the bubble noise to different voice output devices.
[0066] The presentation content control unit 103 synthesizes presentation information with voice information of other users and transmits it to the voice output device 54 (step ST404). The presentation content control unit 103 receives voice information of other users through the communication interface 40 and the network 6. Then, the presentation content control unit 103 transmits information obtained by synthesizing presentation information with voice information to the voice output device 54 through the input / output interface 50.
[0067] The presentation content control unit 103 determines whether the presentation information has been played back by the voice output device 54 for a predetermined time (step ST405). For example, when the playback time of the presentation information is short, the effect of improving the user's speech rate or sound pressure is reduced. Therefore, it is determined whether the voice output device 54 has played back the presentation information for a predetermined time. If it is determined that the voice output device 54 has not played back the presentation information for a predetermined time, the process returns to step ST402. On the other hand, if it is determined that the voice output device 54 has played back the presentation information for a predetermined time, the process ends.
[0068] (Function and Effect) According to the embodiment, in an online meeting or the like, the information presentation device 1 can involuntarily and imperceptibly control the speech rate and sound pressure of the user's speech with respect to the user. As a result, the user can realize speech that is easy to convey to the listener.
[0069] [First Modification Example of the Embodiment] In the first modification example of the embodiment, the information presentation device 1 acquires the user's heart rate and outputs presentation information based on the acquired heart rate.
[0070] (Configuration) The hardware configuration of the information presentation device 1 in the first modification example of the embodiment is the same as that in FIG. 1. Here, it is assumed that the input device 51 includes a wearable terminal or the like and can measure the user's heart rate. Note that the input device 51 is not limited to a wearable terminal, and may include any device as long as it can measure the user's heart rate. Then, the input device 51 outputs the measured heart rate to the information presentation device 1.
[0071] FIG. 10 is a block diagram showing the software configuration of the information presentation device 1 in the first modification example of the embodiment in association with the hardware configuration shown in FIG. 1. The difference from the embodiment is that the control unit 10 includes a heart rate acquisition unit 104.
[0072] The heart rate acquisition unit 104 receives the heart rate from the input device 51 and outputs the received heart rate to the presentation content control unit 103. Further, the heart rate acquisition unit 104 may store the heart rate received from the input device 51 in the data storage unit 30.
[0073] FIG. 11 is a diagram showing an example of the presentation information stored in the presentation content storage unit 303. In the modification example of the embodiment, a plurality of pseudo heart sounds are stored in order to present a pseudo heart sound corresponding to the user's heart rate.
[0074] (Operation) FIG. 12 is a flowchart showing an example of the information presentation operation of the information presentation device 1. The operation of this flowchart is realized by the control unit 10 of the information presentation device 1 reading and executing the program stored in the program storage unit 20.
[0075] Steps ST501 and ST502 in FIG. 12 are the same as steps ST101 and ST102 described with reference to FIG. 3, so the description of these steps is omitted. Here, in step ST502, the user state determination unit 102 outputs a speech speed improvement signal or an improvement signal to the presentation content control unit 103.
[0076] The heart rate acquisition unit 104 acquires the heart rate (step ST503). The heart rate acquisition unit 104 receives the user's heart rate from the input device 51 through the input / output interface 50. The heart rate acquisition unit 104 outputs the acquired heart rate to the presentation content control unit 103.
[0077] The presentation content control unit 103 acquires presentation information (step ST504). For example, when the user's heart rate acquired in step ST503 is 120 beats per minute, a file with a simulated heart sound of 110 beats per minute may be acquired as the presentation information. For example, the presentation content control unit 103 may acquire a simulated heart sound that is about 10% lower than the user's heart rate as the presentation information.
[0078] Step ST505 is the same as step ST104 described with reference to FIG. 3, so the description of this step is omitted.
[0079] The operations described with reference to FIG. 12 can be repeatedly applied. For example, when the heart rate obtained in step ST503 of the first processing is 120 beats per minute, the presentation content control unit 103 acquires an audio file of a heart sound at 110 beats per minute. Then, the presentation content control unit 103 outputs the audio file as presentation information to the audio output device 54. On the other hand, when the heart rate obtained in step ST503 of the second processing has decreased to 110 beats per minute by the first processing, the presentation content control unit 103 acquires an audio file of a heart sound at 100 beats per minute. Then, the presentation content control unit 103 outputs the audio file at 100 beats per minute as presentation information to the audio output device 54. In this way, the presentation content control unit 103 can switch the audio file for auditory presentation according to the current heart rate of the user.
[0080] (Function and Effect) According to the first modification of the embodiment, the control unit 10 acquires the heart rate of the user and outputs presentation information to the audio output device 54 according to the acquired heart rate. Thereby, it is possible to prompt a more effective change in the heart rate as compared with the case of presenting a pseudo heart rate having a heart rate significantly different from the current heart rate to the user as presentation information.
[0081] [Second Modification of the Embodiment] In the second modification of the embodiment, speech rate information or sound pressure information is acquired while the speaker is playing the presentation information, and it is determined whether or not the speech rate information or the sound pressure information has improved within a reference range.
[0082] (Configuration) The configuration of the second modification of the embodiment is the same as the configuration of the embodiment, and thus the description thereof is omitted here.
[0083] (Operation) FIG. 13 is a flowchart showing an example for explaining the operation of step ST104 in more detail. Steps ST601 to ST605 are the same as steps ST401 to ST405 described with reference to FIG. 9, and thus the description of these steps is omitted.
[0084] The presentation content control unit 103 acquires the latest average speech rate or average sound pressure calculated by the user state determination unit 102 (step ST606). For example, when the presentation information is a pseudo-heartbeat sound for improving the speech rate, the presentation content control unit 103 may acquire the average speech rate from the user state determination unit 102. Similarly, when the presentation information is bubble noise for improving the sound pressure, the presentation content control unit 103 may acquire the average sound pressure from the user state determination unit 102.
[0085] The presentation content control unit 103 determines whether at least one of the acquired average speech rate or average sound pressure falls within a reference range (step ST607). Here, the reference range can be arbitrary and may be, for example, the same value as the speech rate threshold or the sound pressure threshold. For example, the reference range may be such that the average speech rate is less than or equal to the speech rate threshold. Alternatively, the reference range may be such that the average sound pressure is greater than or equal to the sound pressure threshold. If it is determined that the average speech rate or average sound pressure is not within the reference range, the process returns to step ST602. On the other hand, if the average speech rate or average sound pressure is within the reference range, the process ends.
[0086] As described above, the presentation content control unit 103 can output the presentation information until at least one of the speech rate or sound pressure falls within the reference range.
[0087] (Function and effect) According to the second modification of the embodiment, when outputting the presentation information to the voice output device 54, it is determined whether the speech rate or sound pressure of the user is improving within the reference range. Thereby, the information presentation device 1 can detect that the speech rate or sound pressure of the user has changed to an appropriate range during the output of the presentation information, and can stop the output of the presentation information.
[0088] [Other embodiments] In the above-described embodiment, an example of outputting presentation information to the voice output device 54 has been described. However, the presentation content control unit 103 may output the presentation information to the display which is the output device 52. Here, when the presentation information is a pseudo-heartbeat sound, the output device 52 changes and displays the change, size, color tone, etc. of an object according to the pseudo-heartbeat sound. Further, when the output device 52 includes a tactile device, it may be presented with the strength and weakness of vibration according to the pseudo-heartbeat sound. Note that when the presentation information is bubble noise, the presentation content control unit 103 may output it to the voice output device 54. That is, the presentation content control unit 103 may transmit the pseudo-heartbeat sound to the output device 52 and transmit the bubble noise to the voice output device. Also, the output device 52 may be a display used in an online meeting or the like, or may be a different display.
[0089] Note that the presentation content control unit 103 may combine these embodiments. That is, the presentation content control unit 103 may of course output the presentation information to the voice output device 54 while also outputting it to the output device 52.
[0090] Also, the method described in the above embodiment can be stored as a program (software means) to be executed by a computer in a storage medium such as a magnetic disk (e.g., a floppy (registered trademark) disk, a hard disk, etc.), an optical disk (CD-ROM, DVD, MO, etc.), a semiconductor memory (ROM, RAM, flash memory, etc.), and can also be transmitted and distributed by a communication medium. Note that the program stored on the medium side includes a setting program for configuring software means (including not only an execution program but also tables and data structures) to be executed by a computer in the computer. The computer that realizes this device reads the program stored in the storage medium, and in some cases, constructs software means by the setting program, and executes the above-described processing by being controlled by this software means. Note that the storage medium referred to in this specification includes not only a storage medium for distribution but also a storage medium such as a magnetic disk or a semiconductor memory provided inside a computer or in a device connected via a network.
[0091] In short, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Further, the embodiments may be combined with each other as appropriate as much as possible, and in that case, the combined effects can be obtained. Furthermore, the above-described embodiments include inventions at various stages, and various inventions can be extracted by appropriate combinations of a plurality of disclosed constituent elements.
Description of Reference Numerals
[0092] 1... Information presentation device 10... Control unit 101... Information acquisition unit 102... User state determination unit 103... Presentation content control unit 104... Heart rate acquisition unit 20... Program storage unit 30... Data storage unit 301... State storage unit 302... Threshold storage unit 303... Presentation content storage unit 40... Communication interface 50... Input / output interface 51... Input device 52... Output device 53... Voice input device 54... Voice output device 6... Network
Claims
1. An acquisition unit that acquires the speech rate and sound pressure of speech at a predetermined interval; A determination unit that determines whether it is necessary to improve at least one of the speech rate or the sound pressure; When it is determined that it is necessary to improve the speech rate, a pseudo-heart sound is acquired, and when it is determined that it is necessary to improve the sound pressure, a bubble noise is acquired, and presentation information including at least one of the pseudo-heart sound or the bubble noise is output. A presentation content control unit; An information presentation device comprising:
2. The information presentation device according to claim 1, wherein the presentation content control unit outputs the pseudo-heart sound to a first device and outputs the bubble noise to a second device.
3. Further comprising a first voice output device and a second voice output device, The information presentation device according to claim 1 or 2, wherein the presentation content control unit outputs voice information from a listener of the speech to the first voice output device and outputs the presentation information to the second voice output device.
4. The information presentation device according to claim 1 or 2, wherein the presentation information is information obtained by synthesizing voice information from a listener of the speech and at least one of the pseudo-heart sound or the bubble noise.
5. Further comprising an acquisition unit that acquires the heart rate of the user who is speaking, The information presentation device according to any one of claims 1 to 4, wherein the presentation content control unit acquires the pseudo-heart sound based on the heart rate.
6. The information presentation device according to any one of claims 1 to 5, wherein the presentation content control unit outputs the presentation information until at least one of the speech rate or the sound pressure is within a reference range.
7. An information presentation method executed by an information presentation device including a processor, The processor obtains the speech rate and sound pressure of speech at a predetermined interval; The processor determines whether it is necessary to improve at least one of the speech rate or the sound pressure; When it is determined that it is necessary to improve the speech rate, the processor obtains a pseudo-heartbeat sound; When it is determined that it is necessary to improve the sound pressure, the processor obtains bubble noise; The processor outputs presentation information including at least one of the pseudo-heartbeat sound or the bubble noise; An information presentation method comprising:
8. An information presentation program for causing a processor to function as each part of the information presentation device according to any one of Claims 1 to 6.
Citation Information
Patent Citations
Uttering condition evaluating device, uttering condition evaluating program, and program storage medium
JP2006267465A
Voice communication apparatus
JP2008294640A
Kalman filtering based speech enhancement using codebook based approach
JP2017194670A
Bathroom mindfulness practice aid system and method, and program
JP2019111307A
Speech translation device, speech translation method, and program therefor
JP2019174784A