Speech speed control method, speech speed control device, and program
The voice speed control method and device dynamically adjust speech speed based on user tension measurements, addressing the inefficiency of manual selection and improving listening comfort by adapting to individual user needs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NT T INC
- Filing Date
- 2024-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing voice speed adjustment technologies require the listener to manually specify or select the speed, which is inefficient and lacks personalization.
A voice speed control method and device that adjusts speech speed based on user tension levels measured by sensors, such as heart rate, electroencephalograph, sweat volume, or autonomic nervous system activity, without requiring listener input.
Automatically adjusts speech speed to accommodate individual user tension levels, enhancing listening comfort by reducing mental or physical strain.
Smart Images

Figure JP2024039348_15052026_PF_FP_ABST
Abstract
Description
Voice Speed Control Method, Voice Speed Control Device, and Program
[0001] The present invention relates to a voice speed control method, a voice speed control device, and a program.
[0002] In communication, the speaking speed varies depending on the person and the situation. The speed at which it is easy to hear, as well as the person and timing, also vary from person to person. Therefore, by adjusting the voice speed according to each individual, it is possible to provide voices that are easy for the listener to hear. Technologies for adjusting the voice speed have been developed (for example, Non-Patent Document 1).
[0003] Toru Tsuki, "Speaking Speed Conversion Technology That Is Convenient and Easy to Hear for Everyone," NHK Broadcasting Technology Research Institute Technical Bulletin, January 2010 issue (https: / / www.nhk.or.jp / strl / publica / giken_dayori / 58 / 5.html) Shoji Makino et al., "Blind Source Separation Based on Independent Component Analysis," 18th Digital Signal Processing Symposium, November 5th, 6th, and 7th, 2003 (https: / / s.makino.w.waseda.jp / reprint / Makino / sm05sice2-9.pdf) Hiroshi Sawada et al., "Blind Separation of Three or More Sound Sources in a Real Environment" (https: / / www.kecl.ntt.co.jp / icl / signal / sawada / mypaper / asj03aki.pdf) Takehito Tōyama et al., "Extraction of Target Sound by Learning Linear Filter Using Similarity between Signals as an Index" (https: / / www.jstage.jst.go.jp / article / ieejeiss / 124 / 5 / 124_5_1134 / _pdf)
[0004] However, in the technology disclosed in Non-Patent Document 1, it is necessary for the listener to specify or select the voice speed.
[0005] An object of the present invention is to provide a technology that can control the voice speed without the listener specifying or selecting it.
[0006] One aspect of the present invention is a voice speed control method that acquires a voice signal, acquires a measurement result of the user's tension level, determines the speed of the voice signal according to the measurement result of the user's tension level, changes the speed of the voice signal based on the determination, and outputs the changed voice signal.
[0007] One aspect of the present invention is a voice speed control device comprising: a voice signal acquisition unit for acquiring a voice signal; a measurement result acquisition unit for acquiring a measurement result of the user's tension level; a voice speed determination unit for determining the speed of the voice signal according to the measurement result of the user's tension level; a voice speed modification unit for changing the speed of the voice signal based on the determination; and a voice signal output unit for outputting the modified voice signal.
[0008] This invention allows the speed of speech to be controlled without the listener having to specify or select a speed.
[0009] This figure shows an example configuration of the voice system 1 according to this embodiment. This figure shows an example configuration of the voice speed control device 11 according to the first embodiment. This is a flowchart showing the operation of the voice speed control device 11 according to the first embodiment. This figure shows an example configuration of the voice speed control device 11 according to the second embodiment. This is a flowchart showing the operation of the voice speed control device 11 according to the second embodiment. This figure shows an example configuration of the voice speed control device 11 according to the third embodiment. This is a flowchart showing the operation of the voice speed control device 11 according to the third embodiment.
[0010] Embodiments of the present invention will be described in detail below with reference to the drawings.
[0011] Figure 1 shows an example of the configuration of the audio system 1 according to this embodiment. The audio system 1 includes an audio signal output device 10, an audio speed control device 11, a sensor 12, and a speaker 13. The audio signal output device 10 outputs an audio signal to the audio speed control device 11. The audio signal output by the audio signal output device 10 is, for example, an audio signal acquired from a medium such as a television or radio. The audio signal output device 10 may convert an audio signal acquired through radio waves or communication lines such as television or radio, or / or an audio acquired from a speaker by a microphone, into an electrical signal, and output the audio signal generated by the conversion to the audio speed control device 11. When the audio signal output device 10 acquires audio, it may perform a process to separate the target audio. Audio separation can be performed, for example, by a blind sound source separation method disclosed in Non-Patent Documents 2-3 or a target sound extraction method disclosed in Non-Patent Document 4.
[0012] The voice speed control device 11 determines the output speed of the voice signal input from the voice signal output device 10 and outputs it to the speaker 13. The output speed of the voice signal determined here is based on the user U's tension level, which is acquired by the sensor 12. The user U's tension level can be rephrased as an indicator of whether the user U can hear the voice without mental or physical strain. Details of the voice speed control device 11 will be described later.
[0013] Sensor 12 measures the user U's level of tension. Sensor 12 is, for example, a heart rate monitor, and the level of tension is indicated by the heart rate. A high heart rate indicates a high level of tension. Sensor 12 is, for example, an electroencephalograph, and the level of tension is indicated by the presence or absence of beta waves. The detection of beta waves indicates a high level of tension. Sensor 12 is, for example, a device that measures the amount of sweat, and the level of tension is indicated by the amount of sweat. A large amount of sweat indicates a high level of tension. Sensor 12 is a device that measures the autonomic nervous system, and the level of tension is indicated by whether or not the sympathetic nervous system is dominant over the parasympathetic nervous system. The dominance of the sympathetic nervous system over the parasympathetic nervous system indicates a high level of tension. Sensor 12 is, for example, a camera, and the level of tension is indicated by eye movements, limb movements, or sweating.
[0014] The speaker 13 receives an audio signal from the audio speed control device 11. The speaker 13 converts the audio signal and outputs it at a determined speed. The speaker 13 can be any device capable of converting and outputting the audio signal input from the audio speed control device 11, such as headphones or earphones. As a result, the audio signal input from the audio signal output device 10 is output from the speaker 13 at the speed determined by the audio speed control device 11, and audio is provided to the user U.
[0015] Furthermore, if the audio signal input from the audio signal output device 10 to the audio speed control device 11 is synchronized with other data such as a video signal, the audio speed control device 11 may also control the output speed of the other data.
[0016] If the audio signal output device 10 is a microphone and acquires audio output from a speaker, the audio speed control device 11 may notify the speaker of the speed of the audio to be output. The speed of the audio output from the audio speed control device 11 is notified to the speaker, for example, by displaying the degree of change in the audio speed as numbers or graphs on a display device. The audio speed control device 11 may also notify the speaker that the output is complete and the audio information that has been output. The audio information that has been output is notified to the speaker, for example, by converting the audio output from the speaker into text and displaying the text on a display device.
[0017] The following describes the speech speed control device 11 in detail. Figure 2 is a diagram showing an example of the configuration of the speech speed control device 11 according to the first embodiment. The speech speed control device 11 includes a speech signal acquisition unit 110, a measurement result acquisition unit 111, a speech speed determination unit 112, a speech speed modification unit 113, and a speech signal output unit 114.
[0018] The audio signal acquisition unit 110 acquires an audio signal from the audio signal output device 10.
[0019] The measurement result acquisition unit 111 acquires the measurement results of the user U's tension level from the sensor 12. The tension level measurement results include data such as heart rate, whether or not beta waves are detected, amount of sweat, and whether or not the sympathetic nervous system is dominant over the parasympathetic nervous system.
[0020] The speech speed determination unit 112 determines the speech speed based on the measurement results of the user U's tension level. The speech speed determination unit 112 determines the speech speed to be slower than normal when the user U is in a state of tension. The speech speed determination unit 112 determines the speech speed to be slower than normal, for example, when the user U's heart rate is above a predetermined threshold. The speech speed determination unit 112 determines the speech speed to be slower than normal, for example, when beta waves are detected from the user U's brain. The speech speed determination unit 112 determines the speech speed to be slower than normal, for example, when the user U's sweating amount is above a predetermined threshold. The speech speed determination unit 112 determines the speech speed to be slower than normal, for example, when the user U's sympathetic nervous system is dominant over the parasympathetic nervous system. Here, "normal speed" refers to the speed of the speech output from the speaker 13 when the speech signal output by the speech signal output device 10 is input to the speaker 13.
[0021] In determining the speech speed, the threshold values used as criteria for heart rate and sweating may be fixed regardless of the user U, or the resting heart rate and sweating amount may be calculated based on previously measured heart rate and sweating amount, and these resting heart rate and sweating amounts may be used as threshold values.
[0022] The speech speed determination unit 112 determines the speech speed by, for example, selecting one speech speed from a plurality of speech speeds. For example, the plurality of speech speeds are two types: "normal speed" and "slower than normal speed". The speech speed determination unit 112 selects "slower than normal speed" if user U is in a state of tension, and selects "normal speed" if user U is not in a state of tension. The speech speed determination unit 112 may also select one speech speed from three or more speech speeds. In this case, it is determined which of the three or more ranges determined by a plurality of thresholds the heart rate or sweat rate falls into. For example, the speech speed determination unit 112 decides to set the speech speed to the first speed if user U's heart rate is above the first threshold, decides to set the speech speed to the second speed if user U's heart rate is below the first threshold but above the second threshold, and decides to set the speech speed to the third speed if user U's heart rate is below the second threshold. The first threshold is greater than the second threshold, the first speed is slower than the second speed, and the second speed is slower than the third speed. The same applies to the sweat rate.
[0023] When the speech speed determination unit 112 selects one speech speed from three or more speech speeds, it determines which frequency band of brainwaves is strongest among several types of frequency bands. Furthermore, for the sympathetic and parasympathetic nervous systems, it determines which of three or more ranges, determined by multiple thresholds, the heart rate variability falls within.
[0024] The audio speed change unit 113 changes the speed at which the audio signal acquired by the audio signal acquisition unit 110 is converted and output, to the speed determined by the audio speed determination unit 112. The audio speed change unit 113 changes the speed at which the audio signal is converted and output by changing the length of time occupied in the time domain of the audio signal.
[0025] The audio signal output unit 114 outputs the audio signal with the modified speed to the speaker 13. The speaker 13 outputs audio at the speed determined by the audio speed determination unit 112.
[0026] Figure 3 is a flowchart showing the operation of the voice speed control device 11 according to the first embodiment. The voice signal acquisition unit 110 acquires a voice signal from the voice signal output device 10 (step S10). The measurement result acquisition unit 111 acquires a measurement result from the sensor 12 (step S11). The voice speed determination unit 112 determines the voice speed based on the measurement result (step S12). The voice speed change unit 113 changes the speed at which the acquired voice signal is converted and output to the speed determined by the voice speed determination unit 112 (step S13). The voice signal output unit 114 outputs the voice signal with the changed output speed to the speaker 13 (step S14). As a result, the speed at which the voice signal output by the voice signal output device 10 is output from the speaker 13 is changed.
[0027] If the speech speed determination unit 112 has determined the speech speed to be slower than normal, and then the measurement result acquisition unit 111 has acquired a measurement result indicating that user U is not in a state of tension, the speech speed determination unit 112 may determine the speech speed to be faster than normal. This allows the system to catch up on the speech that was delayed when the speech speed was slowed down.
[0028] The sensor 12 may include two or more of the following: a heart rate monitor, an electroencephalograph, a device for measuring sweat volume, and a device for measuring the autonomic nervous system. It may measure two or more of the following: heart rate, presence or absence of beta wave detection, sweat volume, and whether the sympathetic nervous system is dominant over the parasympathetic nervous system. In this case, the voice speed determination unit 112 may determine the voice speed based on multiple measurement results of the user U's tension level. If at least one measurement result indicates that the user U is in a state of tension, the voice speed determination unit 112 may determine the voice speed to be slower than normal. If two or more measurement results indicate that the user U is in a state of tension, the voice speed determination unit 112 may determine the voice speed to be even slower.
[0029] The speech speed determination unit 112 may determine the speech speed to be slower than normal if a predetermined number of measurement results indicate that user U is in a state of tension, and to be at the normal speed otherwise. Furthermore, the speech speed determination unit 112 may determine the speech speed to be slower than normal if, among a plurality of measurement results, a specific combination of measurement results all indicate that user U is in a state of tension.
[0030] (Second Embodiment) Figure 4 is a diagram showing an example of the configuration of the voice speed control device 11 according to the second embodiment. The voice speed control device 11 according to the second embodiment includes a user state estimation unit 115 in addition to the voice speed control device 11 according to the first embodiment.
[0031] The user state estimation unit 115 estimates whether user U is in a state of tension based on the measurement results obtained by the measurement result acquisition unit 111. The user state estimation unit 115 estimates whether user U is in a state of tension using an estimation model learned by machine learning. The estimation model is a model that takes the measurement results as input and outputs whether or not the user is in a state of tension, and is a model that has been learned using data that links the subject's measurement results with the presence or absence of tension. The estimation model is stored in the memory unit 119. The user state estimation unit 115 estimates whether or not the user is in a state of tension by inputting the measurement results into the estimation model.
[0032] The machine learning algorithm is not particularly limited. Examples of machine learning algorithms include DNN (Deep Neural Network), SVM (Support Vector Machine), and linear regression. In the second embodiment, the speech speed determination unit 112 determines the speech speed based on the estimation result by the user state estimation unit 115. In the second embodiment, the speech speed determination unit 112 determines the speech speed to a slower speed than normal if, for example, it is estimated that user U is in a state of tension.
[0033] Figure 5 is a flowchart illustrating the operation of the voice speed control device 11 according to the second embodiment. The voice signal acquisition unit 110 acquires a voice signal from the voice signal output device 10 (step S20). The measurement result acquisition unit 111 acquires measurement results from the sensor 12 (step S21). The user state estimation unit 115 estimates whether or not user U is in a state of tension based on the measurement results (step S22). The voice speed determination unit 112 determines the voice speed based on the estimation result by the user state estimation unit 115 (step S23). The voice speed change unit 113 changes the speed at which the acquired voice signal is converted and output to the speed determined by the voice speed determination unit 112 (step S24). The voice signal output unit 114 outputs the voice signal with the changed output speed to the speaker 13 (step S25). As a result, the speed at which the voice signal output by the voice signal output device 10 is output from the speaker 13 is changed.
[0034] Listeners may experience stress due to slow speech speed. In this case, stress can cause an increase in heart rate and other parameters, which in the first embodiment may lead to the listener being judged as being in a state of tension. On the other hand, in the second embodiment, the presence or absence of tension is determined by a model generated by machine learning, allowing the measurement results to be considered comprehensively rather than by a predetermined threshold, thus reducing the possibility of determining a listener as being in a state of tension when the stress is caused by slow speech speed.
[0035] (Third Embodiment) Figure 6 is a diagram showing an example of the configuration of the speech speed control device 11 according to the third embodiment. The speech speed control device 11 according to the third embodiment includes, in addition to the speech speed control device 11 according to the second embodiment, an estimation result presentation unit 116, an answer reception unit 117, and an estimation model retraining unit 118.
[0036] The estimation result presentation unit 116 presents the estimation results from the user state estimation unit 115 to the user U. The estimation result presentation unit 116 presents the estimation results to the user U, for example, by outputting data indicating the estimation results to an external display device or the like.
[0037] User U determines whether the presented estimation result is correct. In other words, User U is presented with an estimation result indicating whether or not they are in a state of tension, and they determine whether or not this is correct for their own state. The response receiving unit 117 receives a response from User U indicating whether or not the estimation result is correct. User U inputs whether or not the estimation result is correct into a predetermined input device, and the response receiving unit 117 obtains the response. This makes it possible to generate new data that links the measurement result with the presence or absence of tension.
[0038] The voice speed control device 11 presents the estimation result and accepts the response at a predetermined timing. The predetermined timing is when there is a larger-than-predetermined fluctuation in the measurement result acquired by the measurement result acquisition unit 111.
[0039] The estimation model relearning unit 118 retrains the estimation model stored in the memory unit 119 using new data that links the measurement results obtained from user U with the presence or absence of tension.
[0040] Figure 7 is a flowchart showing the operation of the voice speed control device 11 according to the third embodiment. The operations in steps S30 to S35 in the third embodiment are the same as the operations in steps S20 to S25 in the second embodiment. In the third embodiment, the estimation result presentation unit 116 presents the estimation result by the user state estimation unit 115 to the user U after the operation in step S32 (step S41). Subsequently, the response reception unit 117 obtains a response from the user U regarding the estimation result (step S42). The estimation model retraining unit 118 retrains the estimation model using new data linked to the measurement results obtained from the user U and the presence or absence of a state of tension (step S43).
[0041] In the flowchart shown in Figure 7, the operations in steps S41 to S43 are performed in parallel with the operations in steps S33 to S35, but the operations in steps S41 to S43 may be performed after the operations in steps S33 to S35.
[0042] In the third embodiment, in addition to the second embodiment, the voice speed control device 11 re-learns the estimation model, so that the estimation accuracy by the estimation model can be further improved.
[0043] As described above, an embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.
[0044] The voice speed control device 11 determines the speed of the voice signal according to the measurement result of the tension of the user U, but is not limited to the tension. The voice speed control device 11 may determine the speed of the voice signal according to the measurement result of the stress / relaxation state of the user U.
[0045] The processing of the voice speed control device 11 in the above-described embodiment may be realized by a computer using software. In that case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the “computer system” shall include hardware such as an OS and peripheral devices. Also, the “computer-readable recording medium” refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built in a computer system. Furthermore, the “computer-readable recording medium” also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, a medium that dynamically holds a program for a short time, and in that case, a volatile memory inside a computer system serving as a server or a client that holds a program for a certain time. Also, the above program may be for realizing a part of the above-described functions, and may further be realized in combination with a program already recorded in the computer system for realizing the above-described functions, and may be realized using a programmable logic device such as an FPGA (Field Programmable Gate Array).
[0046] 1. Voice system, 10. Voice signal output device, 11. Voice speed control device, 12. Sensor, 13. Speaker, 110. Voice signal acquisition unit, 111. Measurement result acquisition unit, 112. Voice speed determination unit, 113. Voice speed change unit, 114. Voice signal output unit
Claims
1. A voice speed control method comprising: acquiring a voice signal; acquiring the result of measuring the user's tension level; determining the speed of the voice signal according to the result of measuring the user's tension level; changing the speed of the voice signal based on the determination; and outputting the changed voice signal.
2. The voice speed control method according to claim 1, wherein the speed of the voice signal is slowed down when the user's level of tension is high.
3. The voice speed control method according to claim 1 or 2, wherein the measurement result of the user's tension level is at least one of the following: the user's heart rate, whether or not beta waves are detected, the amount of sweating, and whether or not the sympathetic nervous system is dominant over the parasympathetic nervous system.
4. The voice speed control method according to claim 1 or 2, wherein, based on the user's measurement results, an estimation model generated by machine learning is used to estimate whether the user is in a state of tension, and the speed of the voice signal is determined based on the estimation result.
5. The voice speed control method according to claim 4, comprising presenting the estimation results to the user, obtaining a response from the user regarding whether or not they are in a state of tension, and retraining the estimation model based on the measurement results and the response.
6. The speech speed control method according to claim 1 or 2, wherein the speech signal is a signal obtained from a speaker's voice, and information indicating a change in the speed of the speech signal is presented to the speaker.
7. A voice speed control device comprising: a voice signal acquisition unit for acquiring a voice signal; a measurement result acquisition unit for acquiring the results of a user's tension level measurement; a voice speed determination unit for determining the speed of the voice signal according to the results of the user's tension level measurement; a voice speed modification unit for changing the speed of the voice signal based on the determination; and a voice signal output unit for outputting the modified voice signal.
8. A program that causes a computer to execute the speech speed control method described in claim 1.