Voice processing system, voice processing method, and voice processing program
The voice processing system addresses the challenge of maintaining smooth online communication by analyzing and adjusting speech in real-time to reduce participant stress, improving interaction quality.
Patent Information
- Application Number
- JP2024127653
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-13
AI Technical Summary
Conventional online communication systems struggle to maintain smooth interaction among multiple participants due to difficulty in understanding and responding to emotional stress levels, especially as participant numbers increase.
A voice processing system that includes an acquisition, determination, identification, and correction processing unit to analyze and adjust speech in real-time, reducing stress levels by modifying factors such as frequency, volume, speaking speed, and intonation based on biometric and emotional data.
Enhances smooth online communication by reducing listener stress through targeted audio adjustments, ensuring clearer and more comfortable speech delivery.
Smart Images

Figure 2026025104000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for controlling speech sounds when multiple users have a conversation. [Background technology]
[0002] In recent years, it has become common for multiple users to communicate online (online conferences, web conferences, etc.) using audio devices (microphone-speaker devices) each equipped with a microphone and a speaker. Conventionally, a technique for online conferences has been known in which the display mode of the facial image of each conference participant is changed in accordance with the emotion of each conference participant estimated based on a facial image extracted from image data of the multiple conference participants (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-6988 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional technology makes it easier to understand participants' emotions in online meetings, where it is more difficult to grasp others' emotions than in face-to-face meetings. However, when the number of participants in a meeting increases, it becomes difficult for the speaker to understand each participant's emotions and change the way they speak, making smooth communication difficult.
[0005] An object of the present disclosure is to provide a voice processing system, a voice processing method, and a voice processing program that enable smooth online communication between multiple users. [Means for solving the problem]
[0006] A voice processing system according to one aspect of the present disclosure includes an acquisition processing unit, a determination processing unit, an identification processing unit, and a correction processing unit. The acquisition processing unit acquires speech speech uttered by a speaker. The determination processing unit determines the stress level of each of multiple listeners listening to the speech speech. The identification processing unit identifies a first listener among the multiple listeners whose stress level is equal to or greater than a threshold, and identifies a first speech speech that causes stress. The correction processing unit performs a correction process to reduce the stress level of the first speech speech for the first listener identified by the identification processing unit.
[0007] Another aspect of the present disclosure is a voice processing method in which one or more processors perform the following steps: acquire speech spoken by a speaker; determine the stress level of each of multiple listeners listening to the speech; identify a first listener among the multiple listeners whose stress level is above a threshold; identify the first speech sound that is a stress factor; and perform a correction process on the first speech sound for the identified first listener to reduce the stress level.
[0008] A voice processing program according to another aspect of the present disclosure is a program for causing one or more processors to execute the following steps: acquire speech spoken by a speaker; determine the stress level of each of multiple listeners listening to the speech; identify a first listener among the multiple listeners whose stress level is above a threshold, and identify the first speech sound that is a stress factor; and perform a correction process on the first speech sound for the identified first listener to reduce the stress level. [Effects of the Invention]
[0009] According to the present disclosure, it is possible to provide a voice processing system, a voice processing method, and a voice processing program that enable smooth online communication between multiple users. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an application example of a conference system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing the configuration of the conference system according to the embodiment of the present disclosure. [Figure 3] FIG. 3 is a diagram illustrating an example of user information used in the conference system according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of audio information used in the conference system according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is an external view showing the configuration of the microphone speaker device according to the embodiment of the present disclosure. [Figure 6] FIG. 6 is a flowchart illustrating an example of a procedure of a conference support process executed in the conference system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples that embody the present disclosure and do not limit the technical scope of the present disclosure.
[0012] The voice processing system according to the present disclosure is applied to a case (e.g., an online conference) in which multiple users in different locations (e.g., a conference room, a home, etc.) hold a conversation (conference) using audio devices (microphone speaker devices) each equipped with a microphone and a speaker. The voice processing system can also be applied to a case in which multiple users hold a conversation in the same location. As an example of the voice processing system according to the present disclosure, a conference system 100 that realizes an online conference will be described.
[0013] Fig. 1 shows an application example of a conference system 100 according to this embodiment. As shown in Fig. 1, users A to D participate in a conference at different locations. Users A to D hold a conversation (online conference) using neckband-type microphone speaker devices 2A to 2D that can be worn around the neck and user terminals 3A to 3D, respectively.
[0014] Each microphone speaker device 2 is wirelessly connected (connected via Bluetooth (registered trademark)) to a user terminal 3, and audio input to the microphone of each microphone speaker device 2 is input to the user terminal 3 and then input from the user terminal 3 to the conference server 1. The audio input to the conference server 1 is transmitted to each microphone speaker device 2 via each user terminal 3 and output (played) from the speaker of each microphone speaker device 2. The microphone speaker device 2 is an example of an audio device of the present disclosure.
[0015] In this way, the conference system 100 has a configuration that enables a plurality of users to hold a conference using the microphone speaker device 2 individually.
[0016] As shown in FIG. 1, the conference system 100 includes a conference server 1, a microphone / speaker device 2, and a user terminal 3. The microphone / speaker device 2 is a wirelessly connected audio device equipped with a microphone 24 and a speaker 25 (see FIG. 5). The conference system 100 includes a plurality of microphone / speaker devices 2, and transmits and receives audio data of users' speech between the plurality of microphone / speaker devices 2. The microphone / speaker devices 2 may be the same type of audio device or different types of audio devices. For example, the plurality of microphone / speaker devices 2 may include a wirelessly connected audio device and a wired connected audio device. The plurality of microphone / speaker devices 2 may also include a neckband-type audio device, a headset-type audio device, and a stationary audio device. The microphone / speaker device 2 may also be built into the user terminal 3.
[0017] The conference server 1 controls the audio (input audio, output audio, etc.) of the microphone / speaker device 2 acquired via the user terminal 3, and executes processing for transmitting and receiving audio between the microphone / speaker device 2 and multiple user terminals 3, for example, when a conference starts. Furthermore, when a specific user (speaker) speaks, if a specific user (listener) listening to the spoken audio feels stressed, the conference server 1 is configured to execute a stress-reducing correction process on the audio output to the user. The conference server 1 alone may constitute the audio processing system of the present disclosure. That is, the audio processing system of the present disclosure may be configured by the conference server 1 alone, or may be configured by the conference server 1, the user terminal 3, and the microphone / speaker device 2.
[0018] The voice processing system of the present disclosure may also have the function of providing various services such as a conference service, a subtitling (transcription) service using voice recognition, a translation service, and a minutes service. In this embodiment, the conference server 1 provides an online conference service. For example, the conference server 1 provides a conference service (online conference service) of a conference application, which is a type of general-purpose software. For example, the conference application is installed in a user terminal 3. A user can start up the user terminal 3 and log in to hold an online conference using the conference application.
[0019] The user terminal 3 is an information processing device (personal computer) owned by a user participating in the conference, and each user can start the conference application on the user terminal 3 and view the conference screen.
[0020] In cases where multiple users hold a conference at the same location, the conference system 100 may include a mixer box (audio transmitting / receiving device) to which the microphone speaker devices 2 are connected. The mixer box has a function of mixing or splitting the audio input from each microphone speaker device 2, for example, and transmits and receives audio (synthesized audio) to and from the conference server 1. In another embodiment, the conference server 1 may have a mixing or splitting function.
[0021] [Conference Server 1] 2, the conference server 1 is a server device including a control unit 11, a storage unit 12, a communication unit 13, etc. For example, the conference server 1 is connected to multiple user terminals 3 via a network, and transmits and receives audio to and from the multiple user terminals 3.
[0022] The communication unit 13 is a communication unit that connects the conference server 1 to the communication network N1 by wire or wirelessly and executes data communication in accordance with a predetermined communication protocol with external devices such as the user terminal 3 via the communication network N1. For example, the communication network N1 may be the Internet, a LAN, a WAN, or a public telephone line.
[0023] The storage unit 12 is a non-volatile storage unit such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory that stores various types of information. The storage unit 12 stores data such as user information D1 related to users participating in the conference and audio information D2 related to speech sounds.
[0024] FIG. 3 shows an example of user information D1. Information such as a user ID, password, and user name is registered in the user information D1. The user ID is user identification information, and the password is a personal identification number set by the user. The user ID and password are used as login information for the conference application. For example, each user registers the user ID, password, and user name when installing the conference application in the user terminal 3. The control unit 11 registers each piece of information in the user information D1 in response to a registration operation by the user. Note that identification information of the microphone speaker device 2 (for example, a Bluetooth (registered trademark) address, a unique number, a name, etc.) may also be registered in the user information D1.
[0025] FIG. 4 shows an example of the voice information D2. Information such as a voice ID, voice data, user ID, and speech time is registered in the voice information D2. The voice ID is voice corresponding to the content of the utterance of the user (speaker) and is identification information of the voice data acquired from the user terminal 3. The voice data is data of the voice (speech) acquired from the microphone speaker device 2. The user ID is identification information of the user corresponding to the utterance. The voice data output from the user terminal 3 is associated with identification information of the user terminal 3 or the user. The speech time is the time when the user speaks, and includes, for example, the start time of the utterance and the end time of the utterance. When the conference starts, the control unit 11 registers each piece of information in the voice information D2 based on the user's utterance acquired from the user terminal 3.
[0026] In the example shown in Figure 4, voice Va with voice ID "101" indicates voice obtained from user A's user terminal 3A, and voice Vb with voice ID "102" indicates voice obtained from user B's user terminal 3B.
[0027] The storage unit 12 also stores control programs such as a conference support program (an example of an audio processing program of the present disclosure) for causing the control unit 11 to execute a conference support process (see FIG. 6 ) described below. For example, the conference support program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, read by a reading device (not shown) such as a CD drive or DVD drive provided in the conference server 1, and stored in the storage unit 12.
[0028] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various types of arithmetic processing. The ROM is a non-volatile storage unit that pre-stores control programs such as a BIOS and an OS that cause the CPU to execute various types of arithmetic processing. The RAM is a volatile or non-volatile storage unit that stores various types of information and is used as a temporary storage memory (work area) for the various types of processing executed by the CPU. The control unit 11 controls the conference server 1 by having the CPU execute various control programs pre-stored in the ROM or the storage unit 12.
[0029] Specifically, as shown in Fig. 2, the control unit 11 includes various processing units such as an acquisition processing unit 111, a determination processing unit 112, a specification processing unit 113, a correction processing unit 114, and an output processing unit 115. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the processing units.
[0030] The acquisition processing unit 111 acquires voices uttered by the user. Specifically, the acquisition processing unit 111 acquires multiple input voices input to the microphones 24 of the multiple microphone speaker devices 2 via the user terminal 3. For example, when a conference starts and a user speaks, the acquisition processing unit 111 acquires the voices (input voices) input to the microphones 24 of the microphone speaker devices 2 of that user by processing the conference application of the user terminal 3. The acquisition processing unit 111 also acquires time information corresponding to the time at which the user's voices were uttered. For example, the acquisition processing unit 111 acquires the time at which the user's voices were input to the microphones 24 of the microphone speaker devices 2, or the time at which the voices were acquired.
[0031] When the acquisition processing unit 111 acquires voice from the user terminal 3, it associates the voice data with the user ID and stores it in the voice information D2 (see FIG. 4). In the example shown in FIG. 1, the acquisition processing unit 111 acquires the speech voices of users A to D who are in different locations from the user terminals 3A to 3D, respectively, and stores each of the input voices Va to Vd in the voice information D2 in association with the user ID.
[0032] The determination processing unit 112 determines the emotions of each of multiple users (listeners) listening to the speech voice. Specifically, the determination processing unit 112 determines the stress level (discomfort level) of each listener with respect to the speech voice of the speaker. For example, the determination processing unit 112 determines the stress level of the listener based on biological information such as the listener's heart rate, pulse rate, and amount of sweat. The determination processing unit 112 may also determine the stress level of the listener based on the listener's facial expression and tone of voice (the quality, tone, speed, etc. of the listener's speech voice). A specific example of a method for determining the stress level of a listener will be described later.
[0033] Figure 5 shows a specific example of the microphone speaker device 2. As shown in Figure 5, the main body 29 of the microphone speaker device 2 has a ring-shaped structure when viewed from above, and has an opening 291 on the front side when viewed from the wearer. In other words, the microphone speaker device 2 has left and right arms when viewed from the user wearing the microphone speaker device 2, and is formed in a U-shape.
[0034] The microphone 24 is disposed on the tip side of the microphone speaker device 2 so as to easily collect the user's voice. The microphone 24 is connected to a microphone board (not shown) built into the microphone speaker device 2. The microphone 24 may be provided on one of the left and right arms, or on both the left and right arms.
[0035] The speakers 25 include a speaker 25L arranged on the left arm and a speaker 25R arranged on the right arm when viewed from the perspective of a user wearing the microphone speaker device 2. The speakers 25L and 25R are arranged near the center of the arms of the microphone speaker device 2 so that the user can easily hear the output sound. The speakers 25L and 25R are connected to a speaker board (not shown) built into the microphone speaker device 2.
[0036] The microphone board is a transmitter board for transmitting audio data to the user terminal 3. The speaker board is a receiver board for receiving audio data from the user terminal 3.
[0037] A power supply 27 and a connection button 28 are arranged on the outside of the main body 29. When the user turns on the power supply 27 and then presses the connection button 28, the microphone speaker device 2 executes pairing processing and connects to the user terminal 3.
[0038] The biosensor 23 is a sensor that reads biometric information such as the heart rate, pulse rate, and amount of sweat of the person wearing the microphone speaker device 2. The biosensor 23 is placed, for example, in a position on the microphone speaker device 2 that comes into contact with the neck of the wearer. The biosensor 23 is an example of a detection unit of the present disclosure. In another embodiment, the biosensor 23 may be configured separately from the microphone speaker device 2.
[0039] The determination processing unit 112 determines the stress level of the listener based on the detection result (biometric information) of the biosensor 23 of the microphone speaker device 2. For example, a threshold value for the increase in heart rate per unit time is set in advance, and when the increase in the listener's heart rate detected by the biosensor 23 is equal to or greater than the threshold value, the determination processing unit 112 determines that the listener is feeling stressed. The determination processing unit 112 also determines the stress level according to the heart rate. For example, when the increase in the heart rate is less than the threshold value, the determination processing unit 112 determines that there is "no stress," and when the increase in the heart rate is equal to or greater than the threshold value, the determination processing unit 112 determines that the stress level increases as the increase in the heart rate increases.
[0040] As another embodiment, for example, the fluctuations in the heart rate of each user may be measured for a certain period of time (1 to 2 minutes), and when the heart rate of the listener detected by the biosensor 23 is equal to or higher than the steady-state heart rate, the determination processing unit 112 may determine that the listener is feeling stressed. Also, as another embodiment, the determination processing unit 112 may determine that the listener is feeling stressed when the listener's heart rate (for example, the listener's average heart rate in normal times) is equal to or higher than a threshold value (for example, the average heart rate when the listener is feeling stressed).
[0041] The same applies to the method of determining the stress level using the pulse rate and the amount of sweating. Known techniques can be applied to the method of determining the stress level based on the heart rate, pulse rate, and amount of sweating.
[0042] In another embodiment, the determination processing unit 112 may determine the stress level based on a facial image of the listener. For example, the determination processing unit 112 may extract a facial image of the listener from an image captured by a camera (an example of an imaging device of the present disclosure) mounted on the microphone / speaker device 2 or the user terminal 3, and analyze the facial expression to determine the stress level. Well-known techniques can be applied as a method for determining the stress level based on the facial expression of the facial image.
[0043] In this way, the determination processing unit 112 determines the stress level of each listener in response to the speaker's speech. For example, when user A, who is the speaker, speaks, the determination processing unit 112 acquires biometric information from the microphone / speaker devices 2C, 2C, and 2D of users B, C, and D, who are listeners, respectively, and determines the stress levels of users B, C, and D.
[0044] The identification processing unit 113 identifies a listener (hereinafter referred to as a "target listener") (first listener in the present disclosure) whose stress level is equal to or greater than a threshold among the multiple listeners. For example, the identification processing unit 113 identifies, among users B, C, and D who are listeners, a user whose heart rate increase is equal to or greater than a threshold as the target listener.
[0045] The correction processing unit 114 performs a correction process to reduce the stress level of the speaker voice output to the target listener identified by the identification processing unit 113. Specifically, the correction processing unit 114 identifies the speech voice (target voice) of the speaker who spoke at the time when the listener felt stressed. For example, the correction processing unit 114 identifies the speech voice of the speaker who spoke when the increase in the listener's heart rate exceeded a threshold. For example, when user A speaks and the increase in user B's heart rate exceeded a threshold, the identification processing unit 113 refers to the voice information D2 (see FIG. 4 ) and identifies the speech voice of user A at the time when the increase in user B's heart rate exceeded the threshold as the target voice. The correction processing unit 114 analyzes the identified target voice and estimates the cause of stress based on the analysis result. Then, the correction processing unit 114 performs the correction process on user A's speech voice according to the estimated cause.
[0046] The output processing unit 115 outputs the input voice acquired by the acquisition processing unit 111 to each user terminal 3. For example, the output processing unit 115 outputs the speech voice of user A to the user terminals 3B to 3D of users B to D, and causes the speech voice to be output (played) from the speakers 25 of the microphone speaker devices 2B to 2D via the user terminals 3B to 3D.
[0047] Furthermore, the output processing unit 115 outputs the speaker voice corrected by the correction processing unit 114 to the user terminal 3 of the target listener (user B in the above example) identified by the identification processing unit 113, and causes it to be output (played) from the speaker 25 of the microphone speaker device 2 via the user terminal 3. That is, the output processing unit 115 outputs the speaker voice on which the correction processing has been performed to the target listener (user B in the above example), and outputs the speaker voice (input voice) on which the correction processing has not been performed to other listeners (users C and D in the above example, if the identification processing unit 113 determines that users C and D are not feeling stressed).
[0048] A specific example of the correction process for the speaker voice will be described below. When there is a target listener whose stress level is equal to or greater than a threshold, the correction processing unit 114 estimates the cause of stress for the target listener from the analysis results of the target voice, and performs correction processing corresponding to the estimated cause only on the speaker voice for the target listener.
[0049] [Example 1] If the correction processing unit 114 estimates from the analysis result of the target voice that the high frequency of the target voice is a cause of stress, it converts the frequency of the speaker's voice (input voice).
[0050] For example, the correction processing unit 114 measures the frequency of the voice and performs a Fourier transform to measure the frequency domain of the voice (pitch of the voice). The correction processing unit 114 determines whether the voice is causing discomfort to the listener by setting a certain threshold in the frequency domain of the voice. The correction processing unit 114 determines whether the target voice is too high or too low.
[0051] Generally, in the frequency range of speech, the frequency that is easiest to hear is said to be approximately 440 Hz. Therefore, for example, discomfort level 1 (2500-3000 Hz), discomfort level 2 (3000-3500 Hz), discomfort level 3 (3500-4000 Hz), discomfort level 4 (4000-4500 Hz), and discomfort level 5 (4500 Hz or higher) are set, and the correction processor 114 determines the discomfort level corresponding to the frequency of the target speech. The correction processor 114 then performs correction processing according to the discomfort level. For example, if the target speech of user A is at discomfort level 1, the correction processor 114 converts the frequency of user A's speaker's voice to approximately 440 Hz. Furthermore, if the target speech of user A is at discomfort level 2, the correction processor 114 may convert the frequency of user A's speaker's voice to a frequency for discomfort level 1 or to approximately 440 Hz. The output processing unit 115 outputs the frequency-converted speaker voice from the microphone speaker device 2 via the user terminal 3 of the target listener.
[0052] The correction processing unit 114 may repeat the conversion process of the frequency of the speaker's voice until the stress level of the listener becomes less than the threshold. For example, the correction processing unit 114 gradually lowers the frequency of the speaker's voice until the stress level of the listener becomes less than the threshold.
[0053] [Example 2] If the correction processing unit 114 estimates from the analysis result of the target voice that the sound pressure (volume) of the target voice is a cause of stress, it converts the sound pressure of the speaker's voice (input voice).
[0054] For example, the correction processing unit 114 determines whether there is a problem with the sound pressure (volume) by setting a certain threshold and determining whether the target sound is too loud or too quiet.
[0055] Generally, the sound pressure (volume) of a normal speaking voice is about 60 dB. Therefore, for example, discomfort level 1 (about 50 dB, about 70 dB), discomfort level 2 (about 40 dB, about 80 dB), discomfort level 3 (about 30 dB, about 90 dB), discomfort level 4 (about 20 dB, about 100 dB), and discomfort level 5 (50 dB or less, 100 dB or more) are set for each sound pressure level. The correction processing unit 114 then determines the discomfort level corresponding to the sound pressure of the target voice. The correction processing unit 114 then performs correction processing according to the discomfort level. For example, if the target voice of user A is at discomfort level 1, the correction processing unit 114 converts the sound pressure of user A's speaker voice to approximately 60 dB. Furthermore, if the target voice of user A is at discomfort level 2, the correction processing unit 114 may convert the sound pressure of user A's speaker voice to a frequency corresponding to discomfort level 1 or to approximately 60 dB. The output processing unit 115 outputs the speaker's voice, whose sound pressure has been converted, from the microphone speaker device 2 via the user terminal 3 of the target listener.
[0056] The correction processing unit 114 may repeat the sound pressure conversion process until the listener's stress level becomes less than the threshold. For example, the correction processing unit 114 gradually brings the sound pressure closer to approximately 60 dB until the listener's stress level becomes less than the threshold.
[0057] [Example 3] If the correction processing unit 114 estimates from the analysis result of the target voice that the fast speaking speed of the target voice is a cause of stress, it converts the speed of the speaker's voice (input voice).
[0058] For example, the correction processing unit 114 uses a voice recognition AI to count the number of characters spoken in a certain period of time and determine whether the speaking speed is appropriate.
[0059] Generally, a speech speed (talk speed) of approximately 300 characters per minute (approximately 50 characters per 10 seconds) is considered ideal. For example, discomfort level 1 (approximately 275 characters, approximately 325 characters), discomfort level 2 (approximately 250 characters, approximately 350 characters), discomfort level 3 (approximately 225 characters, approximately 375 characters), discomfort level 4 (approximately 200 characters, approximately 400 characters), and discomfort level 5 (175 characters or less, 400 characters or more) are set for each level of speech speed per minute. The correction processor 114 then determines the discomfort level corresponding to the speech speed of the target speech. The correction processor 114 then performs correction processing according to the discomfort level. For example, if the target speech of user A is discomfort level 1, the correction processor 114 converts the speech speed of user A's speaker voice to a speed (playback speed) of approximately 300 characters per minute. Furthermore, when the target voice of user A is at discomfort level 2, correction processing unit 114 may convert the speaking speed of user A's speaker voice to a speaking speed for discomfort level 1, or to a speed of about 300 characters. Output processing unit 115 outputs the speaker voice, whose speaking speed has been converted, from microphone speaker device 2 via user terminal 3 of the target listener. Note that correction processing unit 114 changes the speaking speed by adjusting silent sections.
[0060] [Example 4] If the correction processing unit 114 estimates from the analysis result of the target voice that the intonation (way of speaking) of the target voice is a cause of stress, the correction processing unit 114 suppresses the intonation of the speaker's voice (input voice).
[0061] Generally, when a voice rises (high-pitched voice) at the end of a sentence, it sounds like the speaker is just letting the sound go. It is also said that speaking in a high-pitched voice and finishing in a low-pitched voice (reading down) is easier to understand. Speaking with a pronounced intonation is also said to make the words sound more emotional and leave a better impression. For example, the number of times the voice pitch rises at the end of a sentence per minute can be set as discomfort level 1 (less than once), discomfort level 2 (twice), or discomfort level 3 (three or more times). The correction processor 114 then determines the discomfort level corresponding to the intonation of the target voice. The correction processor 114 then performs correction processing according to the discomfort level. For example, if the target voice of user A is at discomfort level 1, the correction processor 114 adjusts the voice pitch to suppress the intonation of user A's speaker voice. The output processor 115 outputs the speaker voice with the adjusted intonation from the microphone speaker device 2 via the user terminal 3 of the target listener.
[0062] The correction processing unit 114 may suppress intonation by adjusting each piece of audio information, such as the pitch (frequency), sound pressure (volume), and speaking speed.
[0063] [Example 5] The correction processor 114 may be configured to customize the adjustment level of the audio correction process. Specifically, the correction processor 114 may perform correction processing suited to the conference and the conference participants. For example, the correction processor 114 may be set to perform weak correction processing in the first conference, and the correction level may be increased depending on the behavior of the conference participants (listeners).
[0064] For example, the correction processing unit 114 sets the adjustment level of the voice correction before the conference. For example, the correction processing unit 114 sets the strength of the correction processing in advance according to the preferences of the participants. Specifically, correction level 1 "weak correction processing", correction level 2 "normal correction processing", and correction level 3 "strong correction processing" are set in advance, and the correction processing unit 114 performs correction processing on the speaker's voice according to the set correction levels.
[0065] [Example 6] When the same participants hold multiple conferences, the correction processing unit 114 may perform the correction processing based on the past (previous) voice analysis results. In the present disclosure, since each conference participant (listener) individually uses the microphone speaker device 2 during the conference, the analysis results of the speech voice and the listener's emotions (stress level, discomfort level) can be easily associated with the listener. Therefore, the correction processing unit 114 can perform the correction processing for each listener using the past analysis results.
[0066] Furthermore, by accumulating data linked to participants, the correction processing unit 114 can grasp the trends in the analysis results of speech sounds and the analysis results of listener emotions when a conference is held with the same participants.
[0067] Furthermore, by utilizing trends known in advance, it is possible to perform voice correction processing immediately after the start of the conference (before obtaining voice analysis and emotion analysis results). Furthermore, since correction processing can be performed immediately after the speaker starts speaking, which is considered to be a cause of stress for listeners, it is possible to respond before the listener feels stressed. Furthermore, by utilizing trends known in advance, it is possible to perform appropriate correction processing without providing a biosensor 23 in the microphone / speaker device 2. Furthermore, by performing correction processing based on trends rather than performing emotion analysis during the conference, it is possible to reduce the processing load on the conference server 1.
[0068] [Example 7] If the correction processing unit 114 estimates from the analysis results of the target voice that the voice quality of the target voice is a cause of stress, the correction processing unit 114 may convert the speaker's voice into a voice with a different sound quality using a voice changer or the like. For example, if the listener is in an extremely stressful state, or if the speaker's voice is extremely unpleasant, and stress cannot be alleviated by partial voice correction processing, the correction processing unit 114 converts the voice into a voice with a different voice quality. For example, the correction processing unit 114 may set the characteristics (unique information) of each listener in advance, such as when the target listener dislikes the voice of a particular speaker, male voices, or female voices, and convert the voice into a voice quality that matches the listener's characteristics.
[0069] As described above, when there is a target listener whose stress level is equal to or higher than a threshold, the correction processor 114 estimates the stress factor of the target listener and performs a correction process (Examples 1 to 7) corresponding to the estimated factor on the speaker's voice. For example, the correction processor 114 estimates the stress factor of the listener based on the speaker's voice and corrects at least one of the frequency, volume, and playback speed of the speaker's voice based on the estimation result. The correction processor 114 may perform any of the correction processes of Examples 1 to 7, or may perform a combination of multiple correction processes. For example, when two listeners whose stress levels are equal to or higher than a threshold have different stress factors, the correction processor 114 performs a correction process according to the stress factor for each listener. The correction processor 114 may also repeatedly perform the correction process by gradually increasing the correction degree until the stress level of the target listener becomes less than the threshold. When repeating the correction process, the correction processor 114 may perform a correction process of a different example.
[0070] In addition, the correction processing unit 114 may perform correction processing on the speaker voices for all listeners, for example, when the proportion of listeners whose stress level is above a threshold reaches a predetermined proportion or more, or when the total number of listeners whose stress level is above a threshold reaches a predetermined number or more.
[0071] The conference server 1 may be a cloud server or a computer (personal computer) installed in the conference room R1.
[0072] [Meeting support processing] FIG. 6 shows an example of the procedure of the conference support process executed by the control unit 11 of the conference server 1.
[0073] The present disclosure can be understood as a conference support method (audio processing method of the present disclosure) that executes one or more steps included in the conference support process. Furthermore, one or more steps included in the conference support process described herein may be omitted as appropriate. Furthermore, the steps in the conference support process may be executed in a different order as long as the same operational effect is achieved. Furthermore, while the present disclosure uses an example in which the control unit 11 executes each step in the conference support process, in other embodiments, one or more processors may execute each step in the conference support process in a distributed manner.
[0074] <Step S1> In step S1, the control unit 11 determines whether an operation to start a conference has been accepted. For example, each user participating in the conference starts a conference application on their own user terminal 3, performs a login operation, and also performs a conference start operation on a setting screen. When the control unit 11 accepts the conference start operation (S1: Yes), it shifts the processing to step S2. The control unit 11 waits until the conference start operation is accepted (S1: No).
[0075] <Step S2> In step S2, the control unit 11 starts a process of acquiring from the user terminal 3 the voice (speaker voice) spoken by the user and input to the microphone speaker device 2. For example, when a conference starts and the voice spoken by user A is input to the microphone 24 of the microphone speaker device 2A, the control unit 11 acquires the speaker voice (input voice) from the user terminal 3A to which the microphone speaker device 2A is connected. Also, when the voice spoken by user B is input to the microphone 24 of the microphone speaker device 2B, the control unit 11 acquires the speaker voice (input voice) from the user terminal 3B to which the microphone speaker device 2B is connected. The control unit 11 registers information related to the acquired input voice in voice information D2 (see FIG. 4).
[0076] <Step S3> In step S3, the control unit 11 starts a process of acquiring biometric information of each user from the user terminal 3. Specifically, when the conference starts, the control unit 11 acquires biometric information (heart rate, pulse rate, amount of sweat, etc.) detected by the biometric sensor 23 provided in the microphone speaker device 2 via the user terminal 3. For example, when user A starts speaking, the control unit 11 starts a process of acquiring the biometric information of users A to D from the user terminals 3A to 3D. Note that the control unit 11 may omit the process of acquiring the biometric information of the speaker (user A).
[0077] <Step S4> In step S4, the control unit 11 determines the stress level of each listener based on the biometric information of each listener, and determines whether or not there is a listener (target listener) whose stress level is above the threshold. The method for determining the stress level based on the biometric information can be the above-mentioned method (well-known method). If the control unit 11 determines that there is a listener (target listener) whose stress level is above the threshold (S4: Yes), the control unit 11 proceeds to step S5. On the other hand, if the control unit 11 determines that there is no listener (target listener) whose stress level is above the threshold (S4: No), the control unit 11 proceeds to step S8.
[0078] <Step S5> In step S5, the control unit 11 identifies the speech (target speech) of a speaker who is speaking at a timing when a listener (target listener) whose stress level is equal to or higher than a threshold feels stress. For example, the control unit 11 identifies the speech (target speech) of a speaker who is speaking when the listener's heart rate, pulse rate, or sweat rate reaches or exceeds a threshold.
[0079] <Step S6> In step S6, the control unit 11 estimates the cause of stress for the target listener. Specifically, the control unit 11 estimates the cause of stress from the analysis results of the target voice. For example, the control unit 11 compares voice information such as the frequency, sound pressure (volume), speaking speed, and intonation (way of speaking) of the target voice with indices to estimate which voice information is the cause of stress. The control unit 11 estimates voice information that is beyond the range that is generally easy for people to hear as the cause of stress.
[0080] <Step S7> In step S7, the control unit 11 performs a correction process to reduce stress on the speaker's voice (input voice) based on the estimated stress factor. For example, if the control unit 11 estimates that the high frequency of the target voice is the stress factor, it converts the frequency of the speaker's voice. If the control unit 11 estimates that the sound pressure (volume) of the target voice is the stress factor, it converts the sound pressure of the speaker's voice. If the control unit 11 estimates that the speaking speed of the target voice is the stress factor, it converts the speed of the speaker's voice. If the control unit 11 estimates that the intonation (way of speaking) of the target voice is the stress factor, it suppresses the intonation of the speaker's voice.
[0081] <Step S8> In step S8, the control unit 11 outputs the input voice (speaker voice) to the user terminal 3 of each listener. For example, if there is no listener (target listener) whose stress level is equal to or greater than the threshold (S4: No), the control unit 11 outputs the input speaker voice as is to the user terminal 3 of each listener. As a result, the voice is reproduced from the speaker 25 of each microphone-speaker device 2.
[0082] On the other hand, if there is a listener (target listener) whose stress level is equal to or higher than the threshold (S4: Yes), the control unit 11 outputs the corrected speaker voice to the user terminal 3 of the target listener, and outputs the uncorrected speaker voice (input voice) to the user terminal 3 of the listener whose stress level is below the threshold. As a result, the corrected speaker voice is reproduced from the microphone / speaker device 2 of the listener whose stress level is equal to or higher than the threshold, and the uncorrected speaker voice is reproduced from the microphone / speaker device 2 of the listener whose stress level is below the threshold. Note that if there is a listener (target listener) whose stress level is equal to or higher than the threshold (S4: Yes) and the cause of stress for the target listener cannot be identified, the control unit 11 outputs the uncorrected speaker voice (input voice) to the target listener as well.
[0083] <Step S9> In step S9, the control unit 11 determines whether or not an operation to end the conference has been accepted. For example, each user participating in the conference performs an operation to end the conference on their own user terminal 3. If the control unit 11 accepts the operation to end the conference (S9: Yes), it ends the conference support process. If the control unit 11 does not accept the operation to end the conference (S9: No), it shifts the process to step S4. The control unit 11 repeatedly executes the processes of steps S4 to S8 until the conference ends. As a result, for example, for a listener whose stress level is equal to or higher than a threshold, the correction process is repeatedly executed on the speaker's voice until the stress level becomes less than the threshold.
[0084] As described above, the conference system 100 according to the present disclosure acquires speech uttered by a speaker, determines the stress level of each of multiple listeners who listen to the speech, identifies a target listener (first listener) among the multiple listeners whose stress level is equal to or greater than a threshold, and performs a correction process to reduce the stress level of the speech output to the identified target listener.
[0085] According to the above configuration, the speech can be individually adjusted according to the emotion (stress) of each listener, thereby enabling smooth online communication among a plurality of users.
[0086] [Other embodiments] Other embodiments of the conference system 100 according to the present disclosure will be described below.
[0087] For example, when there is a listener (target listener) whose stress level is above a threshold, the control unit 11 of the conference server 1 may notify the speaker of information indicating that the listener is feeling stressed. For example, the control unit 11 may cause the speaker's user terminal 3 or microphone / speaker device 2 to display or output instructions such as "You're too excited," "Speak a little quieter," or "Speak more slowly." The control unit 11 may correct the target voice and output information to the speaker encouraging them to improve their speaking style.
[0088] In another embodiment, when there is a listener (target listener) whose stress level is equal to or higher than a threshold, the control unit 11 may stop audio output of the speaker's voice and output text information obtained by converting (transcribes) the speaker's voice to the user terminal 3 of the target listener. This prevents the target listener from listening to stressful voice and allows the target listener to obtain the text information displayed on the user terminal 3. In another embodiment, when the stress level of the target listener does not become less than the threshold as a result of the correction process, the control unit 11 may output text information obtained by converting the target voice to the target listener.
[0089] In another embodiment, the control unit 11 may execute a process to remove an external factor if the listener feels stressed due to a factor other than the speaker's voice (an external factor). For example, if an external factor such as a noisy room, a smelly room, a dark room, or too much work is causing stress to the target listener, the control unit 11 may notify the user terminal 3 of the target listener of a message urging the listener to improve the surrounding environment before holding a conference. Note that if the speaker's voice correction process (Examples 1 to 7) does not reduce the target listener's stress level below a threshold, the control unit 11 may determine that the external factor is the cause of stress and notify the message.
[0090] In another embodiment, the control unit 11 may determine the physical condition of each user participating in the conference based on biological information. For example, the control unit 11 may detect an abnormal heartbeat such as arrhythmia, rapid breathing, or excessive fatigue based on the biological information. The control unit 11 may then notify the user determined to be in poor physical condition and other conference participants to encourage them to see a doctor or take a break.
[0091] In another embodiment, the control unit 11 may perform processing to prevent voice overlap when multiple conference participants are present in the same conference room. For example, when one of multiple conference participants speaks in the same conference room, the spoken voice may be input to the microphone / speaker devices 2 of each conference participant. In this case, the same voice may be input to the conference server 1 in duplicate. Therefore, the control unit 11 may not perform correction processing on the voice of the speaker in the same conference room, but may perform correction processing only on the voice from another conference room and play it back from each user's microphone / speaker device 2. Furthermore, when performing the correction processing on all voices, the control unit 11 may perform noise cancellation on ambient sounds (voices in the same conference room) and play back only the corrected voice from the listener's microphone / speaker device 2.
[0092] In another embodiment, the voice processing system of the present disclosure may be applied to applications other than meetings. For example, when a user is depressed or sad, they may be unable to concentrate on the sound, making it difficult to hear what is being said or understand the content. The microphone speaker device 2 according to this embodiment can be used to perform emotion analysis and correction processing to play voices that match the user's mood, such as when watching television or listening to music. Furthermore, automatically selecting the optimal sound for watching videos or listening to music according to the user's mood improves the user's quality of life. For example, the control unit 11 analyzes the user's emotions and the voices being played. The control unit 11 then adjusts the volume and the audibility of the voices according to the results of the emotion analysis. For example, if the user is sad, the control unit 11 may be unable to concentrate, so the control unit 11 plays voices at a louder, clearer volume. For example, if the user is angry, the control unit 11 plays voices that are calming.
[0093] In another embodiment, each processing unit included in the conference server 1 (acquisition processing unit 111, determination processing unit 112, specification processing unit 113, correction processing unit 114, and output processing unit 115) may be included in each user terminal 3.
[0094] The control unit 11 of the conference server 1 controls the entire conference server 1. The control unit 11 realizes various functions by reading and executing various programs stored in the memory unit 12 (for example, storage or ROM). The control unit 11 may be realized by one or more control devices / arithmetic units (CPUs (Central Processing Units), SoCs (System on a Chip)). The control unit 11 may also be configured by one or more control circuits (electronic circuits).
[0095] As described above, the conference system 100 according to this embodiment may have the following features.
[0096] (Feature 1) In a conference system 100 that transmits and receives voices of multiple participants via a network, a control unit 11 analyzes the stress level of a listener, analyzes the speaker's voice under predetermined conditions, and analyzes the discomfort level of the listener. The control unit 11 determines whether or not the speaker's voice to be transmitted to the listener needs to be converted, and converts the speaker's voice. If the analyzed stress level of the listener is equal to or greater than a predetermined threshold, the control unit 11 converts the speaker's voice to be transmitted to the listener in a direction that reduces the discomfort level of the listener, based on the analysis result of the speaker's voice under the predetermined conditions.
[0097] (Feature 2) The predetermined condition is that the volume of the speaker's voice is outside a predetermined range.
[0098] (Feature 3) The predetermined condition is that the frequency of the speaker's voice is outside a predetermined range.
[0099] (Feature 4) The predetermined condition is that the speaking speed of the speaker's voice is equal to or greater than a predetermined speed.
[0100] (Feature 5) The control unit 11 determines the stress level based on the fluctuation of the heart rate.
[0101] (Feature 6) The system further includes an imaging device for capturing images of the participants, and the heart rates of the listeners are obtained by analyzing the images of the listeners captured by the imaging device.
[0102] (Feature 7) The system further includes a sensor for acquiring the heart rate of the participant, and the heart rate of the listener is acquired by the sensor.
[0103] (Feature 8) The sensor is provided in the headset (microphone speaker device 2).
[0104] (Feature 9) If the discomfort level of the listener does not improve even after converting the voice to be output to the listener, the control unit 11 stops outputting the voice to the listener and displays the result of converting the speaker's voice into text.
[0105] (Feature 10) The control unit 11 gradually increases the voice conversion that suppresses the discomfort level while checking the discomfort level of the listener, and stops increasing the voice conversion when the discomfort level falls within a threshold value.
[0106] (Feature 11) When the listener's discomfort level as a result of the stress analysis is equal to or greater than a predetermined threshold and voice conversion is performed based on the analysis results of predetermined conditions, the control unit 11 notifies the speaker to request improvement.
[0107] (Feature 12) The control unit 11 uniformly converts the voice of a specific speaker based on the past analysis history.
[0108] (Feature 13) The control unit 11 can set the degree of discomfort reduction for each listener.
[0109] (Feature 14) The control unit 11 converts the voices of all speakers whose discomfort levels are outside a predetermined range.
[0110] (Feature 15) The control unit 11 converts the voice of the speaker whose discomfort level is the furthest from the predetermined range.
[0111] [Disclosure Note] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.
[0112] <Appendix 1> an acquisition processing unit that acquires speech uttered by a speaker; a determination processing unit that determines a stress level of each of a plurality of listeners who listen to the speech sound; an identification processing unit that identifies a first listener whose stress level is equal to or greater than a threshold value from among the plurality of listeners; a correction processing unit that executes a correction process to reduce the stress level on the speech voice output to the first listener identified by the identification processing unit; A voice processing system comprising:
[0113] <Appendix 2> the correction processing unit estimates a cause of stress of the listener based on the speech sound, and corrects at least one of a frequency, a volume, and a playback speed of the speech sound based on the estimation result. 10. The speech processing system of claim 1.
[0114] <Appendix 3> the determination processing unit determines the stress level of the listener based on biological information of the listener. 3. The speech processing system according to claim 1 or 2.
[0115] <Appendix 4> the determination processing unit determines the stress level of the listener based on the biological information of at least one of a heart rate, a pulse rate, and an amount of sweat of the listener. 4. The speech processing system of claim 3.
[0116] <Appendix 5> the determination processing unit acquires the biological information of the listener from the audio device, the biological information being detected by a detection unit mounted on the audio device that is worn by the listener and outputs the spoken voice; 5. The speech processing system of claim 4.
[0117] <Appendix 6> the determination processing unit determines the stress level of the listener based on an expression of a facial image in a captured image acquired from an imaging device that captures an image of the listener. 6. A speech processing system according to any one of Supplementary notes 3 to 5.
[0118] <Appendix 7> the correction processing unit repeatedly performs the correction process by successively increasing a correction degree until the stress level of the first listener becomes less than the threshold value. 7. A speech processing system according to any one of Supplementary notes 1 to 6.
[0119] <Appendix 8> an output processing unit that outputs the speech sound to each of the plurality of listeners; the output processing unit outputs the speech sound on which the correction process has been performed to the first listener, and outputs the speech sound on which the correction process has not been performed to the other listeners. 8. A speech processing system according to any one of Supplementary Notes 1 to 7.
[0120] <Appendix 9> an output processing unit that outputs the speech sound to each of the plurality of listeners; the output processing unit outputs, to the first listener, text information obtained by converting the uttered voice into text when the stress level of the first listener does not become less than the threshold value as a result of the correction process. 9. A speech processing system according to any one of Supplementary notes 1 to 8.
[0121] <Appendix 10> The output processing unit outputs information to the speaker to encourage the speaker to improve his or her speaking style. 9. The speech processing system of claim 8.
[0122] <Appendix 11> Acquiring a speech sound uttered by a speaker; determining a stress level of each of a plurality of listeners who listen to the speech sound; Identifying a first listener among the plurality of listeners whose stress level is equal to or greater than a threshold, and identifying a stress factor; performing a correction process for reducing the stress level on the speech voice output to the first listener; An audio processing method executed by one or more processors.
[0123] <Appendix 12> Acquiring a speech sound uttered by a speaker; determining a stress level of each of a plurality of listeners who listen to the speech sound; Identifying a first listener among the plurality of listeners whose stress level is equal to or greater than a threshold, and identifying a stress factor; performing a correction process for reducing the stress level on the speech voice output to the first listener; An audio processing program for causing one or more processors to execute the above. [Explanation of symbols]
[0124] 100: Conference system 1: Conference server 2: Microphone speaker device 3: User device 11: Control section 12: Storage section 13: Communications Department 23: Biometric sensor 24:Mike 25: Speaker 27: Power supply 28: Connect button 111: Acquisition processing unit 112: Judgment processing unit 113: Specific processing unit 114: Correction processing unit 115: Output processing section D1: User information D2: Audio information
Claims
1. an acquisition processing unit that acquires speech uttered by a speaker; a determination processing unit that determines a stress level of each of a plurality of listeners who listen to the speech sound; an identification processing unit that identifies a first listener whose stress level is equal to or greater than a threshold value from among the plurality of listeners; a correction processing unit that executes a correction process to reduce the stress level of the speech voice output to the first listener identified by the identification processing unit; A voice processing system comprising:
2. the correction processing unit estimates a cause of stress of the listener based on the speech sound, and corrects at least one of a frequency, a volume, and a playback speed of the speech sound based on the estimation result. The audio processing system of claim 1 .
3. the determination processing unit determines the stress level of the listener based on biological information of the listener. The audio processing system of claim 1 .
4. the determination processing unit determines the stress level of the listener based on the biological information of at least one of a heart rate, a pulse rate, and an amount of sweat of the listener. The audio processing system of claim 3 .
5. the determination processing unit acquires the biological information of the listener from the audio device, the biological information being detected by a detection unit mounted on the audio device that is worn by the listener and outputs the spoken voice; 5. The audio processing system of claim 4.
6. the correction processing unit repeatedly performs the correction process by gradually increasing a correction degree until the stress level of the first listener becomes less than the threshold value. The audio processing system of claim 1 .
7. an output processing unit that outputs the speech sound to each of the plurality of listeners; the output processing unit outputs the speech sound on which the correction process has been performed to the first listener, and outputs the speech sound on which the correction process has not been performed to the other listeners. The voice processing system according to any one of claims 1 to 6.
8. an output processing unit that outputs the speech sound to each of the plurality of listeners; the output processing unit outputs, to the first listener, text information obtained by converting the uttered voice into text when the stress level of the first listener does not become less than the threshold value as a result of the correction process. The voice processing system according to any one of claims 1 to 6.
9. Acquiring a speech sound uttered by a speaker; determining a stress level of each of a plurality of listeners who listen to the speech sound; Identifying a first listener among the plurality of listeners whose stress level is equal to or greater than a threshold, and identifying a stress factor; performing a correction process for reducing the stress level on the speech voice output to the first listener; An audio processing method executed by one or more processors.
10. Acquiring a speech sound uttered by a speaker; determining a stress level of each of a plurality of listeners who listen to the speech sound; Identifying a first listener among the plurality of listeners whose stress level is equal to or greater than a threshold, and identifying a stress factor; performing a correction process for reducing the stress level on the speech voice output to the first listener; An audio processing program for causing one or more processors to execute the above.
Citation Information
Patent Citations
Information processing device, program, and attendant information notification method
JP2023006988A