Conversation support device, conversation support system, conversation support method, and program
The conversation support device and system enhance emotional expression in voice communication by converting selected text into emotionally charged speech, addressing the limitations of conventional systems in conveying feelings and intentions.
Patent Information
- Application Number
- JP2024034031
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-19
AI Technical Summary
Conventional conversation support systems fail to effectively convey users' intentions and feelings through voice, particularly for hearing-impaired individuals.
A conversation support device and system that includes a voice conversion unit to convert selected text into voice with added emotions, allowing users to specify text portions and input emotions, which are then output accordingly.
Enables clearer communication of intentions and feelings through voice by converting text into emotionally expressive speech.
Smart Images

Figure 2025135936000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a conversation support device, a conversation support system, a conversation support method, and a program. [Background technology]
[0002] A conference support system has been proposed that recognizes speech from participants and displays it on a screen (see, for example, Patent Document 1). Such a system recognizes speech from each participant, converts it into text, and displays it on the screen of a mobile device or the like used by the participant, making it useful when hearing-impaired people are participating in a conference. In such a system, hearing-impaired people have desired to use the system not only to understand the content of speech but also to convey their own intentions and feelings. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6548045 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with conventional technology, it has been difficult to convey one's intentions and feelings more clearly through voice.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a conversation support device, a conversation support system, a conversation support method, and a program that enable users to more clearly communicate their intentions and feelings. [Means for solving the problem]
[0006] (1) In order to achieve the above object, a conversation support device according to one aspect of the present invention includes a display unit that displays input text, a voice output unit that outputs voice into which the text has been converted, and a voice conversion unit that converts the voice, wherein the voice conversion unit recognizes a portion of the text displayed on the display unit that is selected by a user as designated text, and when an emotion is input about the designated text, converts the text into voice that corresponds to the selected emotion, and outputs the converted voice that corresponds to the emotion from the voice output unit.
[0007] (2) In the conversation support device according to one aspect of (1) above, the speech conversion unit may be configured to recognize the entire input text as the specified text when the user does not select any part of the text.
[0008] (3) In a conversation support device according to one aspect of (1) or (2) above, the display unit may be configured to display an image of speech other than that of the user converted into text through voice recognition, a text input area for inputting the text, an emotion-adding button image for instructing the speech to be converted, and an output button image for outputting the converted speech.
[0009] (4) In the conversation support device according to one aspect of (3) above, the display unit may not display the input text in a display area of an image obtained by recognizing speech other than that of the user and converting it into text until the converted speech corresponding to the emotion is output from the speech output unit.
[0010] (5) In order to achieve the above object, a conversation support system according to one aspect of the present invention includes a terminal and a conference support device, wherein the terminal includes a display unit that displays input text, a voice output unit that outputs voice into which the text has been converted, and a voice conversion unit that converts the voice, wherein the voice conversion unit recognizes a portion of the text displayed on the display unit that is selected by a user as designated text, and when an emotion is input about the designated text, converts the text into voice that corresponds to the selected emotion, and outputs the converted voice that corresponds to the emotion from the voice output unit, and also transmits the input text to the conference support device, and when a display image is obtained from the conference support device, displays the obtained display image on the display unit, and the conference support device, after obtaining the input text from the terminal, transmits the text to the terminal as the display image.
[0011] (6) In order to achieve the above object, one aspect of the present invention provides a conversation assistance method in which a display unit displays input text, a speech conversion unit recognizes a portion of the text displayed on the display unit that is selected by a user as designated text, and when an emotion is input about the designated text, the speech conversion unit converts the text into a speech that corresponds to the selected emotion, and outputs the converted speech that corresponds to the emotion from a speech output unit.
[0012] (7) To achieve the above object, one aspect of the present invention provides a program that causes a computer of a conversation support device to display input text on a display unit, recognize a portion of the text displayed on the display unit selected by a user as designated text, and, when an emotion is input about the designated text, convert the text into a voice that corresponds to the selected emotion, and output the converted voice that corresponds to the emotion from a voice output unit. [Effects of the Invention]
[0013] According to (1) to (7) above, you can communicate your intentions and feelings more clearly through voice. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating an overview of a conversation support system according to an embodiment and an image of a conference; [Figure 2] 10A and 10B are diagrams illustrating examples of images and operations displayed on a terminal according to an embodiment. [Figure 3] 10A and 10B are diagrams illustrating examples of images and operations displayed on a terminal according to an embodiment. [Figure 4] 1 is a block diagram illustrating an example of the configuration of a conversation support system according to an embodiment. [Figure 5] 10 is a flowchart of a process of adding emotion to text, converting it into speech, and outputting the speech according to the embodiment. [Figure 6] FIG. 10 is a diagram showing an example of how emotion selection buttons are displayed when selecting one emotion from a plurality of emotions. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).
[0016] [Overview of conversation support system, overview of this embodiment] First, an overview of the conversation support system and the present embodiment will be described. FIG. 1 is a diagram showing an overview of a conversation support system according to this embodiment and an image of a conference. Conversation support system 1 is used, for example, in a conference with two or more participants. Among participants US1 to US3, a participant US2 who is speech- or hearing-impaired (e.g., a hearing-impaired person) may also be participating in the conference. Note that all participants do not have to be in the same conference room; for example, participants may participate from another conference room or from home via a network NW.
[0017] Participants US1 and US3 who are able to speak wear a sound pickup unit 11. For example, participant US2, who is hearing impaired, carries a terminal (smartphone, tablet terminal, personal computer, etc.) 20 (conversation support device). The conference support device 30 performs speech recognition on the voice signals spoken by the participants, converts them into text, and displays the text on a first display device 60 (display device) and the terminal 20. In addition, a PC (personal computer, etc.) 80 that displays explanatory materials is connected to the second display device 70. Participant US2 uses the terminal 20, for example, placed on a table Tb.
[0018] In the prior art, hearing-impaired participants could input their speech as text, which would then be displayed on the other participants' terminals 20 or the first display device 60, thereby communicating the content of their speech to the other participants. However, text alone was frustrating because it was not possible to convey emotions. Furthermore, in the prior art, the input text was simply automatically converted into speech and output, making it impossible to communicate one's emotions. For this reason, in this embodiment, instead of just transmitting text, for example, the text is converted into voice and output in a tone that corresponds to the emotion, so that the participant can convey their emotion to the other participants through their voice. The emotions that can be added include, for example, "joy," "happiness," "dissatisfaction," "anger," "sadness," etc.
[0019] [Device operation and display example] Next, examples of images displayed on the terminal 20 and examples of operations will be described with reference to Fig. 2 and Fig. 3. Fig. 2 and Fig. 3 are diagrams showing examples of images displayed on the terminal according to this embodiment and examples of operations. 2 and 3, image g11 is an icon image representing the first participant. Image g12 is an example image in which the speech of the first participant is recognized and converted into text and displayed. Image g13 is an icon image representing, for example, a second participant who is hearing impaired. Image g14 is an example image in which text entered by the second participant using a software keyboard or an external keyboard is displayed. Image g15 is an example button image for sending the entered text to the conference supporting device 30. Image g16 is an example image of a button (emotion addition button) selected when converting the entered text into an audio signal in which emotion is added. Image g17 is an example button image for outputting an audio signal in which emotion is added.
[0020] (Step 1) As shown in image g10 of Figure 2, first, the second participant operates a software keyboard, for example, to input the text they want to speak. Image g18 shows the second participant's fingers as they input the text. The input text, "That's right. I'm glad you understand," is a continuation of the first participant's response to the second participant's previous text, "So that means...?" The user may also select the text to input from a set of phrases, for example.
[0021] (Step 2) When the second participant wants to express emotion in the input text as an audio signal, rather than simply displaying the text as text on the terminals 20 of the other participants, the second participant selects the button image g16, as shown in image g20 in FIG. 2. Image g21 shows the finger of the second participant selecting the button image g16. Note that a plurality of button images g16 may be prepared, for example, according to emotions. In this case, the button images for selecting a plurality of emotions may be displayed so as to be selectable by the participant pressing and holding the button image g16, for example, by toggling the display.
[0022] (Step 3) Next, as shown in image g30 of FIG. 3, the second participant selects, for example, by tracing a range of a phrase of text that the second participant wants to emphasize (designated text). Image g31 shows the second participant's finger selecting the range to be emphasized. Image g32 shows the selected range. The speech conversion unit 204 recognizes the text selected in this way as the designated text.
[0023] (Step 4) Next, as shown in image g40 of FIG. 3, the second participant selects button image g17, whereby the input text is converted into emphasized voice (with emotion added), and the voice is output from the terminal 20 of the second participant (image g42), and is also transmitted to the conference support device 30, whereby the text is displayed on the terminal 20 of the participant (image g43). Image g41 shows the finger of the second participant selecting button image g17. Note that the conference support device 30 may display the emotion-added text g43 in a different style from other text, for example, in bold.
[0024] In this way, on the display unit 207 of the terminal 20, an image (g12) in which speech other than that of the user is recognized and converted into text, a text input area (g14) for inputting text, an emotion addition button image (g16) for instructing the speech to be converted, and an output button image (g17) for outputting the converted speech are displayed. Note that the image examples in FIGS. 2 and 3 are merely examples, and the displayed image is not limited to these.
[0025] [Example of a conversation support system configuration] Next, an example of the configuration of the conversation support system 1 will be described. Fig. 4 is a block diagram showing an example of the configuration of a conversation support system according to this embodiment. As shown in Fig. 4, the conversation support system 1 includes, for example, a sound collection device 10, a terminal 20, a conference support device 30, an acoustic model / dictionary DB 40, a minutes / speech log storage unit 50, an emotion model 55, a first display device 60, a second display device 70, and a PC 80. The terminal 20 includes terminal 20-1, terminal 20-2, ... Hereinafter, when one of terminal 20-1 and terminal 20-2 is not specified, it will be referred to as "terminal 20".
[0026] The sound collection device 10 includes a sound collection unit 11-1, a sound collection unit 11-2, a sound collection unit 11-3, etc. Hereinafter, when one of the sound collection units 11-1, 11-2, 11-3, etc. is not specified, they will be referred to as "sound collection unit 11."
[0027] The terminal 20 includes, for example, an input unit 201, a processing unit 202, a dependency analysis unit 203, a speech conversion unit 204, an acoustic model / dictionary DB 205, an emotion model 206, a display unit 207, a communication unit 208, and a speaker 209 (speech output unit).
[0028] The conference supporting device 30 includes, for example, an acquisition unit 301, a speech recognition unit 302, a text conversion unit 303, a dependency analysis unit 304, a minutes creation unit 306, a communication unit 307, an authentication unit 308, an operation unit 309, and a processing unit 310. The processing unit 310 includes, for example, a presentation information generation unit 312.
[0029] The sound collection device 10 and the conference support device 30 are connected by wire or wirelessly. The terminal 20 and the conference support device 30 are connected by a wire or wireless network NW. The image capture device 90 and the conference support device 30 are connected by a wire or wireless network NW.
[0030] [Sound recording device] The sound collection device 10 collects audio signals uttered by the participants and outputs the collected audio signals to the conference supporting device 30. The sound collection device 10 may be a single microphone array. In this case, the sound collection device 10 has P microphones arranged at different positions. The sound collection device 10 then generates P-channel audio signals (P is an integer equal to or greater than 2) from the collected sounds and outputs the generated P-channel audio signals to the conference supporting device 30.
[0031] The sound collection unit 11 is a microphone. The sound collection unit 11 collects the audio signals of the participants, converts the collected audio signals from analog signals to digital signals, and outputs the converted digital audio signals to the conference supporting device 30. Note that the sound collection unit 11 may also be configured to output analog audio signals to the conference supporting device 30. Note that the collected audio signals include data on the time of capture.
[0032] [Device] The terminal 20 is, for example, a smartphone, a tablet terminal, a notebook computer, etc. The terminal 20 may include, for example, a motion sensor, a GPS (Global Positioning System), etc.
[0033] The input unit 201 is, for example, a touch panel sensor (including a touch panel pencil) or a keyboard provided on the display unit 207. The input unit 201 detects inputs made by participants and outputs the detection results to the processing unit 202. Note that the user of the terminal 20 inputs text by hand using, for example, a touch panel pencil, or by operating a software keyboard displayed on the display unit 207, or by operating a mechanical keyboard.
[0034] The processing unit 202 acquires the text input by the user based on the result output by the input unit 201. The processing unit 202 performs processing according to the selected button image based on the result output by the input unit 201. The processing unit 202 extracts the selected range from the input text based on the result output by the input unit 201. The processing unit 202 generates transmission information according to the result output by the input unit 201, and outputs the generated transmission information to the communication unit 208. The transmission information includes the input text information and identification information for identifying the terminal 20. The processing unit 202 acquires the text information output by the communication unit 208, converts the acquired text information into image data, and outputs the converted image data to the display unit 207. The image displayed on the display unit 207 will be described later. The processing of the terminal 20 may be performed by the processing unit 310 of the conference supporting device 30. In such a case, the application used for the processing may be located on the cloud, for example.
[0035] The dependency analysis unit 203 performs morphological analysis and dependency analysis on the input text. For the dependency analysis, for example, a shift-reduce method, a spanning tree method, or an SVM (Support Vector Machines) in a stepwise application method of chunk identification is used. If correction is necessary based on the analysis results, the dependency analysis unit 203 corrects the text by referring to the acoustic model and dictionary DB 205.
[0036] The speech conversion unit 204 uses the emotion model 206 to generate a speech signal incorporating, for example, an emphasized emotion for a selected phrase in the text that has undergone dependency analysis. The speech signal is output from the speaker 209, and the input text is transmitted to the conference support device 30 via the communication unit 208. The speech conversion unit 204 may, without using the emotion model 206, emphasize and impart emotion by, for example, outputting the speech signal at a level higher than the normal output level, changing the pitch, or applying an equalizer to emphasize specific frequency components. In this embodiment, "emphasis" refers not to speech signal processing that relatively emphasizes specific speech components to improve quality, but to processing that changes the tone of speech using well-known speech signal processing. The speech conversion unit 204 may impart emotion to speech using, for example, the emotion model 206 that has been trained using training data. The emotion model may be a network such as a DNN (Deep Neural Network) (see, for example, Japanese Patent Application Laid-Open No. 7-72900).
[0037] The acoustic model and dictionary DB 205 stores, for example, an acoustic model, a language model, a word dictionary, etc. An acoustic model is a model based on sound features, and a language model is a model of information about words and their arrangement. A word dictionary is a dictionary with a large vocabulary, such as a large vocabulary word dictionary.
[0038] A model for adding emotions as described above when converting text into a voice signal is stored in the emotion model 206. The acoustic model / dictionary DB 205 and emotion model 206 may be provided on a server (not shown) or may be stored on the cloud.
[0039] The display unit 207 displays the image data output by the processing unit 202. The display unit 207 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, or the like.
[0040] The communication unit 208 receives text information or information on minutes from the conference supporting device 30 and outputs the received information to the processing unit 202. The communication unit 208 transmits the text and transmission information output by the processing unit 202 to the conference supporting device 30.
[0041] The speaker 209 issues the audio signal output by the audio conversion unit 204 .
[0042] [Acoustic model, dictionary database, minutes, and voice log storage] The acoustic model / dictionary DB 40 stores, for example, an acoustic model, a language model, a word dictionary, etc. An acoustic model is a model based on sound features, and a language model is a model of information about words and their arrangement. A word dictionary is a dictionary with a large vocabulary, such as a large vocabulary word dictionary. The conference support device 30 may store words, etc. that are not stored in the speech recognition dictionary 13 in the acoustic model / dictionary DB 40 to update it.
[0043] The minutes and voice log storage unit 50 stores the minutes (including voice signals).
[0044] The acoustic model / dictionary DB 40 and the like may be provided on a server (not shown) or may be placed on the cloud.
[0045] The first display device 60 displays the image data output by the conference supporting device 30. The first display device 60 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, etc. The first display device 60 may be provided in the conference supporting device 30.
[0046] The second display device 70 displays image data output by the PC 80. The second display device 70 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, etc. The second display device 70 may be provided in the PC 80.
[0047] The PC 80 is, for example, any one of a personal computer, a smartphone, a tablet terminal, and the like.
[0048] [Conference support equipment] The conference support device 30 is, for example, any one of a personal computer, a server, a smartphone, a tablet terminal, etc. If the sound collection device 10 is a microphone array, the conference support device 30 further includes a sound source localization unit, a sound source separation unit, and a sound source identification unit. The conference support device 30 performs speech recognition on the voice signals uttered by the participants, for example, at predetermined intervals, and converts the voice signals into text. The conference support device 30 then transmits the text information of the converted speech content to each of the participant's terminals 20.
[0049] The acquisition unit 301 acquires the audio signal output by the sound collection unit 11, and outputs the acquired audio signal to the voice recognition unit 302. If the acquired audio signal is an analog signal, the acquisition unit 301 converts the analog signal into a digital signal, and outputs the converted digital audio signal to the voice recognition unit 302.
[0050] When there are multiple sound collection units 11, the voice recognition unit 302 performs voice recognition for each speaker who uses the sound collection unit 11. The speech recognition unit 302 acquires the speech signal output by the acquisition unit 301. The speech recognition unit 302 detects a speech signal of a speech period from the speech signal output by the acquisition unit 301. The speech period may be detected, for example, by detecting a speech signal equal to or greater than a predetermined threshold value as the speech period, or by detecting the on / off state of the sound collection unit 11. The speech recognition unit 302 may also detect the speech period using other well-known methods. The speech recognition unit 302 performs speech recognition on the detected speech period speech signal using a well-known method by referring to the acoustic model / dictionary DB 40. The speech recognition unit 302 performs speech recognition using, for example, the method disclosed in Japanese Patent Application Laid-Open No. 2015-64554. The speech recognition unit 302 outputs the recognition result and the speech signal to the text conversion unit 303. The speech recognition unit 302 outputs the recognition results and the speech signals to the text conversion unit 303 and the emotion recognition unit 311, by associating them with each other, for example, for each sentence, for each speech phrase, or for each speaker.
[0051] The text conversion unit 303 converts the speech into text based on the recognition result output by the speech recognition unit 302. The text conversion unit 303 outputs the converted text information and the speech signal to the dependency analysis unit 304. Note that the text conversion unit 303 may convert the speech into text by deleting interjections such as "ah," "um," "eh," and "well," etc.
[0052] The dependency analysis unit 304 performs morphological analysis and dependency analysis on the text information output by the text conversion unit 303. For dependency analysis, for example, SVM is used in the shift-reduce method, the spanning tree method, or the stepwise application method of chunk identification. If correction is necessary based on the analysis results, the dependency analysis unit 304 corrects the speech-recognized text information by referring to the acoustic model and dictionary DB 40 and stores the corrected text information in the minutes and speech log storage unit 50. The dependency analysis unit 304 outputs the text information and speech signal resulting from the dependency analysis to the minutes creation unit 306.
[0053] The minutes creation unit 306 creates minutes by dividing them into sections for each speaker, such as a presenter, based on the text information and audio signals output by the dependency analysis unit 304. The minutes creation unit 306 stores the created minutes and the corresponding audio signals in the minutes / audio log storage unit 50. Note that the minutes creation unit 306 may create the minutes by deleting interjections such as "ah," "er," "uh," and "well," etc.
[0054] The communication unit 307 transmits and receives information to and from the terminal 20. Information received from the terminal 20 includes, for example, a request to participate in a conference, information to start text input, information to send text, and instruction information to request the transmission of past minutes. The communication unit 307 extracts, for example, identification information for identifying the terminal 20 from the participation request received from the terminal 20, and outputs the extracted identification information to the authentication unit 308. The identification information is, for example, the serial number, Media Access Control (MAC) address, or Internet Protocol (IP) address of the terminal 20. When the authentication unit 308 outputs an instruction to permit communication participation, the communication unit 307 communicates with the terminal 20 that requested participation in the conference. When the authentication unit 308 outputs an instruction not to permit communication participation, the communication unit 307 does not communicate with the terminal 20 that requested participation in the conference. The communication unit 307 outputs the received information to the processing unit 310. The communication unit 307 transmits the text information or past minutes information output by the processing unit 310 to the terminal 20 that has requested participation.
[0055] The authentication unit 308 receives the identification information output by the communication unit 307 and determines whether or not to permit communication. Note that the conference supporting device 30, for example, accepts registration of the terminals 20 used by participants in the conference and registers them in the authentication unit 308. Depending on the determination result, the authentication unit 308 outputs to the communication unit 307 an instruction to permit or not permit participation in communication.
[0056] The operation unit 309 is, for example, a keyboard, a mouse, or a touch panel sensor provided on the first display device 60. The operation unit 309 detects the operation results of the participants and outputs the detected operation results to the processing unit 310.
[0057] The processing unit 310 performs emotion recognition processing to determine whether or not to change and present the information to be displayed on the terminal 20, the first display device 60, and the second display device 70, and based on the determination result, causes the presented information to be displayed on the terminal 20, the first display device 60, and the second display device 70. Furthermore, when the terminal 20 is operated and a fixed phrase or the like in response to laughter is received, the processing unit 310 causes the fixed phrase to be displayed on the terminals 20, the first display device 60, and the second display device 70 of the other participants. The processing unit 310 reads out the minutes from the minutes / voice log storage unit 50 in response to instruction information requesting transmission of past minutes, and outputs the information of the read out minutes to the communication unit 307. Note that the information of the minutes may include information indicating the speaker, information indicating the result of dependency analysis, etc.
[0058] The presentation information generation unit 312 acquires text information resulting from the dependency analysis output by the dependency analysis unit 304. Based on the acquired information, the presentation information generation unit 312 generates information to be displayed on the terminal 20, the first display device 60, and the second display device 70.
[0059] The voice conversion process may be performed by, for example, the conference supporting device 30. In this case, the processing unit 202 includes a voice conversion unit 313. The terminal 20 adds information indicating the selected range and information indicating that emphasis processing is to be performed to the text, and transmits the text to the conference supporting device 30. The voice conversion unit 313 generates a voice signal with emotion added to the acquired text using the emotion model 55, and transmits the generated voice signal to the terminal 20. Then, the processing unit 202 of the terminal 20 may acquire the voice signal that has been subjected to voice conversion processing by the conference supporting device 30, and issue a notification from the speaker 209.
[0060] Furthermore, when the sound collection device 10 is a microphone array, the conference support device 30 further includes a sound source localization unit, a sound source separation unit, and a sound source identification unit. In this case, the conference support device 30 has the sound source localization unit perform sound source localization using a transfer function generated in advance for the audio signal acquired by the acquisition unit 301. Then, the conference support device 30 performs speaker identification using the localization result of the sound source localization unit. The conference support device 30 performs sound source separation for the audio signal acquired by the acquisition unit 301 using the localization result of the sound source localization unit. Then, the speech recognition unit 302 of the conference support device 30 performs speech segment detection and speech recognition for the separated audio signal (see, for example, Japanese Patent Application Laid-Open No. 2017-9657). The conference support device 30 may also perform reverberation suppression processing.
[0061] [Example of processing procedure] Next, an example of a processing procedure for adding emotion to text, converting it into speech, and outputting the speech will be described. In the following example, an example will be described in which all processing is performed on the terminal 20 side, but part of the processing may also be performed on the conference support device 30 side. Figure 5 is a flowchart of the processing according to this embodiment for adding emotion to text, converting it into speech, and outputting the speech.
[0062] (Step S1) The input unit 201 acquires text input by the user.
[0063] (Step S2) The dependency analysis unit 203 refers to the acoustic model / dictionary DB 205 and performs dependency analysis on the acquired text.
[0064] (Step S3) The processing unit 202 determines whether or not an emotion-added button image (for example, image g16 in FIG. 2) has been selected. If an emotion-added button image has been selected (Step S3; YES), the processing unit 202 proceeds to the processing of Step S4. If an emotion-added button image has not been selected (Step S3; NO), the processing unit 202 repeats the processing of Step S3.
[0065] (Step S4) The processing unit 202 determines whether or not a range (specified text) is selected in the text. If a range is selected in the text (step S4; YES), the processing unit 202 proceeds to the processing of step S5. If a range is not selected in the text (step S4; NO), the processing unit 202 selects all of the input text and proceeds to the processing of step S6.
[0066] (Step S5) The speech conversion unit 204 selects the text in the selected range (designated text) from the input text.
[0067] (Step S6) The speech conversion unit 204 converts the entire input text selected in step S4 or the range of text selected in step S5 into a speech signal with added emotion using the emotion model 206. Note that the speech conversion unit 204 converts the range of text not selected in step S5 into a speech signal without adding emotion.
[0068] (Step S7) The processing unit 202 transmits the acquired text to the conference support device 30.
[0069] (Step S8) The voice conversion unit 204 issues the voice signal converted in step S6 from the speaker 209. Note that the voice signal converted in step S6 may be issued not only from the terminal 20 but also from a speaker (not shown) connected to the conference support device 30, the PC 80, or the like.
[0070] 5 are merely examples, and are not limiting. For example, some processes may be performed simultaneously.
[0071] (Variation) In the above example, an example has been described in which an "emphasized" emotion is added to text and issued as a voice signal, but two or more emotions may be selectable. 6 is a diagram showing an example of the display of emotion selection buttons for selecting one emotion from multiple emotions. When the emotion selection button image g50 is pressed and held by the user, a selection screen for "emphasize" (g61), "agree" (g62), or "disagree" (g63) is displayed, as shown in image g60. The user may select the desired emotion from these button images. In this case, the voice conversion unit 204 adds an emotion to the text in accordance with the selected emotion and converts it into a voice signal.
[0072] Note that the selection button image from among a plurality of emotions may be, for example, a plurality of emotion selection button images g50. In this case, the selection button image may represent, for example, "joy," "sadness," etc., using emoticon facial expressions.
[0073] As described above, in this embodiment, the text input by the user is acquired, the part of the input text to be subjected to the specified emphasis processing is detected, and the content of the emphasis processing is specified. Then, in this embodiment, a voice in which a part of the text is emphasized is output in accordance with the content of the emphasis processing.
[0074] As a result, according to this embodiment, even a person with a hearing impairment can communicate with the other party using voice that expresses their own emotions, simply by touching the part they want to specify and inputting their emotions.
[0075] Note that a program for implementing all or part of the functions of the terminal 20 and the conference support device 30 of the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the terminal 20 and the conference support device 30. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also refers to devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line. Alternatively, some or all of these components may be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), or an SOC (System On Chip), or may be realized by a combination of software and hardware.
[0076] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0077] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0078] 1...conversation support system, 10...sound collection device, 20, 20-1,...terminal, 30...conference support device, 40...acoustic model / dictionary DB, 50...minutes / voice log storage unit, 55...emotion model, 60...first display device, 70...second display device, 80...PC, 11, 11-1,...sound collection unit, 201...input unit, 202...processing unit, 203...dependency analysis unit, 204...voice conversion unit conversion unit, 205...acoustic model / dictionary DB, 206...emotion model, 207...display unit, 208...communication unit, 209...speaker, 301...acquisition unit, 302...speech recognition unit, 303...text conversion unit, 304...dependency analysis unit, 305...minutes creation unit, 307...communication unit, 308...authentication unit, 309...operation unit, 310...processing unit, 312...presentation information generation unit, 313...speech conversion unit
Claims
1. a display unit that displays the input text; a voice output unit that outputs the converted text; a voice conversion unit that converts the voice, The voice conversion unit a portion of the text displayed on the display unit selected by a user is recognized as a designated text, and when an emotion is input for the designated text, the text is converted into a voice corresponding to the selected emotion, and the converted voice corresponding to the emotion is output from the voice output unit; Conversation support device.
2. The voice conversion unit If the user does not select a part of the text, the entire input text is recognized as the specified text. The conversation support device according to claim 1 .
3. The display unit has: an image obtained by recognizing speech other than that of the user and converting it into text; a text input area for inputting the text; an emotion-adding button image for instructing to convert the voice; an output button image for outputting the converted audio; The conversation support device according to claim 1 or 2, wherein the following is displayed:
4. The display unit The input text is not displayed in a display area of an image obtained by recognizing speech of a person other than the user and converting it into text until the converted speech corresponding to the emotion is output from the speech output unit. The conversation support device according to claim 3 .
5. A terminal and a conference support device; The terminal a display unit that displays the input text; a voice output unit that outputs the converted text; a voice conversion unit that converts the voice, The voice conversion unit a portion of the text displayed on the display unit selected by a user is recognized as designated text, and when an emotion is input for the designated text, the text is converted into a voice corresponding to the selected emotion, the converted voice corresponding to the emotion is output from the voice output unit, and the input text is transmitted to the conference support device; When a display image is acquired from the conference support device, the acquired display image is displayed on the display unit; The conference support device After receiving the input text from the terminal, transmit the text to the terminal as the display image; Conversation support system.
6. The display displays the input text, a speech conversion unit recognizes a portion of the text displayed on the display unit selected by a user as a designated text, and when an emotion is input for the designated text, converts the text into a speech corresponding to the selected emotion, and outputs the converted speech corresponding to the emotion from a speech output unit. Conversation support methods.
7. The computer in the conversation support device The input text is displayed on the display. a portion of the text displayed on the display unit selected by a user is recognized as a designated text, and when an emotion is input for the designated text, the text is converted into a voice corresponding to the selected emotion, and the converted voice corresponding to the emotion is output from a voice output unit; program.
Citation Information
Patent Citations
Conference system, conference system control method, and program
JP6548045B2