Conversation support device, conversation support system, conversation support method, and program
The conversation support system enhances contextual understanding by allowing users to designate text and display associated utterances, addressing the challenge of chronological display in conventional systems.
Patent Information
- Application Number
- JP2024009778
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-08-06
AI Technical Summary
Conventional conversation support systems display utterances in chronological order, making it difficult to understand the context of comments on previous utterances.
A conversation support system that allows users to select a designated text, associating subsequent utterances with it, and displays the relationship between them, including thread display options.
Facilitates easier understanding of the context of comments by linking selected text with subsequent utterances, enabling clearer communication.
Smart Images

Figure 2025115298000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a conversation support device, a conversation support system, a conversation support method, and a program. [Background technology]
[0002] A conference support system has been proposed that recognizes speech from participants and displays it on a screen (see, for example, Patent Document 1). This system recognizes speech from each participant, converts it into text, and displays it on the screen of a mobile device or other device used by the participant, making it useful when hearing-impaired people are participating in a conference. This system allows the intentions of hearing-able people to be properly conveyed to hearing-impaired people. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6548045 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with conventional technology, utterances are displayed on the screen in chronological order, so even if you want to go back and comment on an earlier utterance rather than the most recent one, the utterances are only displayed in chronological order, making it difficult to understand the context.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a conversation support device, a conversation support system, a conversation support method, and a program that make it easier to understand the context of a comment when the text spoken or input by a participant is a previous utterance. [Means for solving the problem]
[0006] (1) In order to achieve the above-mentioned object, a conversation support device according to one aspect of the present invention includes an acquisition unit that acquires an utterance, a display unit that displays text based on the acquired utterance, and a control unit that causes the text based on the utterance to be displayed on the display unit, wherein the control unit recognizes a portion of the text selected by a user from among the plurality of texts displayed on the display unit as a designated text, and when an utterance is input in response to the designated text after the designated text has been recognized, causes the utterance in response to the designated text to be displayed on the display unit in association with the designated text.
[0007] (2) In one aspect of the conversation support device according to (1) above, the device may include a text conversion unit that converts the utterance into text when the utterance is a voice signal, wherein the acquisition unit acquires the voice signal when the utterance is a voice signal, and acquires the text when the utterance is text, and the control unit may assign identification information to the converted text or the acquired text, and may associate the utterance in response to the specified text with the specified text and display it on the display unit based on the identification information of the text selected by the user from among the multiple texts displayed on the display unit.
[0008] (3) In the conversation support device according to one aspect of (1) or (2), the control unit: Speech in response to the specified text may be displayed in an area in which the specified text is displayed.
[0009] (4) In the conversation support device according to any one of the above aspects (1) to (3), when an instruction to display a thread is received, the control unit may cause the display unit to display a thread showing the relationship between the utterance in response to the specified text and the specified text.
[0010] (5) In the conversation support device according to any one of the above (1) to (4), the specified text may be the entire sentence or a part of the utterance.
[0011] (6) In order to achieve the above object, a conversation support system according to one aspect of the present invention includes a terminal and a conference support device, wherein the terminal includes an acquisition unit that acquires utterances, a display unit that displays text based on the utterances, and transmits information about the acquired utterances to the conference support device, acquires the text based on the utterances, displays the text acquired from the conference support device on the display unit, recognizes a selected text portion from the displayed plurality of texts as a designated text, and, when an utterance is input in response to the designated text after the designated text has been recognized, displays information indicating the designated text and the utterance in response to the designated text. a processing unit that transmits to the conference support device comment display information that associates information indicating the specified text acquired from the conference support device with utterances in response to the specified text on the display unit, wherein the conference support device comprises: a conversion unit that, when the utterances acquired from the terminal are audio signals, converts the audio signals into text; a comment generation unit that acquires information indicating the specified text and the utterances in response to the specified text and generates the comment display information that associates the utterances in response to the specified text with the specified text; and a processing unit that transmits the generated comment display information to the terminal.
[0012] (7) In order to achieve the above-mentioned object, a conversation assistance method according to one aspect of the present invention is a conversation assistance method in which an acquisition unit acquires an utterance, a display unit displays text based on the acquired utterance, a control unit recognizes a portion of the text selected by a user from among the plurality of texts displayed on the display unit as a designated text, and when an utterance for the designated text is input while the designated text is recognized, the control unit associates the utterance for the designated text with the designated text and displays it on the display unit.
[0013] (8) In order to achieve the above-mentioned object, a conversation assistance program according to one embodiment of the present invention is a program that acquires an utterance, generates text to be displayed based on the acquired utterance, displays the text on a display unit, recognizes a portion of the text selected by a user from among the plurality of texts displayed on the display unit as a designated text, and when an utterance is input in response to the designated text after the designated text has been recognized, associates the utterance in response to the designated text with the designated text and displays it on the display unit. [Effects of the Invention]
[0014] According to the above (1) to (8), when the text spoken or input by a participant is a previous utterance, it becomes easier to understand the context of the comment. Furthermore, according to (4) above, comments and comment subjects can be displayed in a thread, making it easier to understand the context of the comments. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram illustrating an overview of a conversation support system according to an embodiment and an image of a conference; [Figure 2] 10A and 10B are diagrams illustrating examples of images and operations displayed on a terminal according to an embodiment. [Figure 3] 10A and 10B are diagrams illustrating an example of information transmitted to a terminal by a conference support apparatus according to an embodiment. [Figure 4] 1 is a block diagram showing an example of the configuration of a conversation support system according to a first embodiment. [Figure 5] FIG. 2 is a flowchart of processing performed by the conversation support system according to the first embodiment. [Figure 6] 10A and 10B are diagrams showing examples of images displayed on a terminal according to the second embodiment and examples of operations. [Figure 7] FIG. 10 is a flowchart of processing performed by the conversation support system according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings used in the following description, the scale of each component is appropriately changed so that each component can be recognized. In all the drawings for explaining the embodiments, the same reference numerals are used for components having the same functions, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX that has been calculated or processed. "XX" is any element (for example, any information).
[0017] [Overview of conversation support system, overview of this embodiment] First, an overview of the conversation support system and the present embodiment will be described. FIG. 1 is a diagram showing an overview of a conversation support system according to this embodiment and an image of a conference. Conversation support system 1 is used, for example, in a conference with two or more participants. Among participants US1 to US3, a participant US2 who is speech- or hearing-impaired (e.g., a hearing-impaired person) may also be participating in the conference. Note that all participants do not have to be in the same conference room; for example, participants may participate from another conference room or from home via a network NW.
[0018] Participants US1 and US3 who are able to speak wear a sound pickup unit 11. For example, participant US2, who is hearing impaired, carries a terminal (smartphone, tablet terminal, personal computer, etc.) 20 (conversation support device). The conference support device 30 performs speech recognition on the voice signals spoken by the participants, converts them into text, and displays the text on a first display device 60 (display device) and the terminal 20. In addition, a PC (personal computer, etc.) 80 that displays explanatory materials is connected to the second display device 70. Participant US2 uses the terminal 20, for example, placed on a table Tb.
[0019] In the prior art, when a comment is made in voice or text in response to a statement, it is not clear whether the comment is about the immediately preceding utterance or an earlier utterance. For this reason, in the embodiment, after the user selects the comment target (the text of the original utterance of the comment) on the terminal 20, the user inputs a comment on the comment target by voice or text. Then, in the embodiment, the comment is converted into text, and the text of the comment target is also displayed in association with the text in the area displaying the converted text. This makes it easier to understand the context of the comment when the text spoken or input by the participant is a previous utterance.
[0020] First Embodiment [Device operation and display example] An example of an image displayed on the terminal 20 and an example of an operation will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of an image displayed on the terminal according to this embodiment and an example of an operation. 2, images g11 to g14 are images obtained by speech recognition of the first to fourth utterances and converting them into text by the conference supporting device 30. Image g15 is an example of a speech recognition button image selected by the user of the terminal 20 when the user wants to have the speech content of the user recognized by speech recognition. Image g16 is an example of a text input button image selected by the user of the terminal 20 when the user inputs the speech content as text.
[0021] The text of the utterances displayed on the display units of the terminals 20 is transmitted by the conference support device 30 to each of the terminals 20. The conference support device 30 assigns identification numbers for identifying the utterances, for example, sequentially from the start of the conference, and transmits the identification numbers to the text to the terminals.
[0022] Image g20 (before operation) is the first screen. The first screen displays the first to fourth utterances in chronological order. In this example, the user wants to comment on the first utterance, so selects image g11 of the first utterance. Image g21 is an image of the user's finger.
[0023] (Step 1) Image g30 is the second screen. To indicate that the first utterance has been selected as the designated text, image g31 shows an example in which the display, such as the color of the text display area, is changed. The display change may be, for example, to display the text in bold, to add an underline to the text, or to display it in italics. Next, the terminal user speaks a comment and selects the voice recognition button image g15 to have their utterance recognized by voice (image g32).
[0024] Image g40 (after operation) is the third screen. The third screen is an example image displayed after a comment is made by voice on the comment target and the comment voice is recognized and converted into text. Image g41 is an example image displaying the text of the comment on the first utterance. Image g41 is the text of the first utterance selected in the second image. As in image g40, in this embodiment, in order to indicate which of the multiple utterances the comment is being made on, the commented utterance is also displayed in the comment display area g41 of the commented utterance.
[0025] 2, the example in which the user makes a comment by speaking in response to the spoken content has been described, but the present invention is not limited to this. The user may select the text input button image g16 to input a comment in text. Furthermore, in FIG. 2, the terminal 20 may display an image such as an icon or name indicating the speaker in association with the text of each utterance.
[0026] [Example of information sent to the device] Next, an example of information that the conference supporting apparatus 30 transmits to the terminal 20 will be described. 3 is a diagram showing an example of information that the conference support device according to this embodiment transmits to the terminals. As shown in FIG. 3, the conference support device 30 recognizes speech and converts it into text, and then adds identification information in the order of speech and information indicating the speaker to the text, stores the text in, for example, the minutes and voice log storage unit 50 (FIG. 4), and transmits the text to each of the terminals 20.
[0027] When a comment is made on an utterance, as in image g40 of FIG. 2 , the conference supporting device 30 assigns an identification number to the comment, as in the Nth utterance, and also assigns information indicating that it is a comment. The conference supporting device 30 assigns information indicating that the comment is to be displayed within the comment to the original text of the comment. The conference supporting device 30 then stores the "comment text, comment identification information, and information indicating that it is a comment (remark information)," as well as the "text of the original utterance of the comment, identification information of the original text of the utterance of the comment, and information indicating that it is to be displayed within the comment (remark information)," for example, in the minutes and voice log storage unit 50 and transmits them to each terminal 20. Note that when transmitting information regarding a comment, the conference supporting device 30 may also transmit information indicating the original speaker of the comment in addition to information indicating the speaker of the comment. This makes it easier to understand whose utterance the comment is being made on, according to this embodiment.
[0028] The terminal 20 updates the information to be displayed on the display unit 204 (FIG. 4) every time it acquires such information from the conference support device 30. Furthermore, if remark information is added, the terminal 20 displays the original text of the comment in the comment text display area.
[0029] [Example of a conversation support system configuration] Next, an example of the configuration of the conversation support system 1 will be described. Fig. 4 is a block diagram showing an example of the configuration of a conversation support system according to this embodiment. As shown in Fig. 4, the conversation support system 1 includes, for example, a sound collection device 10, a terminal 20, a conference support device 30, an acoustic model / dictionary DB 40, a minutes / speech log storage unit 50, a first display device 60, a second display device 70, and a PC 80. The terminals 20 include terminals 20-1, 20-2, etc. Hereinafter, when one of terminals 20-1 and 20-2 is not specified, it will be referred to as "terminal 20."
[0030] The sound collection device 10 includes a sound collection unit 11-1, a sound collection unit 11-2, a sound collection unit 11-3, etc. Hereinafter, when one of the sound collection units 11-1, 11-2, 11-3, etc. is not specified, they will be referred to as "sound collection unit 11."
[0031] The terminal 20 includes, for example, an input unit 201, a sound collection unit 202, an operation unit 203, a display unit 204, a control unit 205, and a communication unit 206.
[0032] The conference support device 30 includes, for example, an acquisition unit 301, a speech recognition unit 302, a text conversion unit 303, a dependency analysis unit 304, a minutes creation unit 306, a communication unit 307, an authentication unit 308, an operation unit 309, and a processing unit 310. The processing unit 310 includes, for example, a presentation information generation unit 312 and a comment generation unit 313.
[0033] The sound collection device 10 and the conference support device 30 are connected by wire or wirelessly. The terminal 20 and the conference support device 30 are connected by a wire or wireless network NW. The image capture device 90 and the conference support device 30 are connected by a wire or wireless network NW.
[0034] [Sound recording device] The sound collection device 10 collects audio signals uttered by the participants and outputs the collected audio signals to the conference supporting device 30. The sound collection device 10 may be a single microphone array. In this case, the sound collection device 10 has P microphones arranged at different positions. The sound collection device 10 then generates P-channel audio signals (P is an integer equal to or greater than 2) from the collected sounds and outputs the generated P-channel audio signals to the conference supporting device 30.
[0035] The sound collection unit 11 is a microphone. The sound collection unit 11 collects the audio signals of the participants, converts the collected audio signals from analog signals to digital signals, and outputs the converted digital audio signals to the conference supporting device 30. Note that the sound collection unit 11 may also be configured to output analog audio signals to the conference supporting device 30. Note that the collected audio signals include data on the time of capture.
[0036] [Device] The terminal 20 is, for example, a smartphone, a tablet terminal, a notebook computer, etc. The terminal 20 may include, for example, a motion sensor, a GPS (Global Positioning System), etc.
[0037] The input unit 201 is, for example, a touch panel sensor (including a touch panel pencil) or a keyboard provided on the display unit 204. The input unit 201 detects inputs made by participants and outputs the detection results to the control unit 205. Note that the user of the terminal 20 inputs text by hand using, for example, a touch panel pencil, or by operating a software keyboard displayed on the display unit 204, or by operating a mechanical keyboard.
[0038] The sound collection unit 202 is a microphone. The sound collection unit 202 collects the voice signal of the user of the terminal 20.
[0039] The operation unit 203 is, for example, a touch panel sensor provided on the display unit 204. The operation unit 203 detects an operation by the user.
[0040] The display unit 204 displays the image data output by the control unit 205. The display unit 204 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, or the like.
[0041] The control unit 205 acquires the text input by the user based on the result output by the input unit 201. When the voice recognition button image is selected, the control unit 205 acquires the user's voice signal. If the terminal 20 does not include the sound collection unit 202, the control unit 205 may transmit an instruction to the conference supporting device 30 to have the sound collection device 10 collect the user's voice signal when the voice recognition button image is selected. When an utterance to be commented on is selected by operating the operation unit 203, the control unit 205 outputs information indicating the selected utterance to the conference supporting device 30. The control unit 205 causes the display unit 204 to display text or the like based on the information acquired from the conference supporting device 30 (FIG. 3). All or part of the applications used by the control unit 205 for processing may be on the cloud, for example.
[0042] The communication unit 206 receives information such as the text described above from the conference supporting device 30, and outputs the received information to the control unit 205. The communication unit 206 transmits the text and audio signal output by the control unit 205 to the conference supporting device 30.
[0043] [Acoustic model, dictionary database, minutes, and voice log storage] The acoustic model / dictionary DB 40 stores, for example, an acoustic model, a language model, a word dictionary, etc. An acoustic model is a model based on sound features, and a language model is a model of information about words and their arrangement. A word dictionary is a dictionary with a large vocabulary, such as a large vocabulary word dictionary. The conference support device 30 may store words, etc. that are not stored in the speech recognition dictionary 13 in the acoustic model / dictionary DB 40 to update it.
[0044] The minutes and voice log storage unit 50 stores the minutes (including voice signals).
[0045] The acoustic model / dictionary DB 40 and the like may be provided on a server (not shown) or may be placed on the cloud.
[0046] The first display device 60 displays the image data output by the conference supporting device 30. The first display device 60 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, etc. The first display device 60 may be provided in the conference supporting device 30.
[0047] The second display device 70 displays image data output by the PC 80. The second display device 70 is, for example, a liquid crystal display device, an organic EL (electroluminescence) display device, an electronic ink display device, etc. The second display device 70 may be provided in the PC 80.
[0048] The PC 80 is, for example, any one of a personal computer, a smartphone, a tablet terminal, and the like.
[0049] [Conference support equipment] The conference supporting device 30 is, for example, any one of a personal computer, a server, a smartphone, a tablet terminal, etc. If the sound collection device 10 is a microphone array, the conference supporting device 30 further includes a sound source localization unit, a sound source separation unit, and a sound source identification unit. The conference supporting device 30 performs speech recognition on the voice signals uttered by the participants, for example, at predetermined intervals, and converts the voice signals into text. The conference supporting device 30 then transmits the text information of the converted speech content (including identification information and information indicating the speaker, see FIG. 3 ) to each of the participant's terminals 20.
[0050] The acquisition unit 301 acquires the audio signal output by the sound collection unit 11, and outputs the acquired audio signal to the voice recognition unit 302. The acquisition unit 301 also acquires the audio signal transmitted by the terminal 20, and outputs the acquired audio signal to the voice recognition unit 302. If the acquired audio signal is an analog signal, the acquisition unit 301 converts the analog signal into a digital signal, and outputs the converted digital audio signal to the voice recognition unit 302.
[0051] When there are multiple sound collection units 11, the voice recognition unit 302 performs voice recognition for each speaker who uses the sound collection unit 11. The speech recognition unit 302 acquires the speech signal output by the acquisition unit 301. The speech recognition unit 302 detects a speech signal of a speech period from the speech signal output by the acquisition unit 301. The speech period may be detected, for example, by detecting a speech signal equal to or greater than a predetermined threshold value as the speech period, or by detecting the on / off state of the sound collection unit 11. The speech recognition unit 302 may also detect the speech period using other well-known methods. The speech recognition unit 302 performs speech recognition on the detected speech period speech signal using a well-known method by referring to the acoustic model / dictionary DB 40. The speech recognition unit 302 performs speech recognition using, for example, the method disclosed in Japanese Patent Application Laid-Open No. 2015-64554. The speech recognition unit 302 outputs the recognition result and the speech signal to the text conversion unit 303. The speech recognition unit 302 outputs the recognition result and the speech signal to the text conversion unit 303, for example, by associating them with each sentence, each speech phrase, or each speaker.
[0052] The text conversion unit 303 converts the speech into text based on the recognition result output by the speech recognition unit 302. The text conversion unit 303 outputs the converted text information and the speech signal to the dependency analysis unit 304. Note that the text conversion unit 303 may convert the speech into text by deleting interjections such as "ah," "um," "eh," and "well," etc.
[0053] The dependency analysis unit 304 performs morphological analysis and dependency analysis on the text information output by the text conversion unit 303. For dependency analysis, for example, SVM is used in the shift-reduce method, the spanning tree method, or the stepwise application method of chunk identification. If correction is necessary based on the analysis results, the dependency analysis unit 304 corrects the speech-recognized text information by referring to the acoustic model and dictionary DB 40 and stores the corrected text information in the minutes and speech log storage unit 50. The dependency analysis unit 304 outputs the text information and speech signal resulting from the dependency analysis to the minutes creation unit 306.
[0054] The minutes creation unit 306 creates minutes by dividing them into sections for each speaker, such as a presenter, based on the text information and audio signals output by the dependency analysis unit 304. The minutes creation unit 306 stores the created minutes and the corresponding audio signals in the minutes / audio log storage unit 50. Note that the minutes creation unit 306 may create the minutes by deleting interjections such as "ah," "er," "uh," and "well," etc.
[0055] The communication unit 307 transmits and receives information to and from the terminal 20. Information received from the terminal 20 includes, for example, a request to participate in a conference, information to start text input, information to send text, and instruction information to request the transmission of past minutes. The communication unit 307 extracts, for example, identification information for identifying the terminal 20 from the participation request received from the terminal 20, and outputs the extracted identification information to the authentication unit 308. The identification information is, for example, the serial number, Media Access Control (MAC) address, or Internet Protocol (IP) address of the terminal 20. When the authentication unit 308 outputs an instruction to permit communication participation, the communication unit 307 communicates with the terminal 20 that requested participation in the conference. When the authentication unit 308 outputs an instruction not to permit communication participation, the communication unit 307 does not communicate with the terminal 20 that requested participation in the conference. The communication unit 307 outputs the received information to the processing unit 310. The communication unit 307 transmits the text information or past minutes information output by the processing unit 310 to the terminal 20 that has requested participation.
[0056] The authentication unit 308 receives the identification information output by the communication unit 307 and determines whether or not to permit communication. Note that the conference supporting device 30, for example, accepts registration of the terminals 20 used by participants in the conference and registers them in the authentication unit 308. Depending on the determination result, the authentication unit 308 outputs to the communication unit 307 an instruction to permit or not permit participation in communication.
[0057] The operation unit 309 is, for example, a keyboard, a mouse, or a touch panel sensor provided on the first display device 60. The operation unit 309 detects the operation results of the participants and outputs the detected operation results to the processing unit 310.
[0058] The processing unit 310 recognizes the speech of each speaker and converts it into text, and transmits the information, for example, for each utterance, to each terminal 20 via the communication unit 307. When the processing unit 310 acquires information related to a comment from the terminal 20, the processing unit 310 performs processing related to the comment. The processing unit 310 transmits comment display information, which has been processed related to the comment, to each terminal 20 via the communication unit 307. The processing unit 310 reads out the minutes from the minutes / voice log storage unit 50 in response to instruction information requesting transmission of past minutes, and outputs the information of the read out minutes to the communication unit 307. Note that the information of the minutes may include information indicating the speaker, information indicating the result of dependency analysis, etc.
[0059] The presentation information generation unit 312 acquires text information resulting from the dependency analysis output by the dependency analysis unit 304. Based on the acquired information, the presentation information generation unit 312 generates information to be displayed on the terminal 20, the first display device 60, and the second display device 70.
[0060] When the comment generation unit 313 acquires information about a comment from the terminal 20, it performs processing related to the comment. The information about the comment includes, for example, identification information of the original text of the comment, an audio signal or text information about the comment, and identification information for identifying the terminal 20 used by the user who made the comment. When the comment generation unit 313 acquires an audio signal, it outputs the audio signal to the speech recognition unit 302 and acquires text information output by the dependency analysis unit 304. As described above, the comment generation unit 313 assigns the identification information, information indicating the person who made the comment, and information indicating that the comment is a comment to the comment text. Furthermore, the comment generation unit 313 transmits the original text of the comment, the identification information of the original text of the comment, information indicating the speaker of the original text of the comment, and information indicating that the comment should be displayed in the text area of the comment, together with the original text of the comment, the identification information of the original text of the comment, information indicating the speaker of the original text of the comment, and information indicating that the comment should be displayed in the text area of the comment, as comment display information to each terminal 20 via the communication unit 307.
[0061] If the sound collection device 10 is a microphone array, the conference support device 30 further includes a sound source localization unit, a sound source separation unit, and a sound source identification unit. In this case, the conference support device 30 has the sound source localization unit perform sound source localization using a transfer function generated in advance for the audio signal acquired by the acquisition unit 301. Then, the conference support device 30 performs speaker identification using the localization result of the sound source localization unit. The conference support device 30 performs sound source separation for the audio signal acquired by the acquisition unit 301 using the localization result of the sound source localization unit. Then, the speech recognition unit 302 of the conference support device 30 performs speech segment detection and speech recognition for the separated audio signal (see, for example, Japanese Patent Application Laid-Open No. 2017-9657). The conference support device 30 may also perform reverberation suppression processing.
[0062] Note that the speech recognition, text conversion, and dependency analysis may be performed by the terminal 20. In this case, the terminal 20 may include a speech recognition unit 207, a text conversion unit 208, a dependency analysis unit 209, etc., and may execute these on the cloud.
[0063] [Example of processing procedure] Next, a description will be given of an example of a processing procedure of the conversation support system 1. Fig. 5 is a flow diagram of processing performed by the conversation support system according to this embodiment. Note that in Fig. 5, multiple terminals 20 are represented by one terminal 20.
[0064] (Step S1) The sound collection device 10 collects the voice signal of the speaker. Note that the user may input the speech as text by operating the terminal 20 (Step S1A).
[0065] (Step S2) The sound collection device 10 transmits the collected audio signal to the conference support device 30. When text is input at the terminal 20, the terminal 20 transmits the input text to the conference support device 30 (Step S2A).
[0066] (Step S3) The conference supporting apparatus 30 acquires the audio signal transmitted by the sound collection device 10. Alternatively, the conference supporting apparatus 30 acquires the text transmitted by the terminal 20.
[0067] (Step S4) The conference support device 30 performs speech recognition processing, text conversion, and dependency processing on the acquired speech signal.
[0068] (Step S5) The conference supporting device 30 generates text information by adding identification information, etc. to the converted information. Alternatively, the conference supporting device 30 generates text information by adding identification information, etc. to the text acquired from the terminal 20.
[0069] (Step S6) The conference support apparatus 30 transmits the generated text information to each of the terminals 20.
[0070] (Step S7) The terminal 20 acquires the text information from the conference support device 30.
[0071] (Step S8) The terminal 20 displays the acquired text information on the display unit 204.
[0072] (Step S9) The terminal 20 detects that the text to be commented on has been selected from the displayed text of the utterance.
[0073] (Step S10) If the user's comment is an audio signal, the terminal 20 acquires the audio signal by the audio collection unit 202. The audio signal of the comment may be collected by the audio collection device 10 (step S10A). Alternatively, if the comment is input as text, the terminal 20 acquires the input text.
[0074] (Step S11) The terminal 20 transmits the identification information of the text of the original utterance of the selected comment and the comment (audio signal or text) as comment information to the conference supporting device 30. Note that when the sound collection device 10 collects the audio signal of the comment, the terminal 20 may transmit the identification information of the text of the original utterance of the selected comment to the conference supporting device 30, and the sound collection device 10 may transmit the audio signal to the conference supporting device 30 as comment information (step S11A).
[0075] (Step S12) The conference supporting apparatus 30 acquires comment information from the terminal 20. Alternatively, the conference supporting apparatus 30 acquires comment information from the terminal 20 and acquires a voice signal of the comment from the sound collecting device 10.
[0076] (Step S13) If the acquired comment information includes a voice signal, the conference supporting apparatus 30 performs voice recognition processing, text conversion, and dependency processing on the acquired voice signal.
[0077] (Step S14) As described above, the conference support device 30 adds identification information to the comment that has been converted into text, and generates comment display information that further includes an identification number of the original text of the comment to be displayed in the comment display area.
[0078] (Step S15) The conference support apparatus 30 transmits the generated comment display information to each of the terminals 20.
[0079] (Step S16) The terminal 20 acquires comment display information from the conference support device 30.
[0080] (Step S17) The terminal 20 displays the acquired comment display information on the display unit 204. Note that the display is such that the original text of the comment, etc., is displayed in the comment display area, as in image g40 in FIG.
[0081] In the processing described using Figure 5, the processing of the first display device 60, the second display device 70, etc. is omitted, but these display devices may also be configured to display text information and comment display information obtained from the conference support device 30 in the same way as the terminal 20.
[0082] In the above example, the entire text of the comment source is selected as the designated text, but this is not limiting. The text selected as the designated text may be a portion of the utterance. For example, after selecting the text of the third utterance in FIG. 2 (image g13) "I'm free from 2 PM today," as the comment target, a portion of the text, such as "from 2 PM," may be selected. In this case, an example comment may be "Isn't it from 3 PM?"
[0083] As described above, in this embodiment, the utterance text to be commented on is displayed, and the user who wants to comment selects the comment target. Then, in this embodiment, when the user makes an utterance (voice or text input) in the selected state, the selected utterance is also displayed.
[0084] As a result, according to this embodiment, by designating part or all of the input text by touch and then inputting voice (or text) in the designated state, the designated text is displayed together with the input voice text. As a result, according to this embodiment, it becomes easy to know which part the comment is directed to, making it easier to understand the context.
[0085] Second Embodiment In the first embodiment, an example of making a comment once has been described, but there may be cases where a user wants to comment on a comment. For this reason, in this embodiment, an example of making a comment multiple times will be described. The configuration of the conversation support system 1 is the same as the configuration of FIG. 4 in the first embodiment.
[0086] [Device operation and display example] An example of an image displayed on the terminal 20 and an example of an operation will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of an image displayed on the terminal according to this embodiment and an example of an operation.
[0087] Image g100 shows a state in which a second comment is made after the first comment has been made and displayed on the display unit 204 of the terminal 20. Images g101 and g102 are example text images of speech before the comment. Image g103 is an example of a comment, and the text of the comment (image g104) is displayed in the comment display area. Image g105 is an image of the user's finger, showing the state in which the user is selecting to make a second comment. In this example, the comment (image g103) "I don't think that's right" is selected. Image g106 is an example of a voice recognition button image, and image g107 is an example of a text input button image. Image g107 is an example of a thread display button image that displays multiple comments in a thread.
[0088] Image g110 is an example of an image displayed on the terminal 20 after the second comment is made. Image g111 is an example of the text of the second comment, which is a response to the first comment (image g103), "I don't think that's right." For the second comment, the original text image g112 is also displayed in the comment display area. Furthermore, in this display, if the user wishes to display the comments in threads, the thread display button image g107 is selected (image g113).
[0089] Image g120 is an example of a display in which the first two comments are displayed as a thread. As shown in this display, the terminal 20 displays the original text image g121 of the comment as a parent, the text image g122 of the first comment in response to it, and further the text image g123 of the comment in response to the first comment in a thread-like manner. Note that when thread display is in progress, the display state of the thread display button image may be displayed differently from before selection (image g124). Such thread display information is generated by the conference support device 30 and transmitted to the terminal 20.
[0090] Note that the example of the display image shown in FIG. 6 is just an example, and an icon or name indicating the speaker may be displayed in association with each comment.
[0091] [Example of processing procedure] Next, an example of a processing procedure of the conversation support system 1 will be described. Fig. 7 is a flow diagram of processing performed by the conversation support system according to this embodiment. Note that in Fig. 7, multiple terminals 20 are represented as one. The flow in Fig. 7 is processing that starts from a state in which some utterances have already been made and the text of the utterances is displayed on the display unit 204 of the terminal 20.
[0092] (Step S101) The terminal 20 detects that the text to be commented on has been selected from the displayed text of the utterance. Next, if the user's comment is an audio signal, the terminal 20 acquires the audio signal using the sound collection unit 202. Note that the audio signal of the comment may be collected by the sound collection device 10. Alternatively, if the comment is input as text, the terminal 20 acquires the input text.
[0093] (Step S102) The terminal 20 transmits the identification information of the text of the original utterance of the selected comment and the comment (audio signal or text) as comment information to the conference supporting device 30. Note that when the sound collection device 10 collects the audio signal of the comment, the terminal 20 may transmit the identification information of the text of the original utterance of the selected comment to the conference supporting device 30, and the sound collection device 10 may transmit the audio signal to the conference supporting device 30 as comment information.
[0094] (Step S103) The conference supporting apparatus 30 acquires comment information from the terminal 20. Alternatively, the conference supporting apparatus 30 acquires comment information from the terminal 20 and acquires a voice signal of the comment from the sound collecting device 10.
[0095] (Step S104) If the acquired comment information includes a voice signal, the conference supporting device 30 performs voice recognition processing, text conversion, and dependency processing on the acquired voice signal. Next, the conference supporting device 30 adds identification information to the converted comment, and generates comment display information including an identification number of the original text of the comment to be displayed in the comment display area.
[0096] (Step S105) The conference support apparatus 30 transmits the generated comment display information to each of the terminals 20.
[0097] (Step S106) The terminal 20 acquires the comment display information from the conference supporting apparatus 30. Subsequently, the terminal 20 displays the acquired comment display information on the display unit 204.
[0098] (Steps S107 to S112) The terminal 20 and the conference supporting apparatus 30 perform the processes of steps S107 to S112 as the second comment process in the same manner as steps S101 to S106.
[0099] (Step S113) The terminal 20 detects that the thread display button image has been selected. Subsequently, the terminal 20 transmits to the conference supporting device 30 information indicating that thread display has been instructed.
[0100] (Step S114) The conference supporting apparatus 30 acquires information indicating that a thread display instruction has been issued from the terminal 20.
[0101] (Step S115) The conference supporting apparatus 30 generates thread display information to be displayed as, for example, image g120 in FIG.
[0102] (Step S116) The conference support device 30 transmits the generated thread display information to the terminal 20 that transmitted the instruction.
[0103] (Step S117) The terminal 20 acquires thread display information from the conference support device 30.
[0104] (Step S118) The terminal 20 displays the thread display information on the display unit 204, for example, as in image g120 in FIG.
[0105] 7, an example of two comments has been described, but the number of comments may be one or more, and may be three or more. The conversation support system 1 may display a thread even if there is only one comment. Furthermore, when storing the minutes in the minutes and voice log storage unit 50, the conference support device 30 may store the minutes together with information about the number of comments and thread information.
[0106] As described above, in this embodiment, even when a comment is made on a comment, the comment and the commented subject are displayed in association with each other. Furthermore, in this embodiment, the comments are displayed in a thread.
[0107] As a result, according to this embodiment, even if a comment is made multiple times, the relationship between the comment and the commented subject becomes easy to understand, making it easier to understand the context of the comment. Furthermore, according to this embodiment, since thread display is possible, even if a comment is made on a comment, it becomes easy to understand the original remark that the comment originated from and the chronological order of the comments.
[0108] In the above-described embodiments, the examples have been described in which instructions and operations are given on the terminal 20, but they may also be given on the PC 80, for example.
[0109] Note that a program for implementing all or part of the functions of the terminal 20 and the conference support device 30 of the present invention may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform all or part of the processing performed by the terminal 20 and the conference support device 30. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" also includes a WWW system equipped with a homepage provision environment (or display environment). The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. The term "computer-readable recording medium" also refers to devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that acts as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line. Alternatively, some or all of these components may be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), or an SOC (System On Chip), or may be realized by a combination of software and hardware.
[0110] The program may also be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program that realizes part of the above-mentioned functions. Furthermore, the program may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0111] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0112] 1...conversation support system, 10...sound collection device, 20, 20-1,...terminal, 30...conference support device, 40...acoustic model / dictionary DB, 50...minutes / voice log storage unit, 60...first display device, 70...second display device, 80...PC, 11, 11-1,...sound collection unit, 201...input unit, 202...sound collection unit, 203...operation unit, 204...display unit, 205...control unit, 206...communication unit, 301...acquisition unit, 302...speech recognition unit, 303...text conversion unit, 304...dependency analysis unit, 306...minutes creation unit, 307...communication unit, 308...authentication unit, 309...operation unit, 310...processing unit, 312...presentation information generation unit, 313...comment generation unit
Claims
1. an acquisition unit that acquires an utterance; a display unit that displays text based on the acquired utterance; a control unit that causes the display unit to display the text based on the utterance, The control unit a portion selected by a user from among the plurality of texts displayed on the display unit is recognized as a designated text; When an utterance for the designated text is input in a state where the designated text has been recognized, the conversation support device displays the utterance for the designated text on the display unit in association with the designated text.
2. a text conversion unit that converts the speech into text when the speech is a voice signal; the acquisition unit acquires the voice signal when the utterance is a voice signal, and acquires the text when the utterance is text; The control unit assigning identification information to the converted text or the acquired text, and displaying an utterance for the designated text on the display unit in association with the designated text based on the identification information of the text selected by the user from the plurality of texts displayed on the display unit; The conversation support device according to claim 1 .
3. The control unit displaying an utterance in response to the specified text in an area in which the specified text is displayed; The conversation support device according to claim 1 or 2.
4. The control unit When a thread display instruction is given, the display unit displays a thread of a relationship between the utterance in response to the specified text and the specified text. The conversation support device according to claim 1 or 2.
5. The specified text is the entire utterance or a portion thereof. The conversation support device according to claim 1 or 2.
6. A terminal and a conference support device; The terminal an acquisition unit that acquires an utterance; a display unit that displays text based on the utterance; a processing unit that transmits information about the acquired utterance to the conference supporting device, acquires the text based on the utterance, displays the text acquired from the conference supporting device on the display unit, recognizes a selected text portion from the displayed plurality of texts as a designated text, and when an utterance in response to the designated text is input in a state where the designated text has been recognized, transmits information indicating the designated text and the utterance in response to the designated text to the conference supporting device, and displays comment display information on the display unit that associates the information indicating the designated text acquired from the conference supporting device with the utterance in response to the designated text, The conference support device a conversion unit that converts the speech signal into text when the speech acquired from the terminal is a speech signal; a comment generating unit that acquires information indicating the designated text and utterances in response to the designated text, and generates the comment display information in which the utterances in response to the designated text are associated with the designated text; a processing unit that transmits the generated comment display information to the terminal, Conversation support system.
7. an acquisition unit acquires an utterance; a display unit that displays text based on the acquired speech; a control unit that recognizes a portion selected by a user from the plurality of texts displayed on the display unit as a designated text; when an utterance for the designated text is input in a state in which the designated text has been recognized, the control unit causes the utterance for the designated text to be displayed on the display unit in association with the designated text. Conversation support methods.
8. Acquire the utterance, generating text to be displayed based on the acquired speech; Displaying the text on a display unit; a portion selected by a user from among the plurality of texts displayed on the display unit is recognized as a designated text; When an utterance for the designated text is input in a state in which the designated text has been recognized, the utterance for the designated text is displayed on the display unit in association with the designated text. program.
Citation Information
Patent Citations
Conference system, conference system control method, and program
JP6548045B2