Conversation support device, method, and program
Patent Information
- Application Number
- JP2025029325
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-09-07
AI Technical Summary
【0014】 本発明の会話支援装置によれば、聴者の会話のテキストデータを所定の発言単位で時系列に並べて表示させるとともに、聴者の会話の途中でろう者の手話動画または入力テキストデータの入力を受け付けた場合には、ろう者の手話動画または入力テキストデータの入力が終了した時点以降において、聴者の発言の時系列のテキストデータの中に、ろう者の手話動画の翻訳テキストデータまたは入力テキストデータを挿入して時系列に並べて表示させるようにしたので、ろう者と聴者の会話の繋がりを確認することができ、意志疎通を円滑にし、適切に会話を進めることができる。
Smart Images

Figure 2026142295000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a conversation assistance device, method and program for supporting conversations between deaf persons and hearing persons.
Background Art
[0002] Conventionally, for example in conferences, there are cases where both hearing persons (persons with normal hearing) and deaf persons (persons with hearing impairment) participate. In such conferences, when hearing persons and deaf persons communicate with each other, text, as a common language, has been used.
[0003] Specifically, regarding utterances made by hearing persons, audio data is acquired via a microphone, converted into text, and displayed as text. On the other hand, regarding utterances made by deaf persons, text is input by the deaf person using a device such as a personal computer, or the sign language motion of the deaf person is photographed, and the captured sign language video is translated into text data, thereby being displayed as text. This allows hearing persons and deaf persons to hold a conversation while viewing the text display.
[0004] Patent Document 1 proposes a method of acquiring audio data of conversation participants, illustrating the utterance status of the participants in chronological order, and outputting the audio data of a selected portion.
[0005] Patent Document 2 proposes a method in which a sign language video obtained by photographing a deaf person's sign language is translated by a computer and output as text or audio.
[0006] Patent Document 3 proposes a method in which when displaying audio data of utterances during a conference as text in chronological order, the text display of a predetermined utterance is pinned, a question is asked, and an answer is obtained.
Prior Art Literature
Patent Literature
[0007]
Patent Literature 1
[0008] Here, for example, when a meeting is held with two or more hearing people and a deaf person participating, as mentioned above, the hearing people's statements are displayed as text by converting audio data to text, and the deaf people's statements are displayed as text by translating sign language gestures or by text input. However, there are cases where the meeting proceeds through audio communication among the hearing people.
[0009] In this situation, when a deaf person speaks using sign language or text input, it takes time to complete the sign language action and its translation, or for the deaf person to complete their text input. During this time, the hearing people may move on to a different topic in their conversation, resulting in the deaf person's statement being displayed as text with a delay compared to the time of their speech.
[0010] As a result, a problem arises where it becomes impossible to understand what topic a deaf person's statement was referring to.
[0011] Patent documents 1 to 3 do not take into account the timing discrepancy in the text display of deaf people's statements as described above. Patent document 3 proposes a method of having a conversation while displaying the agenda by pinning the items to be discussed in a meeting, but it does not take into account the two-way conversation between deaf and hearing people as described above.
[0012] In view of the above circumstances, the present invention aims to provide a conversation support device, method, and program that can confirm the connection between conversations between deaf and hearing people, facilitate communication, and enable conversations to proceed appropriately. [Means for solving the problem]
[0013] The present invention is a conversation support device that supports conversation between a first speaker who converses using voice and a second speaker who converses without using voice, comprising: a first text data receiving unit that converts the voice data of the first speaker's conversation into text data and receives the first text data; a second text data receiving unit that receives as second text data text data text converted from images of sign language actions uttered by the second speaker or text data input by the second speaker; and the conversation of the first speaker and the sign language actions of the second speaker. The system includes a time information acquisition unit that acquires time information of the start of a conversation by inputting text data, and a display control unit that, based on the time information acquired by the time information acquisition unit, arranges and displays the first text data of the first speaker's conversation in chronological order in predetermined utterance units, and, if the second speaker's conversation begins in the middle of the first speaker's conversation, inserts the second text data into the first text data of the first speaker's utterances in chronological order after the reception of the second text data of the second speaker's conversation has finished, and displays them arranged in chronological order. [Effects of the Invention]
[0014] According to the conversation support device of the present invention, the text data of a hearing person's conversation is displayed in chronological order in predetermined utterance units. Furthermore, if a deaf person's sign language video or input text data is received in the middle of the hearing person's conversation, the translated text data of the deaf person's sign language video or input text data is inserted into the chronological text data of the hearing person's utterances after the input of the deaf person's sign language video or input text data is completed, and displayed in chronological order. This allows the connection between the conversation between the deaf and the hearing person to be confirmed, facilitating smooth communication and enabling the conversation to proceed appropriately. [Brief explanation of the drawing]
[0015] [Figure 1]Block diagram showing a schematic configuration of a conference support system using an embodiment of the conversation support device of the present invention [Figure 2] Flowchart for explaining the processing flow of the conference support system shown in FIG. 1 [Figure 3] Diagram showing an example of a deaf conference screen displayed on a terminal device for the deaf [Figure 4] Diagram showing an example of a conference screen displayed on a shared terminal device [Figure 5] Diagram showing an example of a main conversation screen within a text display frame [Figure 6] Diagram showing a display example of "There is a statement" [Figure 7] Diagram showing an example of pop-up display [Figure 8] Diagram showing an example of a sub-conversation screen [Figure 9] Diagram showing an example of the main conversation screen when returning from the sub-conversation screen MODE FOR CARRYING OUT THE INVENTION
[0016] Hereinafter, a conference support system 1 using an embodiment of the conversation support device of the present invention will be described in detail with reference to the drawings. FIG. 1 is a block diagram showing a schematic configuration of the conference support system 1 of the present embodiment.
[0017] The conference support system 1 of the present embodiment is a system that supports conversation of deaf people in conferences by displaying as text the conversations between hearing people and deaf people. In the present embodiment, description is given as a system that supports conversation in a conference participated by two hearing people and one deaf person, but the number of hearing people and deaf people participating in the conversation is not limited thereto, and may be further increased.
[0018] As shown in Figure 1, the conference support system 1 of this embodiment comprises a conversation support device 10, a terminal device for deaf people 20, and a shared terminal device 30. The conversation support device 10 is a cloud server, and the terminal device for deaf people 20 and the shared terminal device 30 are installed in a conference room used by both hearing and deaf people. In this embodiment, the conversation support device 10 is implemented by a cloud server, but it is not limited to this, and it may be implemented by a local server within the company, or the functions of the conversation support device 10 may be implemented in the shared terminal device 30 of this embodiment and used for both purposes.
[0019] The conversation support device 10, the terminal device for deaf persons 20, and the shared terminal device 30 are connected via a communication line such as an internet connection or a LAN (Local Area Network) connection, and are configured to exchange various information with each other. In Figure 1, only one terminal device for deaf persons 20 is shown, but in reality, it is preferable to provide one terminal device for each deaf person participating in the conversation. Also, in this embodiment, two hearing persons use one shared terminal device 30, but hearing persons may also use individual terminal devices.
[0020] The following provides a more detailed explanation of each component of the conference support system 1.
[0021] As shown in Figure 1, the conversation support device 10 includes a first text data receiving unit 11, a second text data receiving unit 12, a time information acquisition unit 13, and a display control unit 14.
[0022] The first text data receiving unit 11 receives audio data of a listener's conversation output from the shared microphone / speaker 31 connected to the shared terminal device 30, and converts that audio data into text data (hereinafter referred to as the first text data). For the conversion from audio data to text data, known speech recognition techniques using machine learning such as recurrent neural networks (RNNs), convolutional neural networks (CNNs), and transformer models can be used.
[0023] The second text data receiving unit 12 receives sign language videos of deaf people captured using the camera function of the terminal device for deaf people 20, translates the input sign language videos into text data, and accepts it as second text data. As a method for translating sign language videos into text data, for example, a learning model that has been pre-trained to recognize the relationship between various sign language videos and their corresponding text data can be prepared, and the sign language videos received by the second text data receiving unit 12 can be input to this learning model to convert them into text data.
[0024] Furthermore, the second text data receiving unit 12 can also receive input text data entered by a deaf person in the deaf terminal device 20 as second text data. The deaf terminal device 20 in this embodiment is configured to allow sign language photography and text input, and deaf people participate in conversations by using sign language or text input.
[0025] The time information acquisition unit 13 acquires time information for the start of conversations between hearing people and deaf people through sign language actions or text data input. In this embodiment, the time information acquisition unit 13 acquires time information for the start of conversations between hearing people, specifically the time when the first text data receiving unit 11 started receiving audio data. The time information acquisition unit 13 also acquires time information for the start of conversations between deaf people, specifically the time when the second text data receiving unit 12 started receiving sign language videos and input text data.
[0026] The start time for receiving audio data is defined as the time when audio data is received if no audio data has been received for a predetermined period of time. In other words, even if the audio data is interrupted for a period shorter than the above period, it is assumed that a series of statements continues. If the audio data is interrupted for a period longer than the above period, the next audio data is received as audio data based on new statements, and this is considered the start time for receiving audio data. In this embodiment, statements from a predetermined start time for receiving audio data to the start time for receiving the next audio data correspond to one unit of statements according to the present invention.
[0027] Furthermore, in this embodiment, the start time for receiving sign language videos is determined by an operation input by a deaf person on the deaf terminal device 20. Specifically, the start time for receiving sign language videos and the start time for receiving input text data are obtained when the "start / end sign language button" displayed on the deaf terminal device 20 is selected by the deaf person. When the "start / end sign language button" is selected by the deaf person on the deaf terminal device 20, the selection information is output from the deaf terminal device 20 to the conversation support device 10, and the time information acquisition unit 13 detects the time when the above selection information is received and sets this as the start time for receiving sign language videos. Also, the time when the "start / end sign language button" is selected again after the start time for receiving sign language videos is set as the end time for the sign language videos. In this embodiment, the sign language statements from the start time to the end time for receiving sign language videos correspond to one unit of statements in the present invention.
[0028] Furthermore, the start time for receiving sign language videos is not limited to the method described above. For example, the start time for receiving sign language videos may be when the second text data receiving unit 12 detects the start of a sign language action from the captured image, or when it detects a video of a predetermined sign language action (for example, a specific pose or hand movement). Alternatively, the start / stop of a sign language action may be detected by keyboard operation by a deaf person (for example, pressing the space key or enter key), and the second text data receiving unit 12 may acquire the time when that keyboard operation was performed.
[0029] Furthermore, the start time for receiving input text data is obtained in this embodiment by detecting the time when a deaf person starts inputting text on the terminal device 20 for deaf users. In this embodiment, the utterance from the start time to the end time for receiving input text data corresponds to one unit of utterance according to the present invention. The end time for input text data is obtained, for example, by detecting the press of Enter after text input. In addition, while the system initially uses the time information of when the audio data of the listener's conversation was first received as the time information of when the listener's conversation began, it is not limited to this. The system may also use the time information of when the first text data, obtained by converting the above audio data into text, was first received as the time information of when the listener's conversation began. Furthermore, while we have chosen to obtain the time information of when the sign language video was first received as the time information of when the conversation between deaf people began, we are not limited to this. We may also obtain the time information of when the second text data, which is a translation of the above sign language video, was first received as the time information of when the conversation between deaf people began.
[0030] Based on the time information acquired by the time information acquisition unit 13, the display control unit 14 causes the first text data of the hearing person's speech converted by the first text data reception unit 11 and the second text data (translated text data of the sign language video or input text data) received by the second text data reception unit 12 to be displayed in chronological order on the deaf terminal device 20 and the shared terminal device 30.
[0031] However, as mentioned above, even if a deaf person starts using sign language, hearing people may not notice, or it may take time to translate the sign language into text data, causing the hearing conversation to progress ahead and making it difficult to display the deaf person's statements in chronological order at the appropriate time.
[0032] The display control unit 14 of this embodiment can display the inventions of deaf people in chronological order at appropriate timings to address the above-mentioned problems, but the method of controlling its display will be described in detail later.
[0033] The conversation support device 10 includes a CPU (Central Processing Unit), semiconductor memory such as ROM (Read Only Memory) and RAM (Random Access Memory), storage such as a hard disk, and a communication I / F (Interface).
[0034] An embodiment of the conversation support program of the present invention is installed in the storage of the conversation support device 10. When this conversation support program is started by the CPU, the functions of the above-described parts of the conversation support device 10 are executed.
[0035] Furthermore, in this embodiment, the functions of each part are executed by executing a conversation support program using a CPU. However, some or all of the functions executed by the conversation support program may be composed of hardware such as an ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other electrical circuits.
[0036] Next, we will describe the terminal device 20 for deaf people.
[0037] As described above, the terminal device 20 for deaf people is used by deaf people and consists of, for example, a personal computer, but it may also consist of a mobile terminal such as a tablet or smartphone.
[0038] As shown in Figure 1, the terminal device 20 for the deaf includes a control unit 21, a display unit 22, a storage unit 23, an input unit 24, and a shooting unit 25.
[0039] The control unit 21 controls the entire terminal device 20 for deaf people. In particular, the control unit 21 performs functions such as displaying text of conversations between hearing and deaf people, capturing and outputting sign language videos of deaf people, and receiving text input of deaf people's statements and outputting the input text data by activating the conference support application installed in the memory unit 23.
[0040] Furthermore, the control unit 21 displays images captured by the imaging unit 25, images captured by the shared camera 32 connected to the shared terminal device 30, and electronic files opened on the deaf terminal device 20 and the shared terminal device 30 on the display unit 22.
[0041] As described above, the display unit 22 displays second text data of conversations between hearing and deaf individuals, images captured by the camera unit 25 and the shared camera 32, and shared electronic files.
[0042] The memory unit 23 has the aforementioned conference support application installed.
[0043] The input unit 24 accepts various settings inputs from deaf individuals, and in particular accepts text input of deaf individuals' statements.
[0044] The imaging unit 25 includes a CMOS (Complementary Metal Oxide Semiconductor) camera, a CCD (Charge Coupled Devices) camera, and an imaging optical system, and captures the sign language of deaf people. The sign language video data captured by the imaging unit 25 is stored in the storage unit 23 and then output to the conversation support device 10 by the control unit 21.
[0045] The meeting support application may be installed on the storage unit 23 as in this embodiment, or it may be an application provided via a web browser.
[0046] Next, we will describe the shared terminal device 30.
[0047] The shared terminal device 30 is for shared use by the listeners participating in the meeting, and may consist of, for example, a personal computer, but may also consist of a mobile terminal such as a tablet or smartphone.
[0048] As shown in Figure 1, the shared terminal device 30 includes a control unit 33, a display unit 34, a storage unit 35, and an input unit 36.
[0049] The control unit 33 controls the entire shared terminal device 30. In particular, the control unit 33 controls the display of text of conversations between hearing and deaf individuals, the display of images captured by the shared camera 32, and the operation of the shared microphone and speaker 31 by launching the conference support application installed in the memory unit 35.
[0050] Furthermore, the control unit 33 displays on the display unit 34 images of deaf people's sign language captured by the imaging unit 25 of the terminal device for deaf people 20, as well as electronic files opened on the terminal device for deaf people 20 and the shared terminal device 30.
[0051] As described above, the display unit 34 displays second text data of conversations between hearing and deaf people, deaf sign language images, and shared electronic files.
[0052] As described above, the memory unit 35 has the conference support application installed.
[0053] The input unit 36 accepts various settings inputs from the listener.
[0054] The meeting support application may be installed in the storage unit 35 as in this embodiment, or it may be an application provided via a web browser.
[0055] A shared monitor 37 is connected to the shared terminal device 30, and the shared monitor 37 displays the same content as that displayed on the display unit 34 of the shared terminal device 30.
[0056] Next, the processing flow of the conference support system 1 of this embodiment will be explained with reference to the flowchart shown in Figure 2.
[0057] First, the conference support application is launched on the deaf terminal device 20 and the shared terminal device 30, and they connect to the cloud server, which is the conversation support device 10 (S10). Then, the conference support application displays a conference screen for deaf users as shown in Figure 3 on the deaf terminal device 20, and the conference support application displays a conference screen as shown in Figure 4 on the shared terminal device 30.
[0058] The deaf conference screen displays a text display frame T on the left, and a shared screen frame C, a text input box TB, and a sign language start / end button B on the right. The shared screen frame C displays images of deaf individuals taken by the deaf terminal device 20, images of the conference room taken by the shared camera 32, and electronic files opened on the deaf terminal device 20 or the shared terminal device 30.
[0059] The meeting screen is the same as the deaf meeting screen, except for the text input box TB and the sign language start / end button B.
[0060] When the hearing person begins to speak (S12), the hearing person's speech is input into the shared microphone / speaker 31, and the audio data is output from the shared terminal device 30 and input into the conversation support device 10 (S14).
[0061] When voice data is input to the conversation support device 10, the first text data receiving unit 11 converts it into first text data (S16), and the time information acquisition unit 13 acquires the time information of the start of voice data reception (S18).
[0062] The display control unit 14 then links the converted first text data with the time information of the start of audio data reception and outputs it to the deaf terminal device 20 and the shared terminal device 30. The deaf terminal device 20 and the shared terminal device 30 then display the first text data in chronological order in the text display frame T based on the time information linked to the first text data (S20).
[0063] Figure 5 shows an example of the display within the text display frame T of the deaf terminal device 20 and the shared terminal device 30. In the example shown in Figure 5, the statements of hearing individuals "Suzuki" and "Inoue" are transcribed and displayed as text, with time information added to each statement. If the shared microphone / speaker 31 has a speaker recognition function, the name of the person who spoke is also added and displayed.
[0064] When a deaf person has questions, concerns, or opinions regarding a hearing person's statements, they select the sign language start / end button B displayed on the deaf terminal device 20 to confirm, ask questions, or express their opinions. The display control unit 14 of the conversation support device 10 monitors the selection of the sign language start / end button B on the deaf terminal device 20 (S22, NO), and when the sign language start / end button B is selected (S22, YES), it displays "Sign language input in progress" as shown in Figure 5. The time information acquisition unit 13 also detects that the sign language start / end button B has been selected and acquires the time information at that time. In this embodiment, the selection of this sign language start / end button B corresponds to a predetermined operation input in the present invention.
[0065] Then, the deaf person selects the sign language start / end button B and begins sign language (S24). While the deaf person is sign language, the hearing person's conversation continues, and the hearing person's statements are displayed chronologically as shown in Figure 6. To make it clear when the deaf person is speaking, the display control unit 14 displays "There is a statement" based on the time information at the moment the sign language start / end button B is selected, as shown in Figure 6. "There is a statement" is displayed in chronological order with the hearing person's statements. When the sign language start / end button B is selected by the deaf person, the display "There is a statement" may also be output as machine voice from the shared microphone / speaker 31.
[0066] When a deaf person begins using sign language, the deaf terminal device 20 captures the sign language video, and the captured video is output from the deaf terminal device 20 to the conversation support device 10. When the sign language video is input to the conversation support device 10, the second text data receiving unit 12 translates it into text data, and it is received as second text data (S26). When the deaf person has completed the sign language actions, they select the sign language start / end button B displayed on the deaf terminal device 20 again.
[0067] When the sign language start / end button B is selected again (S28, YES), the conversation support device 10 stores the second text data translated in the second text data receiving unit 12 and the time information acquired in S22 when the sign language start / end button B was selected at the start of the sign language operation, linked together (S30).
[0068] Then, the display control unit 14 of the conversation support device 10 compares the time information associated with the first text data of the listener's statement with the time information associated with the second text data, and inserts the second text data between the first text data of multiple listeners' statements so that they are arranged in chronological order (S32).
[0069] Then, the display control unit 14 displays a pop-up message as shown in Figure 7 on the deaf terminal device 20 and the shared terminal device 30, in addition to the main conversation screen shown in Figure 6, depending on the selection of the sign language start / end button B when the sign language action is completed (S34).
[0070] The display control unit 14 displays the second text data described above within the pop-up display screen, and also displays the first text data of the listener's statement, which is linked to the time information before and after the time information of the second text data. Furthermore, as shown in Figure 7, the display control unit 14 displays a "Open Thread" button near the display of the second text data within the pop-up display screen.
[0071] When the "Open Thread" button is selected in the pop-up display screen of the shared terminal device 30 (S36, YES), a sub-conversation screen as shown in Figure 8 is displayed on the deaf terminal device 20 and the shared terminal device 30, in addition to the main conversation screen shown in Figure 6 and the pop-up display shown in Figure 7 (S38). On the sub-conversation screen, text is displayed in the same way as in the pop-up display, and the first text data of what the hearing person said is added sequentially in chronological order immediately after the second text data of what the deaf person said (S40). Also, a "Return to Main Conversation" button is displayed at the bottom of the sub-conversation screen.
[0072] When the conversation regarding the deaf person's statement is finished, the "Return to Main Conversation" button is selected on the shared terminal device 30 (S42, YES), and the user returns to the main conversation screen. At this time, the main conversation screen displays the entire conversation, including the content exchanged in the sub-conversation screen, as shown in Figure 9 (S44).
[0073] In the above explanation, deaf individuals communicate using sign language, but as mentioned above, they may also communicate by entering text into the text input box TB. In this case, the input text data will be displayed as the second text data instead of the translated text data of the sign language video mentioned above.
[0074] Furthermore, in the above description, the time information acquisition unit 13 acquires the time information of the start of sign language video reception when the selection of the sign language start / end button B is detected. However, as an alternative method, for example, the terminal device for deaf persons 20 may accept the selection of a predetermined text data of a statement that a deaf person wishes to comment on or ask a question from among the first text data of a hearing person's statement displayed, and when that selection is detected, the start time of reception of the sign language video or input text data may be acquired.
[0075] In this case, the time information for the start of receiving sign language video or input text data is obtained as the time immediately following (for example, 1 second later) the time information of the hearing person's statement selected by the deaf person. Then, when the second text data is displayed along with the first text data of the hearing person's statement, the second text data is displayed immediately following the first text data of the hearing person's statement selected by the deaf person.
[0076] According to the conference support system 1 of the above embodiment, the first text data of the hearing person's conversation is displayed in chronological order in predetermined utterance units. Furthermore, if a deaf person's conversation begins in the middle of a hearing person's conversation, the second text data of the deaf person's conversation is inserted into the first text data of the hearing person's utterances in chronological order after the reception of the second text data of the deaf person's conversation has finished, and the system is displayed in chronological order. This allows for confirmation of the connection between the conversations of the deaf and hearing people, facilitating smooth communication and enabling the conversation to proceed appropriately.
[0077] Furthermore, in the above-mentioned conference support system 1, the second text data received in the middle of the hearing person's conversation, and the first text data of the hearing person's statement that is chronologically before or after the start of receiving the second text data, are extracted and displayed as a pop-up. This makes it easy to confirm which statement by the hearing person the deaf person is responding to.
[0078] Furthermore, in the sub-conversation screen, the first text data of the hearing person's statement received after the pop-up display is now added and displayed, allowing for smoother conversations regarding the deaf person's statements.
[0079] Furthermore, in the conference support system 1 of the above embodiment, after the sub-conversation screen, the system returns to the main conversation screen and inserts the second text data of the deaf person into the first text data of the hearing person's statements in chronological order, thereby displaying the statements of the entire conversation in chronological order, making it easy to grasp the overall flow of the conversation in the meeting.
[0080] Furthermore, in the conference support system 1 of the above embodiment, when the selection of the sign language start / end button by a deaf person is detected, the system acquires the time information of the start of the sign language video reception. Therefore, the time information of the start of the deaf person's speech can be obtained through simple operation and processing.
[0081] Furthermore, in the conference support system 1 of the above embodiment, when the start of text input by a deaf person is detected, the time information of the start of receiving the input text data is obtained. Therefore, the time information of the start of a deaf person's speech can be obtained through simple operation and processing.
[0082] Furthermore, in the conference support system 1 of the above embodiment, if the system is configured to acquire the start time of receiving a sign language video when the start of a deaf person's sign language is detected or when a pre-set sign language video is detected, the time information of when the deaf person's speech began can be automatically acquired.
[0083] Furthermore, in the conference support system 1 of the above embodiment, if the system accepts the selection of a statement from a hearing person by a deaf person, and acquires the time information immediately following the time information of the first text data of the selected statement as the time information of the start of the conversation by the deaf person, then the deaf person can make questions or other statements in response to any statement made by a hearing person, and these can be displayed as text in chronological order.
[0084] Furthermore, the conference support system 1 of the above embodiment supports conversations between hearing and deaf people, and deaf people include those with hearing impairments, those with hearing loss, those with hearing impairments, those with hearing difficulties, and deaf-mute people. In addition, the second conversationalist who converses without using voice in the present invention is not limited to deaf people, but also includes hearing-mute people and hearing people who are unable to speak due to the surrounding environment being quiet, and other people who need to communicate using text.
[0085] Furthermore, the present invention is not limited to the embodiments described above, and the components can be modified and implemented in practice without departing from the spirit of the invention. Also, various inventions can be formed by appropriate combinations of the multiple components disclosed in the embodiments. For example, all the components shown in the embodiments may be combined as appropriate. It goes without saying that various modifications and applications are possible without departing from the spirit of the invention.
[0086] The following further notes are disclosed regarding the present invention.
[0087] (Note 1) The present invention is a conversation support device that supports conversation between a first speaker who converses using voice and a second speaker who converses without using voice, comprising: a first text data receiving unit that converts the voice data of the first speaker's conversation into text data and receives the first text data; a second text data receiving unit that receives as second text data text data text converted from images of sign language actions uttered by the second speaker or text data input by the second speaker; and the conversation of the first speaker and the sign language actions of the second speaker. The system includes a time information acquisition unit that acquires time information of the start of a conversation by inputting text data, and a display control unit that, based on the time information acquired by the time information acquisition unit, arranges and displays the first text data of the first speaker's conversation in chronological order in predetermined utterance units, and, if the second speaker's conversation begins in the middle of the first speaker's conversation, inserts the second text data into the first text data of the first speaker's utterances in chronological order after the reception of the second text data of the second speaker's conversation has finished, and displays them arranged in chronological order.
[0088] (Note 2) In the conversation support device described in Appendix 1, the display control unit, after the completion of receiving the second text data, can extract the second text data received in the middle of the conversation of the first speaker and the first text data that is chronologically before or after the start of the conversation of the second speaker, and display a different display from the chronological display of the first text data.
[0089] (Note 3) In the conversation support device described in Appendix 2, the display control unit can further display the first text data received after the above-mentioned different displays.
[0090] (Note 4) In the conversation support device described in Appendix 3, when the display control unit receives a predetermined instruction input after the above-mentioned different displays, it can insert the second text data into the first text data of the first speaker's utterance in chronological order and display them side by side.
[0091] (Note 5) In the conversation support device described in any of Appendix 1 to 4, the time information acquisition unit can acquire time information of the start of the conversation of the second speaker when the start of the second speaker's sign language action is detected or when a pre-set sign language video is detected.
[0092] (Note 6) In the conversation support device described in any of Appendix 1 to 4, the time information acquisition unit can acquire time information of the start of the conversation of the second conversationor when a predetermined operation input by the second conversationor is detected.
[0093] (Note 7) In the conversation support device described in any of the appendices 1 to 6, the time information acquisition unit can acquire time information of the start of the conversation by the second conversationalist when the start of text input by the second conversationalist is detected.
[0094] (Note 8) In the conversation support device described in any of Appendix 1 to 7, when the selection of a first text data for a predetermined statement by the first speaker is detected, the time information acquisition unit can acquire the time information immediately following the time information of the selected first text data as the time information of the start of the conversation of the second speaker.
[0095] (Note 9) The present invention provides a conversation support method that supports a conversation between a first speaker who converses using voice and a second speaker who converses without using voice. The method converts the voice data of the first speaker's conversation into text data, receives the first text data, receives text data converted from an image of a sign language gesture uttered by the second speaker or text data input by the second speaker as second text data, obtains time information for the start of the conversation between the first speaker and the second speaker based on their sign language gestures or text data input, and displays the first text data of the first speaker's conversation in chronological order based on the obtained time information. If the second speaker's conversation begins in the middle of the first speaker's conversation, the method inserts the second text data into the first text data of the first speaker's utterances in chronological order after the reception of the second text data of the second speaker's conversation has finished, and displays them in chronological order.
[0096] (Note 10) The present invention provides a conversation support program that supports a conversation between a first speaker who converses using voice and a second speaker who converses without using voice. The program includes the steps of: converting the voice data of the first speaker's conversation into text data and receiving the first text data; receiving text data converted from an image of a sign language gesture uttered by the second speaker or text data input by the second speaker as second text data; acquiring time information for the start of the conversation between the first speaker and the second speaker based on their sign language gestures or text data input; and, based on the acquired time information, displaying the first text data of the first speaker's conversation in chronological order by predetermined utterance units, and, if the second speaker's conversation begins in the middle of the first speaker's conversation, inserting the second text data into the first text data of the first speaker's utterances in chronological order after the reception of the second text data of the second speaker's conversation has finished and displaying them in chronological order. [Explanation of symbols]
[0097] 1. Meeting support system 10 Conversation support device 11. First text data receiving unit 12. Second Text Data Reception Unit 13 Time information acquisition section 14 Display Control Unit 20 Terminal devices for the deaf 21 Control Unit 22 Display section 23 Memory section 24 Input section 25 Photography Department 30 Shared terminal equipment 31 Shared Microphone / Speaker 32 Shared Cameras 33 Control Unit 34 Display section 35 Storage section 36 Input section 37 Shared Monitor B Sign Language Start / End Button C Shared screen frame T Text display frame TB Text Input Box
Claims
1. A conversation support device that assists in conversation between a first speaker who converses using voice and a second speaker who converses without using voice, A first text data receiving unit converts the audio data of the conversation of the first speaker into text data and receives the first text data, A second text data receiving unit receives text data converted from images of sign language gestures uttered by the second speaker, or text data input by the second speaker, as second text data. A time information acquisition unit that acquires time information of the start of the conversation between the first speaker and the second speaker through sign language actions or text data input, A conversation support device comprising: a display control unit that, based on time information acquired by the time information acquisition unit, arranges and displays the first text data of the conversation of the first speaker in chronological order in predetermined utterance units; and, if the conversation of the second speaker begins in the middle of the conversation of the first speaker, inserts the second text data into the first text data of the chronological order of the utterances of the first speaker and displays it in chronological order after the reception of the second text data of the conversation of the second speaker has finished.
2. The conversation support device according to claim 1, wherein the display control unit, after the time when it has finished receiving the second text data, extracts the second text data received in the middle of the conversation of the first speaker and the first text data that are chronologically before or after the start of the conversation of the second speaker, and displays a different display from the time-series display of the first text data.
3. The conversation support device according to claim 2, wherein the display control unit further displays the first text data received after the different display.
4. The conversation support device according to claim 3, wherein when the display control unit receives a predetermined instruction input after the different displays, it inserts the second text data into the first text data of the first speaker's utterance in chronological order and displays them in chronological order.
5. The conversation support device according to claim 1, wherein the time information acquisition unit acquires time information of the start of the conversation of the second speaker when the start of the second speaker's sign language movement is detected or when a pre-set sign language video is detected.
6. The conversation support device according to claim 1, wherein the time information acquisition unit acquires time information of the start of the conversation of the second conversationor when a predetermined operation input by the second conversationor is detected.
7. The conversation support device according to claim 1, wherein the time information acquisition unit acquires time information of the start of the conversation by the second conversationor when the start of text input by the second conversationor is detected.
8. The conversation support device according to claim 1, wherein when the time information acquisition unit detects the selection of the first text data of a predetermined statement by the first speaker, it acquires the time information immediately following the time information of the selected first text data as the time information of the start of the conversation of the second speaker.
9. A conversation support method that supports a conversation between a first speaker who converses using voice and a second speaker who converses without using voice, The audio data of the conversation of the first speaker is converted into text data, and the first text data is received. The system accepts text data converted from images of sign language gestures uttered by the second speaker, or text data entered by the second speaker, as the second text data. The system obtains time information for the start of the conversation between the first speaker and the second speaker, based on their sign language actions or text data input. A conversation support method that, based on the acquired time information, displays the first text data of the conversation of the first speaker arranged in chronological order in predetermined utterance units, and, if the conversation of the second speaker begins in the middle of the conversation of the first speaker, inserts the second text data into the first text data of the chronological order of the utterances of the first speaker and displays them arranged in chronological order after the reception of the second text data of the conversation of the second speaker has finished.
10. A conversation support program that assists in conversations between a first speaker who converses using voice and a second speaker who converses without using voice, The steps include converting the audio data of the conversation of the first speaker into text data and receiving the first text data, The steps include receiving text data converted from an image of a sign language gesture uttered by the second speaker, or text data input by the second speaker, as second text data, The steps include obtaining time information for the start of the conversation between the first speaker and the second speaker through sign language actions or text data input, A conversation support program that causes a computer to perform the following steps: display the first text data of the conversation of the first speaker in chronological order based on the acquired time information, and, if the conversation of the second speaker begins in the middle of the conversation of the first speaker, insert the second text data into the first text data of the chronological order of the first speaker's statements and display them in chronological order after the reception of the second text data of the conversation of the second speaker has finished.
Citation Information
Patent Citations
Sign language conversation support system
JP2017204067A
Conversation support device, conversation support system, conversation support method, and program
JP2021157139A
Communication support system, communication support device, communication support method, and program
JP2024074245A