Display control system, display control method, and program
The system addresses the delay in displaying translation results by using a confirmation request as a trigger, ensuring timely and accurate display of translated speech during video conferences.
Patent Information
- Application Number
- JP2021199424
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-09-22
- Estimated Expiration
- 2041-12-08
AI Technical Summary
Existing technologies delay the display of translation results during video conferences, making it difficult for participants to grasp the translation in a timely manner due to the reliance on the absence of speech input for several seconds as a trigger.
A system that includes a voice data receiving means, confirmation request receiving means, and translation result display control means to timely display translation results by using a confirmation request as a trigger, superimposing character strings representing the translation results on captured images.
Enables immediate display of translation results, reducing the time lag and enhancing participants' understanding of translated speech during video conferences.
Smart Images

Figure 0007742767000001 
Figure 0007742767000002 
Figure 0007742767000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a display control system, a display control method, and a program. [Background technology]
[0002] There is a technology that displays an image in which a character string representing the translation result of a voice is superimposed on an image captured by a capturing unit. As an example of such a technology, Patent Document 1 describes a video conference system that displays on a screen a video signal of video data in which a video signal of a speaker is captured and translated into text information superimposed thereon, the text information being obtained by translating the data of the speaker's voice.
[0003] There is also a technology that uses the absence of recognizable speech input for several seconds as a trigger to start translation of speech that has been input up to that point. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-153408 Summary of the Invention [Problem to be solved by the invention]
[0005] In the technology described in Patent Document 1, if the lack of recognizable speech input for several seconds triggers the start of translation of speech input up to that point, it takes a certain amount of time from the time of speech input until the translation result of that speech is displayed, which makes it difficult for video conference participants to grasp the translation result in a timely manner.
[0006] The present invention has been made in consideration of the above-mentioned problems, and one of its objectives is to provide a display control system, a display control method, and a program that can timely display the translation results of input speech. [Means for solving the problem]
[0007] The display control system of the present invention includes a voice data receiving means for receiving voice data representing voice input by a speaker, a confirmation request receiving means for receiving a confirmation request output in response to a predetermined operation performed by the speaker, a translation control means for controlling the start of translation of the voice represented by the voice data received up to the reception of the confirmation request as a trigger, and a translation result display control means for displaying on a display unit a screen on which an image captured by a photographing unit is arranged with a character string representing the translation result of the voice represented by the voice data received up to the reception of the confirmation request superimposed on an image captured by a photographing unit.
[0008] In one aspect of the present invention, the device further includes a voice recognition result display control means for causing the display unit to display a screen on which an image is placed, in which a character string representing the voice recognition result of the voice represented by the voice data is superimposed on an image captured by the imaging unit, and the voice recognition result display control means causes the display unit to display a screen on which an image is placed, in which a character string representing the voice recognition result of the voice represented by the accepted voice data is superimposed on an image captured by the imaging unit, prior to receiving the confirmation request.
[0009] In one aspect of the present invention, the translation result display control means causes the display unit to display a screen on which an image captured by the imaging unit is superimposed with both a character string representing the speech recognition result of the speech represented by the speech data received up to the time of receiving the confirmation request, and a character string representing the translation result of the speech represented by the speech data received up to the time of receiving the confirmation request.
[0010] In one aspect of the present invention, the translation result display control means further includes an image output unit that outputs an image to a video conferencing system in which a character string is superimposed on an image captured by the imaging unit, and the translation result display control means causes the display unit to display the screen generated by the video conferencing system.
[0011] In another aspect of the present invention, the voice data receiving means receives from the terminal the voice data representing the voice input to the terminal by the speaker, the confirmation request receiving means receives the confirmation request transmitted from the terminal in response to a predetermined operation performed by the speaker on the terminal, the translation result display control means displays on a display unit provided in the terminal a character string representing the translation result of the voice represented by the voice data received up to the time of receiving the confirmation request, and the translation result display control means displays on a display unit provided in the client device a screen on which an image captured by the imaging unit is superimposed with a character string representing the translation result of the voice represented by the voice data received up to the time of receiving the confirmation request.
[0012] Alternatively, the voice data receiving means receives from the client device the voice data representing the voice input to the client device by the speaker, the confirmation request receiving means receives the confirmation request transmitted from the client device in response to a predetermined operation performed by the speaker on the client device, and the translation result display control means displays on the display unit provided in the client device a screen on which an image is placed in which an image captured by the photographing unit is superimposed with a string of characters representing the translation result of the voice represented by the voice data received up until the reception of the confirmation request.
[0013] In one aspect of the present invention, the translation control means controls the start of translation of the speech represented by the speech data received up until the reception of the confirmation request into multiple languages, and the translation result display control means causes the display unit to display a screen on which an image captured by the imaging unit is superimposed with a string of characters representing the translation result of the speech represented by the speech data for each of the multiple languages.
[0014] In addition, the display control method of the present invention includes the steps of accepting voice data representing voice input by a speaker, accepting a confirmation request output in response to a predetermined operation performed by the speaker, controlling the start of translation of the voice represented by the voice data accepted up to the acceptance of the confirmation request as a trigger, and displaying on a display unit a screen on which an image is placed in which a character string representing the translation result of the voice represented by the voice data accepted up to the acceptance of the confirmation request is superimposed on an image captured by a capturing unit.
[0015] In addition, the program of the present invention causes a computer to execute the following steps: accepting voice data representing voice input by a speaker; accepting a confirmation request output in response to a predetermined operation performed by the speaker; using the acceptance of the confirmation request as a trigger to control the start of translation of the voice represented by the voice data accepted up to the acceptance of the confirmation request; and displaying on a display unit a screen on which an image captured by a photographing unit is arranged with a character string representing the translation result of the voice represented by the voice data accepted up to the acceptance of the confirmation request superimposed thereon. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a diagram showing an example of the overall configuration of a video conference translation system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating an example of the rear surface of a terminal according to an embodiment of the present invention. [Figure 3A] FIG. 2 is a diagram illustrating an example of a configuration of a terminal according to an embodiment of the present invention. [Figure 3B] FIG. 2 is a diagram illustrating an example of a configuration of a client device according to an embodiment of the present invention. [Figure 3C] FIG. 2 is a diagram illustrating an example of a configuration of a relay device according to an embodiment of the present invention. [Figure 3D] 1 is a diagram illustrating an example of a configuration of a voice processing system according to an embodiment of the present invention. [Figure 4] FIG. 10 is a diagram illustrating an example of a video conference screen. [Figure 5] FIG. 10 is a diagram illustrating an example of a voice recognition result image. [Figure 6] FIG. 10 is a diagram illustrating an example of a video conference screen. [Figure 7] FIG. 10 is a diagram showing an example of a translation result image. [Figure 8A] FIG. 2 is a functional block diagram showing an example of functions implemented in a terminal, a relay device, and a voice processing system according to an embodiment of the present invention. [Figure 8B] FIG. 2 is a functional block diagram showing an example of functions implemented in a client device according to an embodiment of the present invention. [Figure 9] FIG. 10 is a flowchart showing an example of a flow of processing performed in a relay device according to an embodiment of the present invention. [Figure 10] FIG. 10 is a flowchart showing an example of a flow of processing performed in a relay device according to an embodiment of the present invention. [Figure 11] FIG. 10 is a flowchart showing an example of the flow of processing performed in a client device according to an embodiment of the present invention. [Figure 12] FIG. 10 is a diagram illustrating an example of a video conference screen. [Figure 13] FIG. 10 is a diagram illustrating an example of the configuration of a client device according to a modified example of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0017] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0018] FIG. 1 is a diagram showing an example of the overall configuration of a video conference translation system 1 according to this embodiment. FIG. 2 is a diagram showing an example of the back side of a terminal 10 according to this embodiment. FIG. 3A is a diagram showing an example of the configuration of a terminal 10 according to this embodiment. FIG. 3B is a diagram showing an example of the configuration of a client device 12 according to this embodiment. FIG. 3C is a diagram showing an example of the configuration of a relay device 14 according to this embodiment. FIG. 3D is a diagram showing an example of the configuration of a speech processing system 16 according to this embodiment.
[0019] 1, the videoconference translation system 1 according to this embodiment includes a terminal 10, a client device 12, a relay device 14, a speech processing system 16, and a videoconference system 18. The terminal 10, the client device 12, the relay device 14, the speech processing system 16, and the videoconference system 18 are connected to a computer network 20 such as the Internet. Therefore, the terminal 10, the client device 12, the relay device 14, the speech processing system 16, and the videoconference system 18 are able to communicate with each other via the computer network 20.
[0020] The terminal 10 according to this embodiment is a computer used by a user participating in a video conference such as a remote conference. As shown in Fig. 3A, the terminal 10 according to this embodiment includes, for example, a processor 10a, a storage unit 10b, a communication unit 10c, an operation unit 10d, an image capturing unit 10e, a touch panel 10f, a microphone 10g, and a speaker 10h.
[0021] The processor 10a is a program-controlled device such as a microprocessor that operates according to a program installed in the terminal 10, for example.
[0022] The storage unit 10b is, for example, a storage element such as a ROM or a RAM, etc. The storage unit 10b stores programs executed by the processor 10a, etc.
[0023] The communication unit 10c is a communication interface for exchanging data with the relay device 14 via, for example, the computer network 20. The communication unit 10c may include a wireless communication module that communicates with the computer network 20, such as the Internet, via a mobile phone line including a base station. The communication unit 10c may also include a wireless LAN module that communicates with the computer network 20, such as the Internet, via a Wi-Fi (registered trademark) router or the like.
[0024] The operation unit 10d is an operation member such as a button or a touch sensor that outputs the content of an operation performed by the user to the processor 10a. FIG. 1 shows, as examples of the operation unit 10d, a translation button 10da that is pressed when inputting speech to be translated, a power button 10db for turning the power on and off, and a volume adjustment unit 10dc for adjusting the volume of the speech output from the speaker 10h. The translation button 10da is located below the touch panel 10f provided on the front surface of the terminal 10. The power button 10db and the volume adjustment unit 10dc are located on the right side surface of the terminal 10.
[0025] The photographing unit 10e is, for example, a photographing device such as a digital camera, etc. As shown in Fig. 2, the terminal 10 according to this embodiment is provided with the photographing unit 10e on the rear surface.
[0026] The touch panel 10f is, for example, an integrated device that includes a touch sensor and a display such as a liquid crystal display or an organic EL display. The touch panel 10f is provided on the front surface of the terminal 10 and displays screens generated by the processor 10a.
[0027] The microphone 10g is, for example, a voice input device that converts received voice into an electrical signal. Here, the microphone 10g may be a dual microphone built into the terminal 10 and equipped with a noise canceling function that makes it easy to recognize human voices even in a crowded place.
[0028] The speaker 10h is, for example, an audio output device that outputs audio. Here, the speaker 10h may be a dynamic speaker that is built into the terminal 10 and can be used in noisy places.
[0029] The client device 12 according to this embodiment is a general computer such as a smartphone, a tablet terminal, a personal computer, etc. As shown in Fig. 3B, the client device 12 according to this embodiment includes, for example, a processor 12a, a storage unit 12b, a communication unit 12c, an operation unit 12d, an image capturing unit 12e, a display 12f, a microphone 12g, and a speaker 12h.
[0030] The client device 12 according to this embodiment is used by a user of the terminal 10 when a video conference such as a remote conference is being held. That is, in this embodiment, the user of the terminal 10 and the user of the client device 12 are the same person.
[0031] The processor 12a is a program-controlled device such as a CPU that operates according to a program installed in the client device 12, for example.
[0032] The storage unit 12b is, for example, a storage element such as a ROM or RAM, a solid state drive, a hard disk drive, etc. The storage unit 12b stores programs executed by the processor 12a, etc.
[0033] The communication unit 12c is a communication interface such as a network board, a wireless LAN module, etc. The communication unit 12c exchanges data with the relay device 14 and the video conference system 18 via the computer network 20, for example.
[0034] The operation unit 12d is a user interface such as a keyboard or a mouse, which receives operation input from the user and outputs a signal indicating the content of the input to the processor 12a.
[0035] The image capturing unit 12e is a capturing device such as a digital video camera. The image capturing unit 12e is placed in a position where it can capture an image of the user of the client device 12. The image capturing unit 12e according to this embodiment is capable of capturing moving images.
[0036] The display 12f is a display device such as a liquid crystal display or an organic EL display, and displays various images according to instructions from the processor 12a.
[0037] The microphone 12g is, for example, an audio input device that converts received audio into an electrical signal.
[0038] The speaker 12h is, for example, an audio output device that outputs audio.
[0039] In this embodiment, the relay device 14 is a computer system such as a server computer that relays, for example, voice data representing voice input to the terminal 10, a voice recognition result character string representing the voice recognition result of the voice, a translation result character string representing the translation result of the voice, etc. Note that the video conference translation system 1 may include one or more relay devices 14. As shown in FIG. 3C , the relay device 14 according to this embodiment includes, for example, a processor 14a, a storage unit 14b, and a communication unit 14c.
[0040] The processor 14a is a program-controlled device such as a CPU that operates according to a program installed in the relay device 14, for example.
[0041] The storage unit 14b is, for example, a storage element such as a ROM or RAM, a solid state drive, a hard disk drive, etc. The storage unit 14b stores programs executed by the processor 14a, etc.
[0042] The communication unit 14c is, for example, a communication interface such as a network board, etc. The communication unit 14c exchanges data with the terminal 10, the client device 12, and the voice processing system 16 via the computer network 20, for example.
[0043] The speech processing system 16 is a computer system such as a server computer that performs speech processing such as speech recognition of speech represented by received speech data and translation of the speech. The speech processing system 16 may be composed of one computer or multiple computers. As shown in FIG. 3D, the speech processing system 16 according to this embodiment includes, for example, a processor 16a, a storage unit 16b, and a communication unit 16c.
[0044] The processor 16 a is a program-controlled device such as a CPU that operates according to a program installed in the audio processing system 16 .
[0045] The storage unit 16b is, for example, a storage element such as a ROM or RAM, a solid state drive, a hard disk drive, etc. The storage unit 16b stores programs executed by the processor 16a, etc.
[0046] The communication unit 16c is a communication interface such as a network board, and transmits and receives data to and from the relay device 14 via the computer network 20, for example.
[0047] The video conference system 18 is, for example, a general video conference system that realizes a video conference such as a remote conference with multiple participants. In this embodiment, for example, client software related to the video conference system 18 that operates in cooperation with the video conference system 18 is installed in the client device 12.
[0048] In this embodiment, a video conference such as a remote conference in which multiple participants including users of the terminal 10 and the client device 12 participate is held in advance by the functions of the video conference system 18.
[0049] In this embodiment, a pre-translation language, which is the language of the speech to be input to the terminal 10, and a post-translation language, which is the language into which the speech is translated, are set in advance by the user performing a predetermined operation on the terminal 10. In the following description, it is assumed that Japanese is set as the pre-translation language and English is set as the post-translation language.
[0050] In this embodiment, speech recognition processing is performed on speech input via microphone 10g between the time the user presses and releases a predetermined button (here, for example, translation button 10da) provided on terminal 10. Triggered by the user releasing his / her finger from translation button 10da, translation processing is performed on speech input via microphone 10g between the time the user presses and releases translation button 10da. Hereinafter, the state in which translation button 10da is pressed will be referred to as an input-on state, and the state in which translation button 10da is not pressed will be referred to as an input-off state.
[0051] In this embodiment, for example, while the input is on, the voice recognition process is executed on the voices input from the time when the input is changed from the input off state to the input on state until the present time. Then, a voice recognition result character string representing the voice recognition result for the voices is displayed on the display 12f of the client device 12 and also on the touch panel 10f of the terminal 10.
[0052] 4 is a diagram showing an example of a video conference screen 30, which is a screen for a video conference such as a remote conference, displayed on the display 12f of the client device 12. As shown in FIG. 4, in this embodiment, for example, the video conference screen 30 includes a superimposed image 32 in which a voice recognition result character string is superimposed on a captured image of a user who is a speaker who has input voice into the terminal 10, and is displayed on the display 12f. The captured image according to this embodiment is, for example, an image captured by the image capturing unit 12e. Note that the captured image according to this embodiment may also be an image captured by the image capturing unit 10e.
[0053] Fig. 5 is a diagram showing an example of a voice recognition result image 34 displayed on the touch panel 10f of the terminal 10. As shown in Fig. 5, in this embodiment, the same character strings as those arranged on the video conference screen 30 shown in Fig. 4 are also arranged on the voice recognition result image 34.
[0054] As described above, in this embodiment, while the terminal 10 is in the input-on state, the voice recognition process is successively performed on the voices input from the time when the terminal 10 changed from the input-off state to the input-on state to the present time. Then, every time the voice recognition process is performed, the voice recognition result character string displayed on the touch panel 10f or the display 12f is updated.
[0055] Then, when the user releases his / her finger from the translation button 10da and the terminal 10 enters the input-off state, a confirmation request is sent from the terminal 10 to the relay device 14. Then, a final speech recognition process is performed on the speech input while the terminal 10 was in the input-on state. Then, a translation process is performed on the speech recognition result character string representing the result of the speech recognition process, and a translation result character string is generated by translating the speech recognition result character string. Here, for example, a translation result character string is generated that is an English character string obtained by translating the speech recognition result character string, which is a Japanese character string.
[0056] The speech recognition character string and translation result character string thus generated are displayed on the display 12f of the client device 12 and also on the touch panel 10f of the terminal 10.
[0057] For example, as shown in FIG. 6, a video conference screen 30 is displayed on the display 12f, in which a superimposed image 32 is arranged in which a voice recognition result character string and a translation result character string are superimposed on a photographed image of a user who is the speaker who inputted voice into the terminal 10.
[0058] 7, a translation result image 36 is displayed on the touch panel 10f, in which the same character string as the speech recognition result character string arranged on the video conference screen 30 shown in FIG. 6 and the same character string as the translation result character string arranged on the video conference screen 30 shown in FIG. 6 are arranged.
[0059] For the sake of convenience, FIG. 6 shows a video conference screen 30 in which the translation result character string is easily visible. However, in reality, the displayed translation result character string can be difficult to see depending on the background image (here, for example, a photographed image) on the screen where the translation result character string is displayed, and the user who is speaking may not be able to accurately grasp the translation result.
[0060] In this embodiment, as shown in FIG. 7, a translation result image 36 in which the same character string as the translation result character string shown in FIG. 6 is arranged is displayed on the touch panel 10f of the terminal 10.
[0061] In this way, according to this embodiment, the user can accurately understand the translation result of the voice input by the user.
[0062] For the sake of convenience of explanation, FIGS. 4 and 6 show the video conference screen 30 on which the speech recognition result character string is easily visible. However, in reality, the displayed speech recognition result character string may be difficult to see depending on the background image (here, for example, a photographed image) of the screen on which the speech recognition result character string is placed, and the user who is speaking may not be able to accurately grasp the speech recognition result.
[0063] In this embodiment, as shown in Fig. 5, a speech recognition result image 34 in which the same character string as the speech recognition result character string shown in Fig. 4 is arranged is displayed on the touch panel 10f of the terminal 10. Also, as shown in Fig. 7, a translation result image 36 in which the same character string as the speech recognition result character string shown in Fig. 6 is arranged is displayed on the touch panel 10f of the terminal 10.
[0064] In this way, according to this embodiment, the user can accurately understand the speech recognition result of the speech input by the user.
[0065] Furthermore, in this embodiment, the reception of a confirmation request by relay device 14 is used as a trigger to start translation of the speech represented by the speech data received up until the reception of the confirmation request. This shortens the time from the start of speech input to the translation of the speech, compared to when translation of speech input up to that point is started as a trigger when no recognizable speech is input for several seconds. In this way, according to this embodiment, the translation result of input speech can be displayed in a timely manner.
[0066] The functions of the video conference translation system 1 according to this embodiment and the processes executed by the video conference translation system 1 will be further described below.
[0067] Fig. 8A is a functional block diagram showing an example of functions implemented in the terminal 10, relay device 14, and voice processing system 16 according to this embodiment. Fig. 8B is a functional block diagram showing an example of functions implemented in the client device 12 according to this embodiment.
[0068] It should be noted that the terminal 10, relay device 14, and voice processing system 16 according to this embodiment do not need to implement all of the functions shown in Fig. 8A, and functions other than the functions shown in Fig. 8A may be implemented. It should be noted that the client device 12 according to this embodiment does not need to implement all of the functions shown in Fig. 8B, and functions other than the functions shown in Fig. 8B may be implemented.
[0069] As shown in FIG. 8A, the terminal 10 according to this embodiment functionally includes, for example, an operation input reception unit 40, a voice input reception unit 42, a voice buffer 44, an input transmission unit 46, a character string reception unit 48, and a display control unit 50. The operation input reception unit 40 is implemented mainly using the processor 10a, the operation unit 10d, and the touch panel 10f. The voice input reception unit 42 is implemented mainly using the processor 10a and the microphone 10g. The voice buffer 44 is implemented mainly using the storage unit 10b. The input transmission unit 46 and the character string reception unit 48 are implemented mainly using the communication unit 10c. The display control unit 50 is implemented mainly using the processor 10a and the touch panel 10f.
[0070] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 10a, which is installed in the terminal 10, which is a computer. This program is supplied to the terminal 10 via a computer-readable information storage medium such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.
[0071] As shown in FIG. 8B , the client device 12 according to this embodiment functionally includes, for example, a voice input accepting unit 60, a character string receiving unit 62, a captured image acquiring unit 64, a superimposed image generating unit 66, a video conference client unit 68, a voice output control unit 70, and a display control unit 72. The voice input accepting unit 60 is implemented mainly using the processor 12a and the microphone 12g. The character string receiving unit 62 is implemented mainly using the communication unit 12c. The captured image acquiring unit 64 is implemented mainly using the processor 12a and the image capturing unit 12e. The superimposed image generating unit 66 is implemented mainly using the processor 12a. The video conference client unit 68 is implemented mainly using the processor 12a and the communication unit 12c. The voice output control unit 70 is implemented mainly using the processor 12a and the speaker 12h. The display control unit 72 is implemented mainly using the processor 12a and the display 12f.
[0072] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 12a, which is installed in the client device 12, which is a computer. This program is supplied to the client device 12 via a computer-readable information storage medium, such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.
[0073] 8A, the relay device 14 according to this embodiment functionally includes, for example, an input relay unit 80, an audio buffer 82, and a character string relay unit 84. The input relay unit 80 and the character string relay unit 84 are implemented primarily in the communication unit 14c. The audio buffer 82 is implemented primarily in the storage unit 14b.
[0074] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 14a, which is installed in the relay device 14, which is a computer. This program is supplied to the relay device 14 via a computer-readable information storage medium, such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.
[0075] 8A, the speech processing system 16 according to this embodiment functionally includes, for example, a speech recognition unit 90 and a translation unit 92. The speech recognition unit 90 and the translation unit 92 are implemented mainly using the processor 16a and the communication unit 16c.
[0076] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 16a, which is installed in the speech processing system 16, which is a computer. This program is supplied to the speech processing system 16 via a computer-readable information storage medium such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.
[0077] In this embodiment, the operation input accepting unit 40 of the terminal 10 accepts operation input to the terminal 10, such as an operation in which the user presses the translation button 10da with a finger or an operation in which the user releases the finger from the translation button 10da.
[0078] In this embodiment, for example, the voice input receiving unit 42 of the terminal 10 receives voice input by a speaker via the microphone 10g while the terminal 10 is in an input-on state.
[0079] In this embodiment, the audio buffer 44 of the terminal 10 stores audio data representing audio input via the microphone 10g, for example.
[0080] In this embodiment, the input transmitting unit 46 of the terminal 10 transmits, for example, an operation signal corresponding to the operation input received by the operation input receiving unit 40 to the relay device 14.
[0081] In addition, in this embodiment, the input transmitting unit 46 transmits, for example, audio data representing audio input to the terminal 10 to the relay device 14.
[0082] In this embodiment, for example, in response to the terminal 10 changing from an input-off state to an input-on state, the input transmitting unit 46 transmits a communication start request to the relay device 14. Then, audio data representing audio input via the microphone 10g from the time the terminal 10 changes from the input-off state to the input-on state until communication between the relay device 14 and the terminal 10 is established is accumulated in the audio buffer 44.
[0083] Then, when communication between the relay device 14 and the terminal 10 is established (i.e., the terminal 10 is connected to the relay device 14), the input transmission unit 46 transmits the audio data stored in the audio buffer 44 to the relay device 14. Generally, for example, audio data representing two seconds of audio stored in the audio buffer 44 is transmitted in about 0.1 seconds.
[0084] After all of the voice data stored in the voice buffer 44 has been transmitted to the relay device 14, the input transmitting unit 46 transmits a stream of voice data packets representing the voice received by the voice input receiving unit 42 to the relay device 14 while the terminal 10 is in the input-on state. In this case, the voice data packets are transmitted directly to the relay device 14 in real time without being stored in the voice buffer 44. Note that the voice data packets may include pre-translation language data indicating the pre-translation language and post-translation language data indicating the post-translation language.
[0085] In this embodiment, for example, the input relay unit 80 of the relay device 14 receives voice data transmitted from the input transmission unit 46. Then, the input relay unit 80 transmits the received voice data to the voice recognition unit 90 of the voice processing system 16. For example, the input relay unit 80 receives packets of voice data transmitted as a stream from the input transmission unit 46 and transmits the packets to the voice recognition unit 90 of the voice processing system 16.
[0086] In this embodiment, the speech processing system 16 may include multiple speech recognition units 90 each associated with a different language. The input relay unit 80 may then transmit the received speech data to the speech recognition unit 90 associated with the post-translation language.
[0087] In this embodiment, when the input relay unit 80 receives a packet transmitted from the input transmission unit 46, it temporarily stores the packet in the voice buffer 82. Then, the input relay unit 80 transmits the packet stored in the voice buffer 82 to the voice recognition unit 90 of the voice processing system 16. In this manner, even if a communication error occurs in communication between the voice processing system 16 and the relay device 14, it is possible to retry transmitting the packet.
[0088] In this embodiment, for example, the voice recognition unit 90 of the voice processing system 16 receives packets of voice data transmitted from the input relay unit 80 of the relay device 14 .
[0089] In this embodiment, the speech recognition unit 90 of the speech processing system 16 performs speech recognition processing on the speech represented by the received speech data, and generates a speech recognition result character string representing the speech recognition result of the speech. Here, for example, each time the speech recognition unit 90 receives a packet of speech data, the speech recognition unit 90 may sequentially perform speech recognition processing on the speech data received from the time the terminal 10 is connected to the relay device 14 until the time the packet is received, and generate a speech recognition result character string.
[0090] In this embodiment, for example, the speech recognition unit 90 of the speech processing system 16 transmits the speech recognition result character string generated by the speech recognition unit 90 to the relay device 14. Here, when the speech recognition process is executed sequentially, the generated speech recognition result character string may be transmitted to the relay device 14 every time a speech recognition result character string is generated.
[0091] In this embodiment, the character string relay unit 84 of the relay device 14 receives, for example, the above-mentioned voice recognition result character string.
[0092] Then, in response to the change of terminal 10 from the input on state to the input off state, input transmitting unit 46 transmits a confirmation request to relay device 14. Note that if there is voice data stored in audio buffer 44 when the input on state changes from the input off state, input transmitting unit 46 transmits the voice data stored in audio buffer 44 to relay device 14 and then transmits the confirmation request to relay device 14. Also, if there is no voice data stored in audio buffer 44 when the input on state changes from the input off state, input transmitting unit 46 immediately transmits the confirmation request to relay device 14. Generally, when the input on state changes from the input off state, there is often no voice data stored in audio buffer 44, and at the time the input on state changes from the input off state, almost all of the voice data has already been transmitted.
[0093] In this embodiment, when voice is input to the terminal 10 for a predetermined time (for example, 30 seconds), the reception of voice may be terminated at that timing, and a confirmation request may be transmitted.
[0094] In this embodiment, the input relay unit 80 of the relay device 14 receives a confirmation request output in response to a predetermined operation performed by the speaker (for example, an operation of releasing a finger from the translation button 10da in this example). For example, the input relay unit 80 of the relay device 14 receives a confirmation request transmitted from the input transmitting unit 46 when the speaker releases a finger from the translation button 10da.
[0095] In this embodiment, the string relay unit 84 of the relay device 14, for example, triggers reception of a confirmation request by the input relay unit 80 and controls to start translation of the speech represented by the speech data received up until the reception of the confirmation request. For example, in response to the input relay unit 80 receiving the confirmation request, the string relay unit 84 of the relay device 14 transmits to the translation unit 92 of the speech processing system 16 a speech recognition character string representing the speech recognition result of the speech represented by the speech data received from the time the terminal 10 was connected to the relay device 14 until the time the confirmation request was received.
[0096] In this embodiment, the speech processing system 16 may include multiple translation units 92 each associated with a different language. The string relay unit 84 may then transmit the speech-recognized string to the translation unit 92 associated with the translated language.
[0097] In this embodiment, for example, the translation unit 92 of the speech processing system 16 receives the speech recognition result character string transmitted by the character string relay unit 84. Then, the translation unit 92 of the speech processing system 16 performs translation processing on the received speech recognition result character string. Then, the translation unit 92 generates a translation result character string that represents the result of the translation processing.
[0098] Then, in this embodiment, the translation unit 92 transmits the translation result character string generated as described above to the relay device 14, for example.
[0099] In addition, in this embodiment, for example, the string relay unit 84 of the relay device 14 transmits a voice recognition result character string representing the voice recognition result of the voice represented by the above-mentioned voice data to both the communication unit 10c of the terminal 10 and the communication unit 12c of the client device 12. For example, in response to receiving a voice recognition result character string from the voice recognition unit 90 of the voice processing system 16, the string relay unit 84 transmits the voice recognition character string to both the terminal 10 and the client device 12.
[0100] In addition, in this embodiment, for example, the string relay unit 84 of the relay device 14 transmits a translation result string representing the translation result of the speech represented by the above-mentioned speech data to both the communication unit 10c of the terminal 10 and the communication unit 12c of the client device 12. For example, in response to receiving a translation result string from the translation unit 92 of the speech processing system 16, the string relay unit 84 transmits the translation result string to both the terminal 10 and the client device 12.
[0101] The character string receiving unit 48 of the terminal 10 receives the voice recognition result character string transmitted from the relay device 14, for example, in this embodiment.
[0102] Furthermore, the character string receiving unit 48 of the terminal 10 receives the translation result character string transmitted from the relay device 14 in this embodiment, for example.
[0103] The display control unit 50 of the terminal 10, for example, causes the display unit (for example, the touch panel 10f) of the terminal 10 to display the speech recognition result character string received by the character string receiving unit 48. The display control unit 50 also causes the display unit (for example, the touch panel 10f) of the terminal 10 to display the translation result character string received by the character string receiving unit 48.
[0104] 7, the display control unit 50 may generate a translation result image 36, which is an image in which both the speech recognition result string and the translation result string received by the string receiving unit 48 are arranged. Then, the display control unit 50 may display the translation result image 36 on the touch panel 10f.
[0105] In this embodiment, the display control unit 50 may display the character string received by the character string receiving unit 48 on the touch panel 10f in a color different from the solid-color background. This allows the user to more accurately understand the translation results and speech recognition results of the voice input by the user.
[0106] In this embodiment, the voice input receiving unit 60 of the client device 12 receives the user's voice input via, for example, the microphone 12g. Then, the voice input receiving unit 60 outputs voice data representing the input voice to the video conference client unit 68.
[0107] In this embodiment, the character string receiving unit 62 of the client device 12 receives, for example, a voice recognition result character string transmitted from the relay device 14.
[0108] Furthermore, the character string receiving unit 62 of the client device 12 receives the translation result character string transmitted from the relay device 14, for example, in this embodiment.
[0109] In this embodiment, the photographed image acquisition unit 64 acquires a photographed image, which is an image photographed by the photographing unit 12e, for example.
[0110] In this embodiment, the superimposed image generating unit 66 generates a superimposed image 32, which is an image obtained by superimposing, on the captured image, for example, a voice recognition result character string received by the character string receiving unit 62. In addition, in this embodiment, the superimposed image generating unit 66 generates a superimposed image 32, which is an image obtained by superimposing, on the captured image, for example, a translation result character string received by the character string receiving unit 62.
[0111] Here, as shown in FIG. 6, the superimposed image generating unit 66 may generate a superimposed image 32, which is an image in which both the translation result string and the voice recognition result string received by the string receiving unit 62 are superimposed on the above-mentioned captured image.
[0112] Then, the superimposed image generating unit 66 outputs the generated superimposed image 32 to the video conference client unit 68, for example, in this embodiment.
[0113] In this embodiment, the video conference client unit 68 of the client device 12 cooperates with the video conference system 18, for example, to execute various processes related to the video conference.
[0114] The video conference client unit 68 may, for example, output a superimposed image 32 in which the character string received by the character string receiving unit 62 is superimposed on the above-mentioned captured image to the video conference system 18. For example, the video conference client unit 68 may output the superimposed image 32 received from the superimposed image generating unit 66 to the video conference system 18.
[0115] Furthermore, the video conference client unit 68 may output the audio data received from the audio input receiving unit 60 to the video conference system 18, for example.
[0116] Then, the video conference client unit 68 outputs the video conference screen 30 shown in FIGS. 4 and 6, which is generated by the video conference system 18 in this embodiment, to the display control unit 72.
[0117] In addition, the video conference client unit 68 outputs to the audio output control unit 70, for example, audio data representing the audio of the speaker in the video conference, which is generated by the video conference system 18 in this embodiment.
[0118] In this embodiment, for example, the audio output control unit 70 of the client device 12 outputs the audio represented by the audio data received from the video conference client unit 68 from the speaker 12h.
[0119] In this embodiment, for example, the display control unit 72 of the client device 12 causes the display 12f to display a screen on which an image is arranged in which a character string representing the speech recognition result of the speech represented by the audio data is superimposed on an image captured by the imaging unit 12e. Here, the display control unit 72 may cause the display 12f to display a screen on which an image is arranged in which a character string representing the speech recognition result of the speech represented by the accepted audio data is superimposed on an image captured by the imaging unit 12e before accepting the confirmation request. For example, the display control unit 72 of the client device 12 causes the display 12f of the client device 12 to display a screen on which an image is arranged in which a character recognition result character string received by the character string receiving unit 62 is superimposed on the above-mentioned captured image.
[0120] In this embodiment, the display control unit 72 also causes the display 12f to display a screen on which an image is arranged in which a character string representing the translation result of the voice represented by the voice data received up until the reception of the confirmation request is superimposed on an image captured by the imaging unit 12e. For example, the display control unit 72 of the client device 12 causes the display 12f of the client device 12 to display a screen on which an image is arranged in which a translation result character string received by the character string receiving unit 62 is superimposed on the above-mentioned captured image.
[0121] Here, the display control unit 72 may cause the display 12f to display a video conference screen 30 in which a superimposed image 32 is arranged, in which both the translation result string and the voice recognition result string received by the string receiving unit 62 are superimposed on the above-mentioned captured image, as shown in FIG. 6.
[0122] The display control unit 72 may also cause the display 12f to display a screen generated by the video conference system 18. For example, the display control unit 72 may cause the display 12f to display the video conference screen 30 received from the video conference client unit 68.
[0123] Here, an example of the flow of the audio data relay process executed by the relay device 14 will be described with reference to the flow diagram shown in FIG.
[0124] In this processing example, the input relay unit 80 monitors the reception of a communication start request transmitted from the input transmitting unit 46 of the terminal 10 (S101).
[0125] When the input relay unit 80 receives a communication start request from the input transmission unit 46 of the terminal 10, the input relay unit 80 establishes communication between the relay device 14 and the terminal 10 (S102).
[0126] Then, the input relay unit 80 monitors the reception of a packet of audio data (S103). When the input relay unit 80 receives a packet of audio data, the input relay unit 80 stores the received packet in the audio buffer 82 (S104).
[0127] Then, the input relay unit 80 transmits the packets stored in the voice buffer 82 in the process shown in S104 to the voice recognition unit 90 of the voice processing system 16, and returns to the process shown in S103.
[0128] The processes shown in S103 to S105 are continued until the process shown in S207, which will be described later, is executed.
[0129] Next, an example of the flow of the character string relay process executed by the relay device 14 will be described with reference to the flowchart shown in FIG.
[0130] In this processing example, the string relay unit 84 monitors the reception of a speech recognition result character string transmitted from the speech recognition unit 90 of the speech processing system 16 (S201). When the string relay unit 84 receives the speech recognition result character string, it transmits the received speech recognition result character string to the string receiving unit 62 of the client device 12 (S202).
[0131] Then, the character string relay unit 84 checks whether the input relay unit 80 has received a confirmation request (S203). If the reception of the confirmation request is not confirmed (S203: N), the process returns to S201. If the reception of the confirmation request is confirmed (S203: Y), the character string relay unit 84 transmits a speech recognition result character string representing the speech recognition result of the speech represented by the speech data received up until the reception of the confirmation request to the translation unit 92 of the speech processing system 16 (S204).
[0132] Then, the character string relay unit 84 receives the translation result character string obtained by translating the speech recognition result character string transmitted in the process shown in S203, transmitted from the translation unit 92 of the speech processing system 16 (S205).
[0133] Then, the string relay unit 84 transmits the confirmation flag, the translation result string received in the process shown in S205, and the speech recognition result string representing the speech recognition result of the speech represented by the speech data received up until the reception of the confirmation request to the string receiving unit 62 of the client device 12 (S206).
[0134] Then, the character string relay unit 84 disconnects the communication between the relay device 14 and the terminal 10 (S207), and the process shown in this processing example is terminated. By executing the process shown in S207, the processes shown in S103 to S105 are also terminated.
[0135] Next, an example of the flow of the process of generating the superimposed image 32 executed by the client device 12 will be described with reference to the flowchart shown in Fig. 11. In this process example, the following processes shown in S301 to S305 are repeatedly executed at the frame rate at which the image capturing unit 12e captures images. In this embodiment, for example, the processes shown in S301 to S305 may be executed at intervals of 1 / 30 seconds. Note that the execution intervals of the processes shown in S301 to S305 may be longer (or shorter) than 1 / 30 seconds. Furthermore, the execution intervals may be adjustable by the user.
[0136] First, the photographed image acquisition unit 64 acquires a photographed image in the frame (S301).
[0137] Then, the superimposed image generating unit 66 checks whether or not the character string receiving unit 62 has received a confirmation flag since the process shown in S202 was last executed (S302).
[0138] If it is confirmed that the confirmation flag has not been received (S302: N), the superimposed image generating unit 66 generates a superimposed image 32 in which the latest voice recognition result character string received by the character string receiving unit 62 is superimposed on the captured image acquired by the process shown in S301 (S303).
[0139] If it is confirmed that the confirmation flag has been received (S302: Y), the superimposed image generation unit 66 generates a superimposed image 32 in which the latest speech recognition result string and the latest translation result string received by the string receiving unit 62 are superimposed on the captured image acquired in the process shown in S301 (S304).
[0140] Then, the superimposed image generating unit 66 outputs the superimposed image 32 generated in the process shown in S303 or S304 to the video conference client unit 68 (S305), and the process returns to the process shown in S301.
[0141] In this embodiment, the displayable area within the captured image may be set by the user. For example, the displayable area may be selectable from among the upper section, the lower section, the entire image, etc. Also, the displayable area for the speech recognition result character string and the displayable area for the translation result character string may be set separately. For example, FIGS. 4 and 6 show an example of the video conference screen 30 when the lower section is set as the displayable area for the speech recognition result character string. Also, FIG. 6 shows an example of the video conference screen 30 when the entire image is set as the displayable area for the translation result character string.
[0142] In this embodiment, a line break may be inserted at a predetermined number of characters for a character string in a language such as Japanese that does not have spaces separating words, or a word wrap process may be performed at a predetermined number of characters for a character string in a language such as English that has spaces separating words.
[0143] To improve readability, the character size of the translation result character string may be larger than the character size of the speech recognition result character string.
[0144] In this embodiment, it is not necessary to superimpose both the translation result character string and the voice recognition result character string on the captured image. For example, when the translation result character string is superimposed on the captured image, the voice recognition result character string may not be superimposed on the captured image.
[0145] In this embodiment, the character size of the speech recognition result character string may be fixed, and the character size of the translation result character string may be variable.
[0146] In this case, the maximum character size of the characters included in the translation result character string may be a size calculated by multiplying the height of the screen by a predetermined ratio, and the character size of the translation result character string may be made smaller as the number of characters in one line increases.
[0147] The character size of the speech recognition result character string may be variable, or the character size of the translation result character string may be fixed.
[0148] In this embodiment, the number of displayable characters corresponding to the size of the displayable area may be predetermined. When a speech recognition result character string with more characters than the displayable number of characters is superimposed on a captured image, the speech recognition result character string may be reduced to fit within the height of the displayable area before being superimposed on the captured image. When a translation result character string with more characters than the displayable number of characters is superimposed on a captured image, the translation result character string may be reduced to fit within the height of the displayable area before being superimposed on the captured image.
[0149] Furthermore, in this embodiment, the string relay unit 84 may perform control so as to start translation of the speech represented by the speech data received up to now, triggered by a predetermined time (e.g., 1.5 seconds) being interrupted in the reception of speech data packets by the input relay unit 80. For example, the string relay unit 84 of the relay device 14 may transmit to the translation unit 92 of the speech processing system 16 a speech recognition character string representing the speech recognition result of the speech represented by the speech data received up to now since the terminal 10 was connected to the relay device 14, in response to a predetermined time (e.g., 1.5 seconds) being interrupted in the reception of speech data packets by the input relay unit 80.
[0150] A list (log) of the speech recognition result character strings and the translation result character strings may be displayed on a screen (for example, a browser) separate from the video conference screen 30. This log may be stored in a storage medium such as the storage unit 12b of the client device 12. A translation result character string obtained by translating the speech recognition result character string into a language different from the above-mentioned post-translation language may be displayed on the browser.
[0151] Furthermore, the functions of the terminal 10 of the video conference translation system 1 may be implemented in the client device 12.
[0152] 12, the client device 12 may have a function to display a translation button 94 on the display 12f. The translation button 94 may be displayed on the display 12f in addition to the video conference screen 30. The client device 12 may be configured to switch between the input-on state and the input-off state described above each time the speaker performs a predetermined operation, such as a click operation, on the translation button 94. A speech recognition result character string and a translation result character string for the speech input while the input is in the input-on state may be displayed on the video conference screen 30.
[0153] Fig. 13 is a functional block diagram showing an example of functions implemented in a client device 12 according to a modified example of the embodiment described with reference to Figs. 1 to 11. The client device 12 according to this embodiment does not need to implement all of the functions shown in Fig. 13, and may also implement functions other than the functions shown in Fig. 13.
[0154] As shown in FIG. 13 , the client device 12 according to this modification functionally includes, for example, an operation input reception unit 40, an audio buffer 44, an input transmission unit 46, an audio input reception unit 60, a character string reception unit 62, a captured image acquisition unit 64, a superimposed image generation unit 66, a video conference client unit 68, an audio output control unit 70, and a display control unit 72. The operation input reception unit 40 is implemented mainly using the processor 12a and the operation unit 12d. The audio buffer 44 is implemented mainly using the storage unit 12b. The input transmission unit 46 and the character string reception unit 62 are implemented mainly using the communication unit 12c. The audio input reception unit 60 is implemented mainly using the processor 12a and the microphone 12g. The captured image acquisition unit 64 is implemented mainly using the processor 12a and the image capture unit 12e. The superimposed image generation unit 66 is implemented mainly using the processor 12a. The video conference client unit 68 is mainly implemented by the processor 12a and the communication unit 12c. The audio output control unit 70 is mainly implemented by the processor 12a and the speaker 12h. The display control unit 72 is mainly implemented by the processor 12a and the display 12f.
[0155] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 12a, which is installed in the client device 12, which is a computer. This program is supplied to the client device 12 via a computer-readable information storage medium, such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.
[0156] In this embodiment, the operation input accepting unit 40 causes the display 12f to display, for example, a translation button 94. Then, in this embodiment, the operation input accepting unit 40 accepts an operation input such as clicking the translation button 94.
[0157] In this embodiment, the voice input receiving unit 60 receives the user's voice input via, for example, the microphone 12g. Then, the voice input receiving unit 60 outputs voice data representing the input voice to the video conference client unit 68.
[0158] In the present embodiment, for example, in response to the client device 12 changing from an input-off state to an input-on state, the input transmitting unit 46 transmits a communication start request to the relay device 14. Then, audio data representing audio input via the microphone 12g from the time the client device 12 changes from an input-off state to an input-on state until communication between the relay device 14 and the terminal 10 is established is not only output to the video conference client unit 68 but also accumulated in the audio buffer 44.
[0159] Furthermore, in response to the client device 12 changing from an input-on state to an input-off state, the input transmitting unit 46 transmits a confirmation request to the relay device 14.
[0160] Other functions of the audio buffer 44 and the input transmitting unit 46 are similar to those described above with reference to Fig. 8A, and therefore description thereof will be omitted. Also, functions of the character string receiving unit 62, the captured image acquiring unit 64, the superimposed image generating unit 66, the video conference client unit 68, the audio output control unit 70, and the display control unit 72 are similar to those described above with reference to Fig. 8B, and therefore description thereof will be omitted. Note that in this modified example, the relay device 14 does not transmit a character string to the terminal 10.
[0161] 12 and 13 , the input relay unit 80 may receive, from the client device 12, voice data representing the voice input by the speaker to the client device 12. In addition, the input relay unit 80 may receive a confirmation request transmitted from the client device 12 in response to a predetermined operation performed by the speaker on the client device 12.
[0162] The display control unit 72 may then cause the display 12f provided on the client device 12 to display a screen in which an image is placed in which a character string representing the translation result of the voice represented by the voice data received up until the time of receiving the confirmation request is superimposed on an image captured by the photographing unit 12e.
[0163] In this embodiment, multiple languages may be set as post-translation languages. The string relay unit 84 may then control the start of translation of the speech represented by the speech data received before the confirmation request is received into the multiple set languages. In this case, for example, the string relay unit 84 may transmit the speech-recognized character string to multiple translation units 92 that are respectively associated with the multiple post-translation languages.
[0164] Then, the display control unit 72 may cause the display 12f to display a screen in which an image in which the translation result character string for each of the set languages is superimposed on the captured image is arranged.
[0165] For example, a translation result string obtained by translating the voice recognition result string into English may be displayed in the lower part of the captured image, and a translation result string obtained by translating the voice recognition result string into Chinese may be displayed in the upper part of the captured image.
[0166] Then, once it has been confirmed that the translation result character strings for all post-translation languages have been displayed, these translation result character strings may be erased from the screen.
[0167] The present invention is not limited to the above-described embodiment.
[0168] For example, the division of roles among the terminal 10, the client device 12, the relay device 14, the speech processing system 16, and the video conference system 18 is not limited to those described above. For example, the speech recognition result character string may be translated in the speech processing system 16 without passing through the relay device 14.
[0169] For example, the client device 12 may receive, from the relay device 14, audio data transmitted from the terminal 10 to the relay device 14. Then, the client device 12 may output to the video conference system 18 the audio data received from the relay device 14, instead of the audio data representing the audio input via the microphone 12g.
[0170] Furthermore, the specific character strings and numerical values described above and the specific character strings and numerical values in the drawings are examples, and the present invention is not limited to these character strings and numerical values. [Explanation of symbols]
[0171] 1 Video conference translation system, 10 terminal, 10a processor, 10b memory unit, 10c communication unit, 10d operation unit, 10da translation button, 10db power button, 10dc volume control unit, 10e photographing unit, 10f touch panel, 10g microphone, 10h speaker, 12 client device, 12a processor, 12b memory unit, 12c communication unit, 12d operation unit, 12e photographing unit, 12f display, 12g microphone, 12h speaker, 14 relay device, 14a processor, 14b memory unit, 14c communication unit, 16 voice processing system, 16a processor, 16b memory unit, 16c communication unit, 18 video conference system, 20 computer network, 30 video conference screen, 32 superimposed image, 34 speech recognition result image, 36 translation result image, 40 Operation input reception unit, 42 voice input reception unit, 44 voice buffer, 46 input transmission unit, 48 character string reception unit, 50 display control unit, 60 voice input reception unit, 62 character string reception unit, 64 captured image acquisition unit, 66 superimposed image generation unit, 68 video conference client unit, 70 voice output control unit, 72 display control unit, 80 input relay unit, 82 voice buffer, 84 character string relay unit, 90 voice recognition unit, 92 translation unit, 94 translation button.
Claims
1. a voice data receiving means for receiving voice data representing a voice input by a speaker; a confirmation request receiving means for receiving a confirmation request output in response to a predetermined operation performed by the speaker; a translation control means for controlling the start of translation of the speech represented by the speech data received up until the reception of the confirmation request, using the reception of the confirmation request as a trigger; a translation result display control means for displaying on a display unit a screen on which an image is arranged in which a character string representing the translation result of the speech represented by the speech data received up until the reception of the confirmation request is superimposed on an image taken by a photographing unit; and A display control system comprising:
2. a voice recognition result display control means for controlling the display unit to display a screen on which an image is arranged in which a character string representing a voice recognition result of the voice represented by the voice data is superimposed on an image captured by the image capturing unit, the voice recognition result display control means, before the acceptance of the confirmation request, causes the display unit to display a screen on which an image is arranged in which a character string representing a voice recognition result of the voice represented by the accepted voice data is superimposed on an image taken by the image taking unit.
2. The display control system according to claim 1.
3. the translation result display control means causes the display unit to display a screen on which an image is arranged in which both a character string representing the speech recognition result of the speech represented by the speech data received up to the time of receiving the confirmation request and a character string representing the translation result of the speech represented by the speech data received up to the time of receiving the confirmation request are superimposed on an image captured by the imaging unit.
3. The display control system according to claim 1 or 2.
4. an image output unit that outputs an image, in which a character string is superimposed on the image captured by the image capturing unit, to a video conference system; the translation result display control means causes the screen generated by the video conference system to be displayed on the display unit; 4. The display control system according to claim 1, wherein the display control system is a display control system for displaying a display image.
5. the voice data receiving means receives, from the terminal, the voice data representing the voice input to the terminal by the speaker; the confirmation request receiving means receives the confirmation request transmitted from the terminal in response to a predetermined operation performed by the speaker on the terminal; the translation result display control means displays, on a display unit provided in the terminal, a character string representing a translation result of the speech represented by the speech data received up until the reception of the confirmation request; the translation result display control means causes a display unit provided in the client device to display a screen on which an image is arranged in which a character string representing the translation result of the speech represented by the speech data received up until the reception of the confirmation request is superimposed on an image captured by the image capturing unit; 5. The display control system according to claim 1, wherein the display control system is a display control system for displaying a display image.
6. the voice data receiving means receives, from the client device, the voice data representing the voice input to the client device by the speaker; the confirmation request receiving means receives the confirmation request transmitted from the client device in response to a predetermined operation performed by the speaker on the client device; the translation result display control means causes the display unit of the client device to display a screen on which an image is arranged in which a character string representing a translation result of the speech represented by the speech data received up until the reception of the confirmation request is superimposed on an image captured by the image capturing unit, 5. The display control system according to claim 1, wherein the display control system is a display control system for displaying a display image.
7. the translation control means controls to start translation of the speech represented by the speech data received up until the reception of the confirmation request into a plurality of languages; the translation result display control means causes the display unit to display a screen on which an image is arranged in which a character string representing a translation result of the speech represented by the speech data for each of the plurality of languages is superimposed on an image captured by the imaging unit; 7. The display control system according to claim 1, wherein the display control system is a display control system for displaying a display image.
8. receiving speech data representing speech input by a speaker; receiving a confirmation request output in response to a predetermined operation performed by the speaker; a step of controlling the translation of the speech represented by the speech data received up until the reception of the confirmation request to be started, using the reception of the confirmation request as a trigger; displaying on a display unit a screen on which an image is arranged in which a character string representing a translation result of the voice represented by the voice data received up until the reception of the confirmation request is superimposed on an image captured by a photographing unit; A display control method comprising:
9. receiving speech data representing speech input by a speaker; a step of receiving a confirmation request output in response to a predetermined operation performed by the speaker; a step of controlling the start of translation of the speech represented by the speech data received up until the reception of the confirmation request, triggered by the reception of the confirmation request; a step of displaying on a display unit a screen on which an image is arranged in which a character string representing the translation result of the voice represented by the voice data received up until the reception of the confirmation request is superimposed on an image captured by a photographing unit; A program characterized by causing a computer to execute the above.
Citation Information
Patent Citations
Voice conversation translation device, voice conversation translation method and voice conversation translation program
JP2007080097A
Conference system, information processor, conference supporting method, information processing method, and computer program
JP2011182125A
Method and system for adding translation in videoconference
JP2011209731A
Translation system, translation processor, and translation processing program
JP2015153408A
Voice translation device, method, and program
JP2016062357A