Speech processing device, speech processing method, and program
Patent Information
- Application Number
- JP2024105113
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2044-06-28
Smart Images

Figure 0007914165000001 
Figure 0007914165000002 
Figure 0007914165000003
Abstract
Description
Technical Field
[0001] The present invention relates to a voice processing device, a voice processing method, and a program.
Background Art
[0002] In recent years, with the development of AI (Artificial Intelligence), the accuracy of speech recognition using AI has improved, but recognition results may still contain misrecognition.
[0003] In relation to this, the following Patent Document 1 discloses a technique for correcting recognition results obtained by speech recognition. This technique is used, for example, when correcting captions generated based on speech recognition results in video data including audio data.
Prior Art Literature
Patent Literature
[0004]
Patent Document 1
Summary of the Invention
Problem to be Solved by the Invention
[0005] However, with the technique described in Patent Document 1 above, even if misrecognition occurring in speech recognition for a certain piece of audio data is corrected, the same misrecognition may occur in speech recognition for another piece of audio data. In this case, a checker needs to perform the same correction every time the same misrecognition occurs, resulting in a large workload for the misrecognition correction work.
[0006] In view of the above problems, an object of the present invention is to provide a voice processing device, a voice processing method, and a program capable of reducing the load imposed on correction work of misrecognition in speech recognition.
Means for Solving the Problem
[0007] To solve the above-mentioned problems, an audio processing device according to one aspect of the present invention provides a first language Speaker The audio data representing the content of the utterance is processed by speech recognition in the first language. Speaker A speech conversion unit that converts the spoken content into text data, and a unit that divides the text data into multiple texts. The terminal that the checker uses for monitoring Display and modify the text data in units of the divided text. From the aforementioned checker A monitoring processing unit that receives the data, a translation unit that translates the corrected text data into multiple second languages different from the first language, and obtains text data showing the translation result with the corrections reflected for each of the multiple second languages, A subtitle processing unit that displays text data showing the corrected translation result as subtitles on a display device used by the audience to check the subtitles. The monitoring processing unit is an audio processing device that, when displaying second text data obtained by speech recognition performed after the modification of the first text data has been carried out, converts the second text to the modified first text and displays it if the second text contains the first text that has been modified in the first text data.
[0008] A speech processing method according to one aspect of the present invention is a method of processing in a first language Speaker The audio data representing the content of the utterance is processed by speech recognition in the first language. Speaker A speech conversion process that converts the spoken content into text data, and a process that divides the text data into multiple texts. The terminal that the checker uses for monitoring Display and modify the text data in units of the divided text. From the aforementioned checker A monitoring process that receives the data, a translation process that translates the corrected text data into multiple second languages different from the first language, and obtains text data showing the translation result with the corrections reflected for each of the multiple second languages, A subtitle processing process that displays text data showing the corrected translation result as subtitles on a display device used by the audience to view the subtitles.The monitoring process is a computer-based speech processing method that, when displaying second text data obtained by speech recognition performed after the modification of the first text data, if the second text data contains a second text corresponding to the first text modified in the first text data, converts the second text to the modified first text and displays it.
[0009] A program according to one aspect of the present invention enables a computer to use a first language Speaker The audio data representing the content of the utterance is processed by speech recognition in the first language. Speaker A speech conversion means that converts the content of the utterance into text data, and divides the text data into multiple texts. The terminal that the checker uses for monitoring Display and modify the text data in units of the divided text. From the aforementioned checker A monitoring processing means for receiving data, and a translation means for translating the corrected text data into multiple second languages different from the first language, and obtaining text data showing the translation result with the corrections reflected for each of the multiple second languages, Subtitle processing means that displays text data showing the corrected translation result as subtitles on a display device used by the audience to check the subtitles. The monitoring processing means functions as a program that, when displaying second text data obtained by speech recognition performed after the modification of the first text data has been carried out, converts the second text to the modified first text and displays it if the second text contains the first text that has been modified in the first text data. [Effects of the Invention]
[0010] According to the present invention, the workload associated with correcting misrecognitions in speech recognition can be reduced. [Brief explanation of the drawing]
[0011] [Figure 1] This figure shows an example of the configuration of a subtitle display system according to the first embodiment. [Figure 2]It is a block diagram showing an example of the functional configuration of the speech processing apparatus according to the first embodiment. [Figure 3] It is a diagram showing an example of the monitoring screen according to the first embodiment. [Figure 4] It is a diagram showing an example of the correction operation procedure according to the first embodiment. [Figure 5] It is a sequence diagram showing an example of the processing flow in the caption display system according to the first embodiment. [Figure 6] It is a diagram showing an example of the configuration of the caption display system according to the second embodiment. [Figure 7] It is a block diagram showing an example of the functional configuration of the speech processing apparatus according to the second embodiment. [Figure 8] It is a diagram showing an example of the monitoring screen according to the second embodiment. [Figure 9] It is a diagram showing an example of the correction operation procedure according to the second embodiment. [Figure 10] It is a diagram showing an example of changing conversion priority according to the second embodiment. [Figure 11] It is a sequence diagram showing an example of the processing flow in the caption display system according to the second embodiment. [Figure 12] It is a sequence diagram showing an example of the flow of display preparation processing according to the second embodiment. MODE FOR CARRYING OUT THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0013] <<1. First Embodiment>> The first embodiment will be described with reference to FIGS. 1 to 5. A caption display system will be described below. The caption display system is a system for translating a user's utterance content into a plurality of languages different from the language used by the user (multilingual translation) and displaying captions in each language (multilingual captions). In the following, we will describe the first embodiment using as an example the distribution of multilingual translated subtitles of the content spoken by a user (speaker) at a lecture to users (attendees) attending the lecture. In the first embodiment, the language spoken by the speaker (first language) is any one language (for example, any language other than Japanese), and the language of the multilingual translated subtitles (second language) is any multiple languages (for example, multiple languages including Japanese).
[0014] <1-1. Configuration of the subtitle display system> Referring to Figure 1, the configuration of the subtitle display system according to the first embodiment will be described. Figure 1 is a diagram showing an example of the configuration of the subtitle display system according to the first embodiment.
[0015] As shown in Figure 1, the subtitle display system 1 comprises a sound collection device 10, an audio processing device 20, a simultaneous interpretation engine 21, a machine translation engine 22, a monitoring terminal 30, and a display device 40. Each device and terminal is connected via wired, wireless, or network connection, enabling the transmission and reception of various types of information. Examples of networks include LAN (Local Area Network), WAN (Wide Area Network), telephone networks (mobile phone networks, fixed-line telephone networks, etc.), regional IP (Internet Protocol) networks, and the Internet.
[0016] (1) Sound collection device 10 The sound collection device 10 is a device that collects sound produced when the speaker speaks. The sound collection device 10 is, for example, a microphone. The sound collection device 10 is connected to the audio processing device 20 via a wired or wireless connection for communication purposes. When the sound collection device 10 collects the speaker's voice, it converts the collected voice into data and transmits the converted audio data to the audio processing device 20.
[0017] (2) Audio processing device 20 The audio processing device 20 is a device that processes the results of simultaneous interpretation to display them as subtitles. The audio processing device 20 is implemented by devices such as one or more servers (e.g., cloud servers) or PCs (personal computers). Various processes are executed on these devices by a program that enables them to function as the audio processing device 20. The audio processing device 20 performs various processes based on the audio data received from the sound collection device 10. For example, the audio processing device 20 performs processes such as speech conversion, monitoring, machine translation, and subtitle display.
[0018] The speech conversion process is the process of converting speech data into text data. In the speech conversion process, the speech processing device 20 performs speech recognition and machine translation on the speech data received from the sound collection device 10 using the functions of the simultaneous interpretation engine 21 described later, and obtains text data showing the translation result. In the speech conversion process, one text data expressed in a first language is machine translated into multiple text data expressed in multiple second languages different from the first language.
[0019] The monitoring process is performed so that the checker can monitor the results of the speech conversion process on the audio data (speech conversion results). During the monitoring process, the audio processing device 20 displays a monitoring screen on the monitoring terminal 30. The monitoring screen is a screen that accepts the selection of whether or not to display the text data converted from the audio data, and the modification of the text data.
[0020] The checker can monitor the speech conversion results by checking the speech conversion results displayed on the monitoring screen at the monitoring terminal 30. The checker determines whether the text data converted from the speech data during the speech conversion process can be displayed as subtitles (displayability) through monitoring. The criteria for determining whether or not to display include, for example, whether or not there are misrecognitions or misconversions in the speech conversion results, or whether or not the meaning of the text data is understandable. If there are misrecognitions or misconversions in the speech conversion results, or if the meaning of the text data is not understandable, the checker determines that the text data cannot be displayed as subtitles. On the other hand, if there are no misrecognitions or misconversions in the speech conversion results, and the meaning of the text data is understandable, the checker determines that the text data can be displayed as subtitles.
[0021] The checker, based on the judgment result, selects whether to display the data on the monitoring screen and makes corrections as necessary. If it determines that the data can be displayed, the checker selects "displayable" on the monitoring screen. On the other hand, if it determines that the data cannot be displayed, the checker selects "displayable" on the monitoring screen and corrects the text data so that it can be displayed. Once the checker has completed its corrections, the display status of the corrected text data is automatically changed to "displayable" by the speech processing device 20.
[0022] Furthermore, the monitoring screen may be equipped with an automatic display / display selection function. This automatic selection function automatically selects whether to display or hide text data after a predetermined time has elapsed since the data was displayed on the monitoring screen. The predetermined time can be set by any user (e.g., checker or administrator) on the monitoring screen (e.g., 3 seconds). In addition, any user on the monitoring screen can set whether to automatically select display or hide.
[0023] Furthermore, the voice conversion results displayed on the monitoring screen consist of text data in only one language. The checker only needs to monitor one of the text data in the second language out of the multiple text data in the second language converted by the voice conversion process on the monitoring screen. Note that the text data in the second language to be displayed on the monitoring screen can be appropriately selected, for example, depending on the languages that the checker can handle.
[0024] Machine translation processing is the process of translating text data. In machine translation processing, the speech processing device 20 machine-translates the text data that has been determined to be undisplayable and corrected in the monitoring process using the functions of the machine translation engine 22 described later, and obtains text data showing the translation result. In machine translation processing, one text data expressed in a first language is machine-translated into multiple text data expressed in multiple second languages different from the first language. Therefore, through machine translation processing, the speech processing device 20 can obtain text data that reflects the corrections in multiple other languages from one text data that has been corrected in one language.
[0025] The subtitle display process is the process of displaying subtitles. In the subtitle display process, the audio processing device 20 displays text data representing the translation result in the machine translation process as subtitles on the display device 40. The audio processing device 20 may display only subtitles in one specified second language on the display device 40, or it may display subtitles in multiple second languages on the display device 40.
[0026] (3) Simultaneous interpretation engine 21 The simultaneous interpretation engine 21 is an engine (program) that simultaneously interprets a first language into a second language. The simultaneous interpretation engine 21 converts speech data of the first language input from the speech processing device 20 into text data of the first language using speech recognition, and then machine-translates (converts) the text data of the first language into text data of the second language. The simultaneous interpretation engine 21 machine-translates one first language into multiple different second languages. That is, the simultaneous interpretation engine 21 generates multiple text data of second languages from one text data of the first language. The functions of the simultaneous interpretation engine 21 may be provided by a device or terminal different from the speech processing device 20, or they may be provided by the speech processing device 20.
[0027] (4) Machine translation engine 22 The machine translation engine 22 is an engine (program) that machine translates a first language into a second language. The machine translation engine 22 machine translates (converts) text data in the first language input from the speech processing device 20 into text data in the second language. The machine translation engine 22 machine translates one first language into multiple different second languages. That is, the machine translation engine 22 generates multiple text data in the second language from one text data in the first language. The functions of the machine translation engine 22 may be provided by a device or terminal different from the speech processing device 20, or they may be provided by the speech processing device 20.
[0028] (5) Monitoring terminal 30 The monitoring terminal 30 is a terminal used by the checker for monitoring purposes. The monitoring terminal 30 is, for example, a PC, smartphone, or tablet. Various processes are executed on this terminal by a program that enables it to function as a monitoring terminal 30. The monitoring terminal 30 displays a monitoring screen based on screen information received from the voice processing device 20 and accepts various operations related to monitoring by the checker.
[0029] The monitoring terminal 30 displays a monitoring screen, for example, through an application for using the monitoring function (hereinafter also referred to as the "monitoring app"). The checker can monitor the simultaneous interpretation results by operating the monitoring screen displayed on the monitoring terminal 30 via the monitoring app. The functionality of the monitoring application may be provided by installing the monitoring application on each device (i.e., a native application), or it may be provided by a web system (i.e., a web application). In the case of a web application, the monitoring application is managed on a server, and its functionality is provided via a web browser.
[0030] (6) Display device 40 The display device 40 is a device that displays subtitles. The display device 40 may be a display device such as a screen 41, or a device having a display such as a smartphone 42. The display device 40 is communicatively connected to the audio processing device 20 and displays subtitles based on screen information received from the audio processing device 20.
[0031] The display device 40 displays a subtitle display screen, for example, through an application for using the subtitle display function (hereinafter also referred to as the "subtitle display app"). By operating the subtitle display screen displayed on the display device 40 using the subtitle display app, the audience can check the subtitles showing the simultaneous interpretation results. The functionality of the subtitle display application may be provided by installing the application on each device (i.e., a native application), or it may be provided by a web system (i.e., a web application). In the case of a web application, the subtitle display application is managed on a server, and its functionality is provided via a web browser.
[0032] <1-2. Functional Configuration of the Audio Processing Device> The configuration of the subtitle display system 1 according to the first embodiment has been described above. Next, the functional configuration of the audio processing device 20 according to the first embodiment will be described with reference to Figures 2 to 4. Figure 2 is a block diagram showing an example of the functional configuration of the audio processing device 20 according to the first embodiment. As shown in Figure 2, the voice processing device 20 includes a communication unit 210, a storage unit 220, a first control unit 230, and a second control unit 240.
[0033] (1) Communications Section 210 The communication unit 210 has the function of sending and receiving various types of information. The communication unit 210 is communicatively connected to the sound collection device 10, the simultaneous interpretation engine 21, the machine translation engine 22, the monitoring terminal 30, and the display device 40, and sends and receives various types of information.
[0034] (2) Storage section 220 The storage unit 220 has the function of storing various types of information. The storage unit 220 is composed of storage media provided as hardware by the audio processing device 20, such as an HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, EEPROM (Electrically Erasable Programmable Read Only Memory), RAM (Random Access read / write Memory), ROM (Read Only Memory), or any combination of these storage media.
[0035] The memory unit 220 stores, for example, conversion candidate information. This conversion candidate information indicates conversion candidates for text data. Based on the correction history performed by the checker, the conversion candidate information accumulates text that can be converted. Conversion candidates are accumulated, for example, on a per-lecture basis. These candidates may be deleted at the end of a lecture, or they may remain accumulated. Conversion candidates that remain accumulated may be used for other lectures.
[0036] (3) First control unit 230 The first control unit 230 has the function of controlling processing related to simultaneous interpretation. The first control unit 230 is realized, for example, by causing the CPU (Central Processing Unit) or GPU (Graphics Processing Unit) provided as hardware in the speech processing device 20 to execute a program. As shown in Figure 2, the first control unit 230 includes an audio data acquisition unit 231 and an audio conversion unit 232.
[0037] (3-1) Audio data acquisition unit 231 The voice data acquisition unit 231 has the function of acquiring voice data. The voice data acquisition unit 231 acquires voice data received by the communication unit 210 from the sound collection device 10 and inputs it to the voice conversion unit 232.
[0038] (3-2) Voice conversion unit 232 The speech conversion unit 232 has the function of performing speech conversion processing. The speech conversion unit 232 performs speech conversion processing using the simultaneous interpretation engine 21. In the speech conversion processing, the speech conversion unit 232 inputs speech data representing the content of the speaker's utterance in a first language, acquired by the speech data acquisition unit 231, to the simultaneous interpretation engine 21, and acquires text data output from the simultaneous interpretation engine 21 as the translation result. In this way, the speech conversion unit 232 can convert speech data representing the content of the speaker's utterance in a first language into text data representing the translation results into multiple second languages different from the first language, through simultaneous interpretation. The speech conversion unit 232 inputs multiple text data in a second language obtained through the speech conversion process to the monitoring processing unit 241.
[0039] (4) Second control unit 240 The second control unit 240 has the function of controlling the processing related to monitoring and subtitle display. The second control unit 240 is implemented, for example, by causing the CPU or GPU provided as hardware in the audio processing device 20 to execute a program. As shown in Figure 2, the second control unit 240 includes a monitoring processing unit 241, a machine translation unit 242, and a subtitle processing unit 243.
[0040] (4-1) Monitoring Processing Unit 241 The monitoring processing unit 241 has the function of performing monitoring processing. During monitoring processing, the monitoring processing unit 241 transmits screen information from the communication unit 210 to the monitoring terminal 30 and displays a monitoring screen on the monitoring terminal 30. The monitoring processing unit 241 receives selection operations from the checker regarding whether or not to display text data, and correction operations for text data, via the monitoring screen displayed on the monitoring terminal 30.
[0041] The monitoring processing unit 241 displays the text data of one language that has been pre-designated as the target of monitoring from among the multiple second language text data input from the speech conversion unit 232 on the monitoring screen. For this reason, the monitoring processing unit 241 accepts corrections to the text data for the one designated second language among the multiple second languages. The one language designated in advance for monitoring is, for example, the language preferred by the checker. In the first embodiment, as an example, let's assume the checker is Japanese and the language preferred by the checker is Japanese.
[0042] The monitoring processing unit 241 displays a UI (User Interface) on the monitoring screen to accept an operation to select whether or not to display the text data displayed on the monitoring screen on the display device 40. This UI may be a button, for example, but may also be a checkbox, a dropdown menu, etc. When the monitoring processing unit 241 receives a display / display selection operation, it controls the display of subtitles for the text data that is the target of the operation. Under the control of the monitoring processing unit 241, subtitles for text data for which "display / display" is selected on the monitoring screen are not displayed on the display device 40, and subtitles for text data for which "display / display" is selected on the monitoring screen are displayed on the display device 40.
[0043] The monitoring processing unit 241 displays a UI for displaying text data on the monitoring screen. This UI is, for example, a text field. The monitoring processing unit 241 may display the text data in the text field on the monitoring screen and accept the option to enable or disable the display through operations on the text field. For example, the monitoring processing unit 241 may switch the display status to disabled when the inside of the text field is selected, and switch the display status to enabled when the outside of the text field is selected after the inside of the text field has been selected.
[0044] Furthermore, the monitoring processing unit 241 accepts modification operations on the text data through operations performed on the text field where the text data is displayed. For example, when the inside of a text field is selected by the checker, the monitoring processing unit 241 accepts modifications to the text data displayed in the selected text field. The checker can modify the text data, for example, by manual input.
[0045] The monitoring processing unit 241 may display conversion candidates near the location to be modified in the text data on the monitoring screen when a location to be modified is selected. The monitoring processing unit 241 then inserts the text selected by the checker from the conversion candidates into the location to be modified. In this way, the checker can modify text data not only by manual input but also by selecting from conversion candidates.
[0046] The monitoring processing unit 241 compares the text data before and after correction and adds any text that is not among the conversion candidates to the new conversion candidates. In this case, the monitoring processing unit 241 adds the new conversion candidates to the conversion candidate information stored in the storage unit 220. As a result, the conversion candidate information accumulates text that can be converted based on the correction history performed by the checker.
[0047] If the text data displayed in the text field does not require modification and the option to display is selected, the monitoring processing unit 241 inputs multiple text data in a second language obtained by the speech conversion processing by the speech conversion unit 232 to the subtitle processing unit 243. On the other hand, if the text data displayed in the text field requires modification and the option to display is selected as unavailable, the monitoring processing unit 241 inputs the modified text data to the machine translation unit 242.
[0048] Furthermore, if the display device 40 selects "not displayable" for text data for which subtitles are already displayed, the monitoring processing unit 241 hides the subtitles for the selected text data and makes it possible to accept correction operations.
[0049] Now, with reference to Figure 3, the monitoring screen according to the first embodiment will be described. Figure 3 is a diagram showing an example of the monitoring screen according to the first embodiment.
[0050] The buttons B1, B2, and pull-down PD on monitoring screen G1 shown in Figure 3 are UI elements for configuring the automatic selection of whether to display or not. Button B1 is used to set the system to automatically select "possible" (○) as the display option. Button B2 is used to set the system to automatically select "not possible" (×) as the display option. Pull-down PD is a drop-down menu for setting a predetermined time interval between the display of text data on monitoring screen G1 and the automatic selection of whether to display or not. In the example shown in Figure 3, button B1 is set to ON, button B2 to OFF, and the predetermined time is set to 3 seconds. In this case, after 3 seconds have elapsed since the text data was displayed on the monitoring screen, "OK" is automatically selected as the display option. On the other hand, if button B1 is OFF and button B2 is ON, after 3 seconds have elapsed since the text data was displayed on the monitoring screen, "NG" is automatically selected as the display option.
[0051] The text fields F1 to F4 on the monitoring screen G1 in Figure 3 are UI elements for the checker to modify the displayed text data. The monitoring processing unit 241 displays the text data of one language that has been pre-specified as the target of monitoring from among the multiple second language text data input from the speech conversion unit 232 in the corresponding text field. In the example shown in Figure 3, four text data points converted from four audio data points are displayed chronologically in text fields F1 to F4.
[0052] Buttons B3 and B4 on the monitoring screen G1 in Figure 3 are UI elements for the checker to select whether or not to display text data. Button B3 is for selecting "not displayable" (×). Button B4 is for selecting "displayable" (○). Buttons B3 and B4 are displayed for each text field F. Note that the selection of buttons B3 and B4 can be switched not only by the checker, but also by the monitoring processing unit 241 after a predetermined time has elapsed, at the start of text data modification, or after text data modification.
[0053] In the example shown in Figure 3, buttons B3-1 to B3-4 and buttons B4-1 to B4-4 are displayed in text fields F1 to F4, respectively. In the example of buttons B3-1 and B4-1 in text field F1, the checker determines that the text data displayed in text field F1 is accurate and displayable, and therefore button B4-1 is selected. In the example of buttons B3-2 and B4-2 in text field F2, the checker determined that the text data displayed in text field F2 was meaningless, and therefore button B3-2 was selected. In the example of buttons B3-3 and B4-3 for text field F3, the checker determines that the text data displayed in text field F3 is accurate and displayable, and therefore button B4-3 is selected. In the example of buttons B3-4 and B4-4 for text field F4, the checker determined that the text data displayed in text field F4 needed to be corrected and could not be displayed, so button F3-4 was selected. However, after the text data was corrected, button B4-4 was selected under the control of the monitoring processing unit 241.
[0054] Now, with reference to Figure 4, the correction operation procedure according to the first embodiment will be described. Figure 4 is a diagram showing an example of the correction operation procedure according to the first embodiment. Figure 4 shows an example in which the checker corrects the text data displayed in text field F5. Note that, for buttons B3-5 and B4-5, button B4-5 is selected as the initial selection.
[0055] As shown in Figure 4, first the checker examines the text data displayed in text field F5 and, if correction is needed, selects (touches) an arbitrary location in text field F5 (step S1). In the example shown in Figure 4, it is assumed that the location to be corrected (location P) has been selected.
[0056] After the checker selects text field F5, the monitoring processing unit 241 selects text field F5 (for example, by displaying a thick border), switches the display permission from buttons B4-5 to B3-5, displays cursor K inside text field F5, and displays window W1 showing conversion candidates near the area to be corrected (step S2). If the position of cursor K is not aligned with the area to be corrected, the checker moves the position of cursor K to the area to be corrected. If the subtitles for the text data displayed in text field F5 are already displayed on the display device 40, the subtitles for the text data displayed in text field F5 will be hidden (deleted) when the checker selects a section to be corrected.
[0057] The checker deletes the text to be corrected (step S3). In the example shown in Figure 4, the checker deletes the text indicating "concatenation".
[0058] The checker selects the correct text from the conversion candidates shown in window W1 (step S4). In the example shown in Figure 4, the checker selects the text that means "cooperation".
[0059] After the checker selects the correct text, the monitoring processing unit 241 inserts the selected text into the corrected location within the text field F5 (step S5). If the correct text is not among the conversion candidates displayed in window W1, the checker can manually enter the correct text into the corrected location.
[0060] The checker selects the area outside of text field F5 because the correction is complete (step S6). As a result, the monitoring processing unit 241 returns text field F5 to a deselected state (removes the bold border display) and switches the display enable / disable selection from button B3-5 to button B4-5. Furthermore, the monitoring processing unit 241 inputs the corrected text data into the machine translation unit 242.
[0061] (4-2) Machine Translation Section 242 The machine translation unit 242 has the function of performing machine translation processing. The machine translation unit 242 performs machine translation processing using the machine translation engine 22. In the machine translation processing, the machine translation unit 242 inputs the text data corrected by the monitoring processing unit 241 to the machine translation engine 22 and obtains the text data output from the machine translation engine 22 as the translation result. As a result, the machine translation unit 242 can translate the corrected text data into a second language that was not specified during the correction, and obtain text data showing the translation result with the corrections reflected for each of the multiple second languages. The machine translation unit 242 inputs the corrected text data obtained through the monitoring process and the text data in multiple second languages obtained through the machine translation process to the subtitle processing unit 243.
[0062] (4-3) Subtitle processing unit 243 The subtitle processing unit 243 has the function of performing subtitle display processing. In the subtitle display processing, the subtitle processing unit 243 displays subtitles on the display device 40 using multiple second language text data converted by the speech conversion unit 232 and input from the monitoring processing unit 241, or multiple second language text data input from the machine translation unit 242. When using multiple second language text data input from the monitoring processing unit 241, the subtitle processing unit 243 can display text data indicating translation results that the monitoring process determined do not require correction as subtitles on the display device 40. On the other hand, when using multiple second language text data input from the machine translation unit 242, the subtitle processing unit 243 can display text data indicating translation results that the monitoring process determined require correction and that reflect the corrections as subtitles on the display device 40.
[0063] <1-3. Processing Flow> The functional configuration of the audio processing device 20 according to the first embodiment has been described above. Next, with reference to Figure 5, the processing flow in the subtitle display system 1 according to the first embodiment will be described. Figure 5 is a sequence diagram showing an example of the processing flow in the subtitle display system 1 according to the first embodiment.
[0064] As shown in Figure 5, first, the sound collection device 10 transmits the collected audio data to the audio processing device 20 (step S101). The audio data acquisition unit 231 of the audio processing device 20 acquires the audio data that the communication unit 210 receives from the sound collection device 10.
[0065] Next, the speech conversion unit 232 of the speech processing device 20 requests simultaneous interpretation from the simultaneous interpretation engine 21 regarding the speech data acquired by the speech data acquisition unit 231 (step S102). The speech conversion unit 232 transmits the speech data to the simultaneous interpretation engine 21 via the communication unit 210.
[0066] Next, the simultaneous interpretation engine 21 performs simultaneous interpretation (speech recognition and machine translation) of the audio data received from the speech processing device 20 and transmits the result of the simultaneous interpretation to the speech processing device 20 (step S103). The result of the simultaneous interpretation is text data in multiple second languages.
[0067] Next, the monitoring processing unit 241 of the audio processing device 20 performs the display processing of the monitoring screen (step S104). The monitoring processing unit 241 transmits screen information to the monitoring terminal 30 via the communication unit 210 and displays the monitoring screen.
[0068] The monitoring terminal 30 displays a monitoring screen based on the screen information received from the voice processing device 20 (step S105). After the monitoring screen is displayed, the monitoring terminal 30 accepts the text data correction by the checker on the monitoring screen and sends correction information indicating the correction to the voice processing device 20 (step S106).
[0069] The monitoring processing unit 241 determines whether or not there has been a modification to the text data, depending on whether or not the communication unit 210 receives modification information from the monitoring terminal 30 (step S107). If there is a modification (step S107 / YES), the process proceeds to step S108. On the other hand, if there is no modification (step S107 / NO), the process proceeds to step S110.
[0070] If the process proceeds to step S108, the machine translation unit 242 of the speech processing device 20 requests the machine translation engine 22 to perform machine translation on the text data corrected by the checker (step S108). The machine translation unit 242 transmits the text data to the machine translation engine 22 via the communication unit 210.
[0071] Next, the machine translation engine 22 machine-translates the text data received from the speech processing device 20 and transmits the result of the machine translation back to the speech processing device 20 (step S109). The result of the machine translation is text data in multiple second languages. After transmission, the process proceeds to step S111.
[0072] If the process proceeds to step S110, the monitoring processing unit 241 determines whether the text data can be displayed (step S110). If the data can be displayed (step S110 / YES), the process proceeds to step S111. On the other hand, if the data cannot be displayed (step S110 / NO), the process terminates.
[0073] If the process proceeds to step S111, the subtitle processing unit 243 of the audio processing device 20 executes subtitle display processing (step S111). The subtitle processing unit 243 transmits text data in multiple second languages to the display device 40 via the communication unit 210. The display device 40 displays the text data in the second language received from the audio processing device 20 as subtitles. (Step S112).
[0074] The processing flow according to the first embodiment has been described above. As described above, the speech processing device 20 according to the first embodiment includes: a speech conversion unit 232 that converts speech data representing the content of a user's speech in a first language into text data representing the translation results into a plurality of second languages different from the first language by simultaneous interpretation; a monitoring processing unit 241 that accepts corrections to the text data for one designated second language from among the plurality of second languages; and a machine translation unit 242 that translates the corrected text data into a second language not specified at the time of correction and obtains text data representing the translation results reflecting the corrections for each of the plurality of second languages.
[0075] With this configuration, in multilingual simultaneous interpretation, if the simultaneous interpretation result contains misrecognition or mistranslation, the checker can correct the simultaneous interpretation results in the other languages by correcting only the simultaneous interpretation result in one of the multilingual languages of the translated word. Therefore, the speech processing device 20 according to the first embodiment makes it possible to easily correct the simultaneous interpretation results in multilingual simultaneous interpretation.
[0076] Furthermore, if the simultaneous interpretation of the speaker's speech contains misinterpretations or mistranslations, the checker can correct these before the simultaneous interpretation (subtitles) is presented to the audience. As a result, the audience will not be presented with the simultaneous interpretation containing the misinterpretations or mistranslations, but only with the corrected version that does not contain any misinterpretations or mistranslations. Therefore, the voice processing device 20 according to the first embodiment enables the user to correctly understand the content of the simultaneous translation.
[0077] <<2. Second Embodiment>> The first embodiment has been described above. Next, the second embodiment will be described with reference to Figures 6 to 12. In the second embodiment, a method for correcting text data on the monitoring screen will be described that differs from that of the first embodiment. In the following, explanations that overlap with the explanation of the first embodiment will be omitted as appropriate. In the second embodiment, the language spoken by the speaker (first language) is any one language (for example, Japanese), and the language of the multilingual translated subtitles (second language) is any multiple languages (for example, languages other than Japanese).
[0078] <2-1. Configuration of the subtitle display system> Referring to Figure 6, the configuration of the subtitle display system 1a according to the second embodiment will be described. Figure 6 is a diagram showing an example of the configuration of the subtitle display system 1a according to the second embodiment. As shown in Figure 6, the subtitle display system 1a comprises a sound collection device 10, an audio processing device 20a, a machine translation engine 22, an audio recognition engine 23, a conversion candidate API (Application Programming Interface) 24, a monitoring terminal 30, and a display device 40.
[0079] (1) Sound collection device 10 The sound collection device 10 according to the second embodiment is the same as the sound collection device 10 according to the first embodiment, so its description will be omitted.
[0080] (2) Audio processing device 20a The audio processing device 20a according to the second embodiment performs different processing compared to the audio processing device 20 according to the first embodiment due to a different method of correction on the monitoring screen. Based on the audio data received from the sound collection device 10, the audio processing device 20a further performs morphological analysis processing, conversion priority processing, and automatic conversion processing, in addition to audio conversion processing, monitoring processing, machine translation processing, and subtitle display processing.
[0081] (3) Machine translation engine 22 The machine translation engine 22 according to the second embodiment is the same as the machine translation engine 22 according to the first embodiment, so its description will be omitted.
[0082] (4) Speech recognition engine 23 The speech recognition engine 23 is an engine (program) that performs speech recognition on speech data. The speech recognition engine 23 converts speech data in a first language input from the speech processing device 20a into text data in the first language through speech recognition. The functions of the speech recognition engine 23 may be provided by a device or terminal different from the speech processing device 20a, or they may be provided by the speech processing device 20a.
[0083] (5) Conversion candidate API24 The conversion candidate API 24 is an API that obtains conversion candidates for the reading of text data. Based on information indicating the reading of the text data in the first language input from the speech processing device 20a, the conversion candidate API 24 obtains conversion candidates corresponding to the reading of the text data in the first language. The functionality of the conversion candidate API 24 may be provided by a device or terminal different from the voice processing device 20a, or it may be provided by the voice processing device 20a.
[0084] (6) Monitoring terminal 30 The monitoring terminal 30 according to the second embodiment is the same as the monitoring terminal 30 according to the first embodiment, so its description will be omitted.
[0085] (7) Display device 40 The display device 40 according to the second embodiment is the same as the display device 40 according to the first embodiment, so its description will be omitted.
[0086] <2-2. Functional Configuration of the Voice Processing Device> The configuration of the subtitle display system 1a according to the second embodiment has been described above. Next, the functional configuration of the audio processing device 20a according to the second embodiment will be described with reference to Figures 7 to 10. Figure 7 is a block diagram showing an example of the functional configuration of the audio processing device 20a according to the second embodiment. As shown in Figure 7, the audio processing device 20a comprises a communication unit 210a, a storage unit 220a, a first control unit 230a, and a second control unit 240a.
[0087] (1) Communications section 210a The communication unit 210a is connected to the sound collection device 10, the machine translation engine 22, the speech recognition engine 23, the conversion candidate API 24, the monitoring terminal 30, and the display device 40 in a communication manner, and transmits and receives various types of information.
[0088] (2) Memory unit 220a The memory unit 220a differs from the memory unit 220 of the first embodiment in that it also stores priority information. The priority information indicates the priority of automatic conversion at the morphological unit level obtained by morphological analysis of the text data. This priority is calculated based on past revision history monitored by the checker.
[0089] (3) First control unit 230a As shown in Figure 7, the first control unit 230a includes an audio data acquisition unit 231a and an audio conversion unit 232a.
[0090] (3-1) Audio data acquisition unit 231a The audio data acquisition unit 231a according to the second embodiment is the same as the audio data acquisition unit 231 according to the first embodiment, so its description will be omitted.
[0091] (3-2) Voice conversion unit 232a The speech conversion unit 232a differs from the speech conversion unit 232 in the first embodiment in that it performs only speech recognition and not simultaneous interpretation. The speech conversion unit 232a performs speech recognition processing as a speech conversion process using the speech recognition engine 23. In the speech recognition process, the speech conversion unit 232a inputs speech data representing the speaker's utterance in a first language, acquired by the speech data acquisition unit 231a, to the speech recognition engine 23, and acquires text data output from the speech recognition engine 23 as the speech recognition result. In this way, the speech conversion unit 232a can convert speech data representing the user's utterance in a first language into text data representing the user's utterance in a first language through speech recognition. The speech conversion unit 232a inputs the text data of the first language obtained by speech recognition processing to the monitoring processing unit 241a.
[0092] (4) Second control unit 240a As shown in Figure 7, the second control unit 240a includes a monitoring processing unit 241a, a machine translation unit 242a, a subtitle processing unit 243a, a morphological analysis unit 244, a conversion priority processing unit 245, and an automatic conversion processing unit 246.
[0093] (4-1) Monitoring processing unit 241a The monitoring processing unit 241a accepts correction operations on text data from the checker via the monitoring screen displayed on the monitoring terminal 30 during the monitoring process, but does not accept the operation to select whether or not to display the text data. Furthermore, the monitoring processing unit 241a divides the text data into multiple texts and displays them on the monitoring screen, and accepts corrections to the text data for each divided text unit. For this reason, the method of correcting text data on the monitoring screen differs between the first embodiment and the second embodiment.
[0094] The monitoring processing unit 241a displays a UI for displaying text data on the monitoring screen. This UI is, for example, a text field. On the monitoring screen, the monitoring processing unit 241a displays multiple text fields for a single piece of text data. Based on the results of morphological analysis by the morphological analysis unit 244, which will be described later, the monitoring processing unit 241a divides the text data into morphological units. A morphological unit is, for example, a part of speech unit. In this case, the monitoring processing unit 241a divides the text data into parts of speech units based on information indicating the results of the morphological analysis unit 244 dividing the text data into parts of speech units (hereinafter also referred to as "part of speech information"). The part of speech information includes, for example, the divided text and information indicating the reading of each text. After splitting, the monitoring processing unit 241a displays the split text in its respective text field. The monitoring processing unit 241a accepts modification operations for each of the multiple text fields for a single piece of text data. This allows the checker to modify the text data at the morphological level.
[0095] When a text to be modified is selected on the monitoring screen, the monitoring processing unit 241a displays conversion candidates in the vicinity of the selected text. The monitoring processing unit 241a obtains conversion candidates using the conversion candidate API 24. The monitoring processing unit 241a inputs the part-of-speech information obtained by the morphological analysis unit 244 to the conversion candidate API 24 and obtains information indicating the conversion candidates output from the conversion candidate API 24 (hereinafter also referred to as "conversion candidate information"). The conversion candidate API 24 refers to the reading of the text included in the part-of-speech information and obtains homonyms as conversion candidates. As a result, the monitoring processing unit 241a can display conversion candidates on the monitoring screen based on the obtained conversion candidate information. If a conversion candidate displayed on the monitoring screen is selected, the monitoring processing unit 241a replaces the text to be corrected with the text selected from the conversion candidates. In this way, the checker can correct text data at the morphological level by selecting a conversion candidate.
[0096] Furthermore, the monitoring processing unit 241a displays a text input field along with the conversion candidates, and accepts corrections by entering text into the input field. This allows the checker to appropriately correct the text data by entering appropriate text into the input field, for example, if there is no suitable conversion candidate among the conversion candidates.
[0097] The monitoring processing unit 241a may, when displaying text data on the monitoring screen, convert the text data based on past correction records before displaying it. For example, suppose the text data to be displayed contains text that corresponds to text that was previously corrected in other text data. The corresponding text is, for example, a homophone. In this case, the monitoring processing unit 241a converts the text in the text data to be displayed to the previously corrected text before displaying it. That is, when the monitoring processing unit 241a displays second text data obtained by speech recognition performed after the correction of the first text data has been made, if the second text data contains second text that corresponds to the first text that was corrected in the first text data, it converts the second text to the corrected first text before displaying it. As a result, texts with a history of past corrections are automatically converted by the monitoring processing unit 241a, thereby reducing the load on the checker during monitoring.
[0098] The monitoring processing unit 241a may display not only the text data converted from the audio data on the monitoring screen, but also the text data translated into different languages. If the text data converted from the audio data is modified on the monitoring screen, the monitoring processing unit 241a uses the machine translation engine 22 to obtain text data showing the translation result that reflects the modification, and updates the translation display with the obtained text data. The translation displayed on the monitoring screen by the monitoring processing unit 241a may be a translation of one of several second languages, or it may be a translation of multiple languages.
[0099] Now, with reference to Figure 8, a monitoring screen according to the second embodiment will be described. Figure 8 is a diagram showing an example of a monitoring screen according to the second embodiment.
[0100] As an example, the monitoring screen G2 in Figure 8 displays, for two audio data sets, the text data obtained through speech recognition processing and the translation results obtained through machine translation of each text data set.
[0101] The first audio data is, for example, audio representing the utterance "We assign a conductor at the company." This audio data is converted into text data, "We assign a conductor at the company," by the speech recognition processing of the speech conversion unit 232a. The monitoring processing unit 241a divides this text data into parts of speech based on the results of morphological analysis and then displays it in text field F11. In the example shown in Figure 8, the text data is divided into seven parts of speech (words): "company," "at," "is," "conductor," "attach," and "do," and is displayed in text field F11. Furthermore, the monitoring processing unit 241a displays text data in text field F12 that shows the result of the text data being translated by the machine translation unit 242a. In the example shown in Figure 8, the text data is translated to "There is a conductor in the company." and displayed in text field F12. Note that the text data displayed in text field F11 is the first text data, which is the text data obtained by speech recognition from audio data representing the speaker's utterances between a given time T1 and T2. Assume that time T2 is later than time T1. The text data displayed in text field F12 is text data that shows the translation result into any of the pre-set languages (multiple second languages) that are the target of multilingual translation. In the second embodiment, as an example, it is set to display the translation result from Japanese to English.
[0102] The second audio data is, for example, audio representing the utterance, "We mainly attach our company emblem with an important formula." This audio data is converted into text data, "We mainly attach our company emblem with an important formula," by the speech recognition processing of the speech conversion unit 232a. The monitoring processing unit 241a divides this text data into parts of speech based on the results of morphological analysis and then displays it in text field F13. In the example shown in Figure 8, the text data is divided into nine parts of speech (words): "company emblem," "is," "mainly," "important," "form," "attach," and "do," and these are displayed in text field F13. Furthermore, the monitoring processing unit 241a displays text data in text field F14 that shows the result of the text data being translated by the machine translation unit 242a. In the example shown in Figure 8, the text data is translated to "We put the company emblem mainly at important ceremonies." and displayed in text field F14. Note that the text data displayed in text field F13 is the second set of text data, which is the text data obtained by speech recognition from audio data representing the speaker's utterances between a given time T3 and T4. Assume that time T3 is after time T2, and time T4 is after time T3.
[0103] Now, with reference to Figure 9, the modification procedure according to the second embodiment will be described. Figure 9 is a diagram showing an example of the modification procedure according to the second embodiment. Figure 9 shows an example of modifying the text data in text field F11 in Figure 8.
[0104] As shown in Figure 9, first the checker examines the text data displayed in text field F11 and selects (touches) the text that needs to be corrected (step S11). In the example shown in Figure 9, it is assumed that "conductor" is selected as the text to be corrected.
[0105] After the checker selects text, the monitoring processing unit 241a displays a conversion candidate window W11, a text field F15, an add button B5, and a delete button B6 near the selected text (step S12). The checker selects appropriate text for correction from the conversion candidates displayed in the conversion candidate window W11. If there is no appropriate text among the conversion candidates, the checker enters appropriate text into the text field F15 (input field) and presses the add button B5. If the checker wants to delete the text selected in the text field F11, it presses the delete button B6. If the text selected from text field F11 by the checker has been automatically converted, the text before the conversion will be displayed in text field F15. In this case, the checker can revert the text selected in text field F11 back to the text before the conversion by pressing the add button B5.
[0106] After the checker performs the correction operation, the monitoring processing unit 241 displays the text selected in text field F11 with the corrected text (step S13). Furthermore, the monitoring processing unit 241a displays the translated text data of the corrected text data shown in text field F11 in text field F12. In the example shown in Figure 9, "conductor" has been corrected to "company logo." As a result, in the example shown in Figure 8, the text recognized as "conductor" in the speech recognition of the second audio data is automatically converted to "company logo" before being displayed in text field F13.
[0107] (4-2) Machine Translation Section 242a The machine translation unit 242a translates the text data converted from the audio data of the first language by the speech recognition processing of the speech conversion unit 232a into a second language different from the first language, and obtains text data showing the translation result. As a result, the machine translation unit 242a can obtain text data showing the translation result, which is displayed on the monitoring screen along with the text data converted from the audio data. Furthermore, when the text data displayed on the monitoring screen is corrected, the machine translation unit 242a translates the corrected text data and obtains text data showing the translation result with the corrections reflected. This allows the machine translation unit 242a to obtain text data with the corrections reflected, which can be checked by the checker on the monitoring screen and displayed as subtitles on the display device 40.
[0108] (4-3) Subtitle processing unit 243a The subtitle processing unit 243a according to the second embodiment is the same as the subtitle processing unit 243 according to the first embodiment, so its description will be omitted.
[0109] (4-4) Morphological analysis section 244 The morphological analysis unit 244 has the function of performing morphological analysis. The morphological analysis unit 244 performs morphological analysis on text data converted from speech data. Through morphological analysis, the morphological analysis unit 244 divides the text data into multiple texts, for example, by part of speech.
[0110] (4-5) Conversion priority processing unit 245 The conversion priority processing unit 245 has the function of processing priority information. Priority information is information indicating the priority of the text to be used for conversion in automatic text conversion. The priority is determined based on the past revision history. When the text data displayed on the monitoring screen is modified, the conversion priority processing unit 245 changes the priority (weight) according to the content of the modification.
[0111] Now, with reference to Figure 10, the modification of the conversion priority according to the second embodiment will be described. Figure 10 is a diagram showing an example of the modification of the conversion priority according to the second embodiment. The table on the left in Figure 10 shows an example of the priority before modification, and the table on the right shows an example of the priority after modification.
[0112] Each table in Figure 10 shows the original text, the converted text, and the priority (weight) as priority information. This priority information indicates that if the speech-recognized text data contains the original text, it will be converted to the converted text with the highest priority. The table on the left shows that the priority for converting "park" to "lecture" is "1.05", the priority for converting "park" to "performance" is "1.00", the priority for converting "park" to "oral presentation" is "0.95", and the priority for converting "lecture" to "performance" is "1.00". Multiple priority options are shown for "park". If "park" is included in the speech-recognized text data, it will be automatically converted to "lecture", which has the highest priority. When the speech-recognized text data is corrected by the checker, the conversion priority processing unit 245 changes the priority information. For example, suppose the speech-recognized text data contains "lecture" and the checker corrects "lecture" to "performance". In this case, the conversion priority processing unit 245 changes the table shown on the left side of Figure 10 to the table shown on the right side. Because "lecture" has been corrected to "performance", the conversion priority processing unit 245 increases the priority of converting to "performance" and decreases the priority of converting to anything other than "performance".
[0113] (4-6) Automatic conversion processing unit 246 The automatic conversion processing unit 246 has the function of automatically converting text. Before displaying text data that has been split into multiple texts, the automatic conversion processing unit 246 refers to priority information. If the text to be automatically converted is included in multiple texts, the automatic conversion processing unit 246 selects a text from the texts used in past corrections according to priority and converts the target text with the selected text. The texts used in past corrections are the corrected texts from when the checker corrected the text data in the past, and are the converted texts in the priority information shown in Figure 10. The text selected by the automatic conversion processing unit 246 according to its priority is, for example, the text with the highest priority. If only one priority piece of information exists for a single text, the automatic conversion processing unit 246 selects the text indicated by that priority piece of information and performs the automatic conversion. If multiple priority pieces of information exist for a single text, the automatic conversion processing unit 246 selects the text indicated by the priority piece of information with the highest priority and performs the automatic conversion.
[0114] <2-3. Processing Flow> The functional configuration of the audio processing device 20a according to the second embodiment has been described above. Next, the processing flow according to the second embodiment will be described with reference to Figures 11 to 12.
[0115] (1) Processing flow in subtitle display system 1a Referring to Figure 11, the processing flow in the subtitle display system 1a according to the second embodiment will be described. Figure 11 is a sequence diagram showing an example of the processing flow in the subtitle display system 1a according to the second embodiment.
[0116] As shown in Figure 11, first, the sound collection device 10 transmits the collected audio data to the audio processing device 20a (step S201). The audio data acquisition unit 231a of the audio processing device 20a acquires the audio data that the communication unit 210a receives from the sound collection device 10.
[0117] Next, the voice conversion unit 232a of the voice processing device 20a performs voice conversion processing (step S202). The voice conversion unit 232a transmits the voice data acquired by the voice data acquisition unit 231a to the voice recognition engine 23 via the communication unit 210a and requests voice recognition. The voice recognition engine 23 performs voice recognition on the voice data received from the voice processing device 20a and transmits the result of the voice recognition back to the voice processing device 20a.
[0118] Next, the machine translation unit 242a of the speech processing device 20a performs machine translation processing (step S203). The machine translation unit 242a transmits the text data acquired by the speech conversion unit 232a to the machine translation engine 22 via the communication unit 210a and requests machine translation. The machine translation engine 22 machine translates the text data received from the speech processing device 20a and transmits the result of the machine translation back to the speech processing device 20a.
[0119] Next, the audio processing device 20a performs display preparation processing (step S204). Display preparation processing is the process of preparing to display the monitoring screen. Details of the display preparation processing will be described later.
[0120] Next, the monitoring processing unit 241a of the audio processing device 20a performs the display processing of the monitoring screen (step S205). The monitoring processing unit 241a transmits screen information to the monitoring terminal 30 via the communication unit 210a and displays the monitoring screen.
[0121] The monitoring terminal 30 displays a monitoring screen based on the screen information received from the voice processing device 20a (step S206). After the monitoring screen is displayed, the monitoring terminal 30 accepts the text data correction by the checker on the monitoring screen and sends correction information indicating the correction to the voice processing device 20a (step S207).
[0122] The monitoring processing unit 241a determines whether or not there has been a modification to the text data, depending on whether or not the communication unit 210a receives modification information from the monitoring terminal 30 (step S208). If there is a modification (step S208 / YES), the process proceeds to step S209. On the other hand, if there is no modification (step S208 / NO), the process proceeds to step S212.
[0123] If the process proceeds to step S209, the machine translation unit 242a of the speech processing device 20a performs machine translation (step S209). The machine translation unit 242a sends the text data corrected by the checker to the machine translation engine 22 via the communication unit 210a and requests machine translation. The machine translation engine 22 machine translates the text data received from the speech processing device 20a and sends the result of the machine translation back to the speech processing device 20a.
[0124] Next, the monitoring processing unit 241a performs a monitoring screen update process (step S210). Based on the results of the machine translation obtained by the machine translation unit 242a, the monitoring processing unit 241a updates the translation display on the monitoring screen.
[0125] Next, the conversion priority processing unit 245 of the voice processing device 20a updates the priority (step S211). Based on the correction information received by the communication unit 210a from the monitoring terminal 30, the conversion priority processing unit 245 updates the priority of the priority information corresponding to the correction information.
[0126] If the process proceeds to step S212, the subtitle processing unit 243a of the audio processing device 20a executes subtitle display processing (step S212). The subtitle processing unit 243a transmits text data in multiple second languages to the display device 40 via the communication unit 210a. The display device 40 displays the text data in the second language received from the audio processing device 20a as subtitles. (Step S213)
[0127] (2) Flow of display preparation process Referring to Figure 12, the flow of the display preparation process according to the second embodiment will be described. Figure 12 is a sequence diagram showing an example of the flow of the display preparation process according to the second embodiment.
[0128] As shown in Figure 12, first, the morphological analysis unit 244 of the speech processing device 20a performs morphological analysis on the text data acquired by the speech conversion unit 232a (step S301).
[0129] Next, the automatic conversion processing unit 246 of the voice processing device 20a acquires priority information stored in the memory unit 220a (step S302).
[0130] The automatic conversion processing unit 246 refers to priority information and checks whether there are any parts of speech (text) with a high priority for automatic conversion in the morphologically analyzed text data (step S303). If there are parts of speech with a high priority for automatic conversion (step S303 / YES), the process proceeds to step S304. On the other hand, if there are no parts of speech with a high priority for automatic conversion (step S303 / NO), the process proceeds to step S305.
[0131] If the process proceeds to step S304, the automatic conversion processing unit 246 performs the automatic conversion (step S304). After the automatic conversion, the process proceeds to step S305.
[0132] If the process proceeds to step S305, the monitoring processing unit 241a of the speech processing device 20a transmits the part-of-speech information obtained by the morphological analysis unit 244 to the conversion candidate API 24 via the communication unit 210a (step S305).
[0133] The conversion candidate API 24 obtains conversion candidate information based on the part-of-speech information received from the speech processing device 20a and transmits it to the speech processing device 20a (step S306).
[0134] The processing from steps S302 to S306 is performed on a part-of-speech basis (text unit) based on the morphological analysis results from step S301. Therefore, the processing from steps S302 to S306 is repeated until conversion candidate information is obtained for all parts of speech. Once conversion candidate information has been obtained for all parts of speech, the display preparation process is completed.
[0135] The processing flow according to the second embodiment has been described above. As described above, the voice processing device 20a according to the second embodiment includes a voice conversion unit 232a that converts voice data representing the content of a user's speech into text data by voice recognition, and a monitoring processing unit 241a that divides the text data into multiple texts for display and accepts modifications to the text data for each divided text unit. When the monitoring processing unit 241a displays the second text data obtained by voice recognition performed after the modification of the first text data has been made, if the second text data contains a second text that corresponds to the first text that has been modified in the first text data, it converts the second text to the modified first text and displays it.
[0136] With this configuration, when the checker corrects a misrecognition that occurred during speech recognition of a given audio file, any subsequent instances of the same misrecognition occurring during speech recognition of a different audio file will be automatically converted to correct text. This eliminates the need for the checker to perform the same correction every time the same misrecognition occurs. Therefore, the speech processing device 20a according to the second embodiment makes it possible to reduce the workload associated with correcting misrecognitions in speech recognition.
[0137] Furthermore, if the simultaneous interpretation of the speaker's speech contains misinterpretations or mistranslations, the checker can correct these before the simultaneous interpretation (subtitles) is presented to the audience. As a result, the audience will not be presented with the simultaneous interpretation containing the misinterpretations or mistranslations, but only with the corrected version that does not contain any misinterpretations or mistranslations. Therefore, the voice processing device 20a according to the second embodiment enables the user to correctly understand the content of the simultaneous translation.
[0138] <<3. Variant Example>> The embodiments have been described above. Next, modifications of the embodiments described above will be explained. The modifications described below may be applied to the embodiments individually or in combination. Furthermore, the modifications may be applied in place of the configuration described in the embodiments, or they may be applied in addition to the configuration described in the embodiments.
[0139] The first and second embodiments described above may be implemented in combination. For example, the monitoring screen of the first embodiment may display a translation, similar to the monitoring screen of the second embodiment. Furthermore, the monitoring screen of the second embodiment may display a UI that allows the checker to select whether or not to display text data, similar to the monitoring screen of the first embodiment.
[0140] Furthermore, while the first embodiment described above included an example where the first language is any language other than Japanese and the second language is one of several languages including Japanese, the embodiment is not limited to such an example. For example, the first language may be Japanese and the second language may be a language other than Japanese. Furthermore, while the second embodiment described above describes an example where the first language is Japanese and the second language is multiple languages other than Japanese, the embodiment is not limited to such an example. For example, the first language may be a language other than Japanese, and the second language may be multiple languages including Japanese. Furthermore, in each of the embodiments described above, the display device 40 may be capable of displaying subtitles in the first language.
[0141] The above describes some modified examples of the embodiments. Furthermore, some or all of the functions of the subtitle display system 1,1a and the audio processing device 20,20a in the above-described embodiment may be implemented by a computer. In that case, the functions may be implemented by recording a program for implementing these functions on a computer-readable recording medium, and then loading and executing the program recorded on this recording medium into a computer system. Hereinafter, "computer system" includes hardware such as the OS and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. In addition, "computer-readable recording media" may also include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs over networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside computer systems that act as servers or clients in such cases. Furthermore, the above program may be for the purpose of implementing some of the functions described above, or it may be for the purpose of implementing the above functions in combination with a program already recorded in the computer system, or it may be implemented using a programmable logic device such as an FPGA (Field Programmable Gate Array).
[0142] Although embodiments of this invention have been described in detail above with reference to the drawings, the specific configuration is not limited to those described above, and various design changes can be made without departing from the spirit of this invention. [Explanation of symbols]
[0143] 1,1a…Subtitle display system, 10…Sound collection device, 20,20a…Speech processing device, 21…Simultaneous interpretation engine, 22…Machine translation engine, 23…Speech recognition engine, 24…Conversion candidate API, 30…Monitoring terminal, 40…Display device, 41…Screen, 42…Smartphone, 210,210a…Communication unit, 220,220a…Storage unit, 230,230a…First control unit, 231,231a…Speech data acquisition unit, 232,232a…Speech conversion unit, 240,240a…Second control unit, 241,241a…Monitoring processing unit, 242,242a…Machine translation unit, 243,243a…Subtitle processing unit, 244…Morphological analysis unit, 245…Conversion priority processing unit, 246…Automatic conversion processing unit
Claims
1. A speech conversion unit that converts speech data representing the content of a speaker's utterance in a first language into text data representing the content of the speaker's utterance in the first language by speech recognition, A monitoring processing unit that divides the aforementioned text data into multiple texts and displays them on a terminal used by the checker for monitoring, and receives corrections to the aforementioned text data from the checker on a per-divided text basis, A translation unit that translates the corrected text data into multiple second languages different from the first language, and obtains text data showing the translation result with the corrections reflected for each of the multiple second languages, A subtitle processing unit that displays text data showing the corrected translation result as subtitles on a display device used by the audience to check the subtitles. Equipped with, When the monitoring processing unit displays the second text data obtained by speech recognition performed after the modification of the first text data, if the second text data contains the second text corresponding to the first text modified in the first text data, it converts the second text to the modified first text and displays it. Voice processing device.
2. Before the text data, which has been divided into multiple texts, is displayed, the automatic conversion processing unit refers to priority information indicating the priority of automatic conversion based on past revision history, and if the text to be automatically converted is included in the multiple texts, it selects a text from the texts used in past revisions according to the priority, and converts the target text with the selected text. The audio processing device according to claim 1, further comprising:
3. If the displayed text data is modified, a conversion priority processing unit changes the priority according to the modification. The audio processing device according to claim 2, further comprising:
4. The translation unit translates the text data converted from the audio data representing the content of the utterance in the first language into a second language different from the first language, and obtains text data representing the translation result. The monitoring processing unit displays the text data converted from the audio data and the translated text data, and displays a monitoring screen on the monitoring terminal that can accept modification operations on the text data converted from the audio data. The voice processing device according to claim 1.
5. When the text data converted from the audio data displayed on the monitoring screen is corrected, the translation unit translates the corrected text data and obtains text data showing the translation result that reflects the correction. The monitoring processing unit updates the translation display with text data showing the translation result that reflects the corrections. The audio processing device according to claim 4.
6. When the monitoring processing unit detects a text to be modified on the monitoring screen, it displays conversion candidates near the selected text and replaces the text to be modified with the text selected from the conversion candidates. The audio processing device according to claim 4.
7. A morphological analysis unit performs morphological analysis on the text data converted from the audio data. Furthermore, The monitoring processing unit divides the text data into morphological units and displays them based on the results of the morphological analysis. The voice processing device according to claim 1.
8. A speech conversion process that converts speech data representing the content of a speaker's utterance in a first language into text data representing the content of the speaker's utterance in the first language by speech recognition, A monitoring process that divides the aforementioned text data into multiple texts and displays them on a terminal used by the checker for monitoring, and accepts corrections to the aforementioned text data from the checker for each divided text unit, A translation process that involves translating the corrected text data into multiple second languages different from the first language, and obtaining text data showing the translation results that reflect the corrections for each of the multiple second languages, A subtitle processing process that displays text data showing the corrected translation result as subtitles on a display device used by the audience to view the subtitles. Includes, The monitoring process, when displaying the second text data obtained by speech recognition performed after the modification of the first text data, if the second text data contains the first text that has been modified in the first text data, converts the second text to the modified first text and displays it. A method of audio processing performed by a computer.
9. Computers, A speech conversion means that converts speech data representing the content of a speaker's utterance in a first language into text data representing the content of the speaker's utterance in the first language by speech recognition, A monitoring processing means that divides the aforementioned text data into multiple texts and displays them on a terminal used by the checker for monitoring, and receives corrections to the aforementioned text data from the checker on a per-divided text basis, Translation means that translates the corrected text data into a plurality of second languages different from the first language, and obtains text data showing the translation result with the corrections reflected for each of the plurality of second languages, Subtitle processing means that displays text data showing the corrected translation result as subtitles on a display device used by the audience to check the subtitles. To make it function as, When the monitoring processing means displays the second text data obtained by speech recognition performed after the modification of the first text data, if the second text data contains the second text corresponding to the first text modified in the first text data, it converts the second text to the modified first text and displays it. program.
Citation Information
Patent Citations
Device and method for speech recognition, storage medium, and program
JP2005234236A
Voice recognition device, voice recognition system, and voice recognition program
JP2011197410A
Word registering apparatus, and computer program for the same
JP2014048506A
Text correction device, text correction method and text correction program
JP2019148681A
Text correction device and text correction method
JP2020197592A