Audio processing device, audio processing method, and program

The speech processing device facilitates easy correction of misinterpretations in simultaneous interpretation by converting, monitoring, and translating speech into multiple languages, ensuring accurate translation results.

JP2026006104AActive Publication Date: 2026-01-16TOPPAN HOLDINGS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024104877
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-16
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

Existing simultaneous interpretation systems suffer from misrecognitions and mistranslations, requiring manual correction in multiple languages, which is cumbersome and inefficient.

Method used

A speech processing device and method that includes a speech conversion unit for simultaneous interpretation into multiple languages, a monitoring unit for correcting text data in one specified language, and a translation unit to reflect corrections in all languages, facilitating easy correction and translation of misinterpreted text.

Benefits of technology

Enables easy correction of simultaneous interpretation results across multiple languages, ensuring accurate translation without misrecognitions or mistranslations, improving user understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026006104000001_ABST
    Figure 2026006104000001_ABST
Patent Text Reader

Abstract

To provide a voice processing device, a voice processing method, and a program capable of easily correcting a simultaneous interpretation result in multilingual simultaneous interpretation.SOLUTION: A voice conversion unit configured to convert voice data indicating a speech content of a user in a first language into text data indicating a translation result into a plurality of second languages different from the first language by simultaneous interpretation; A translation unit configured to translate the text data after the correction into the second language that is not designated at the time of the correction and acquire text data indicating a translation result in which the correction is reflected for each of the plurality of second languages.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an audio processing device, an audio processing method, and a program. [Background technology]

[0002] In recent years, advances in AI (Artificial Intelligence) have improved the accuracy of AI-based speech recognition and machine translation, but recognition and translation results may still contain misrecognitions or mistranslations.

[0003] In this regard, Patent Document 1 below discloses a technique for correcting the recognition results of speech recognition. This technique is used, for example, when correcting subtitles generated based on the speech recognition results for video data that includes audio data. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-148681 Summary of the Invention [Problem to be solved by the invention]

[0005] Simultaneous interpretation is possible by combining speech recognition and machine translation. However, even with simultaneous interpretation, the recognition or translation results may contain misrecognitions or mistranslations. For this reason, in multilingual simultaneous interpretation, if the simultaneous interpretation results contain at least one misrecognition or mistranslation, the checker must correct the simultaneous interpretation results for each language, which is not easy.

[0006] In view of the above-mentioned problems, an object of the present invention is to provide a speech processing device, a speech processing method, and a program that can easily correct the results of simultaneous interpretation in simultaneous interpretation of multiple languages. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems, a speech processing device according to one embodiment of the present invention is a speech processing device that includes: a speech conversion unit that converts speech data indicating the content of a user's utterance in a first language into text data indicating the translation results into a plurality of second languages ​​different from the first language through simultaneous interpretation; a monitoring processing unit that accepts corrections to the text data for one specified second language from among the plurality of second languages; and a translation unit that translates the corrected text data into a second language that was not specified at the time of correction and obtains text data indicating the translation results in which the corrections are reflected for each of the plurality of second languages.

[0008] A speech processing method according to one embodiment of the present invention is a speech processing method executed by a computer, including: a speech conversion process for converting speech data indicating the content of a user's speech in a first language into text data indicating the translation results into multiple second languages ​​different from the first language by simultaneous interpretation; a monitoring process for accepting corrections to the text data for one specified second language from among the multiple second languages; and a translation process for translating the corrected text data into a second language that was not specified at the time of correction, and obtaining text data indicating the translation results in which the corrections are reflected for each of the multiple second languages.

[0009] A program according to one embodiment of the present invention causes a computer to function as: a speech conversion means for converting, by simultaneous interpretation, speech data indicating the content of a user's speech in a first language into text data indicating the translation results into a plurality of second languages ​​different from the first language; a monitoring processing means for accepting corrections to the text data for one specified second language from among the plurality of second languages; and a translation means for translating the corrected text data into a second language that was not specified at the time of correction, and obtaining text data indicating the translation results in which the corrections are reflected for each of the plurality of second languages. [Effects of the Invention]

[0010] According to the present invention, simultaneous interpretation results in multiple languages ​​can be easily corrected. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram illustrating an example of a configuration of a subtitle display system according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of a functional configuration of a voice processing device according to a first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of a monitoring screen according to the first embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of a correction operation procedure according to the first embodiment. [Figure 5] FIG. 3 is a sequence diagram showing an example of a processing flow in the subtitle display system according to the first embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of the configuration of a subtitle display system according to a second embodiment. [Figure 7] FIG. 10 is a block diagram showing an example of a functional configuration of a voice processing device according to a second embodiment. [Figure 8] FIG. 10 is a diagram showing an example of a monitoring screen according to the second embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of a correction operation procedure according to the second embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of change of conversion priority according to the second embodiment. [Figure 11] FIG. 10 is a sequence diagram showing an example of a processing flow in the subtitle display system according to the second embodiment. [Figure 12] FIG. 10 is a sequence diagram showing an example of the flow of a display preparation process according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0013] <<1. First Embodiment>> A first embodiment will be described with reference to Figures 1 to 5. A subtitle display system will be described below. The subtitle display system is a system for translating a user's speech into multiple languages ​​different from the language used by the user (multilingual translation) and displaying subtitles in each language (multilingual subtitles). In the following, a first embodiment will be described, taking as an example a case in which subtitles obtained by translating the speech content (lecture content) of a user (lecture speaker) giving a lecture at a lecture into multiple languages ​​are distributed to users (listeners) listening to the lecture. In the first embodiment, it is assumed that the language spoken by the lecturer (first language) is any one language (e.g., any language other than Japanese), and the languages ​​of the translated subtitles (second languages) are any multiple languages ​​(e.g., multiple languages ​​including Japanese).

[0014] <1-1. Configuration of the subtitle display system> The configuration of the subtitle display system according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of the subtitle display system according to the first embodiment.

[0015] As shown in FIG. 1, the subtitle display system 1 includes a sound collection device 10, a sound processing device 20, a simultaneous interpretation engine 21, a machine translation engine 22, a monitoring terminal 30, and a display device 40. Each device and terminal are connected to each other via a wired connection, a wireless connection, or a network connection so that various information can be transmitted and received. The network may be, for example, a local area network (LAN), a wide area network (WAN), a telephone network (such as a mobile phone network or a fixed telephone network), a regional Internet Protocol (IP) network, or the Internet.

[0016] (1) Sound collection device 10 The sound collection device 10 is a device that collects the sound generated when a speaker speaks, and is, for example, a microphone. The sound collection device 10 is connected to the sound processing device 20 via a wired or wireless connection so that they can communicate with each other. When the sound collection device 10 collects the voice of the speaker, it converts the collected voice into data and transmits the converted voice data to the sound processing device 20.

[0017] (2) Audio processing device 20 The voice processing device 20 is a device that performs processing to display the results of simultaneous interpretation as subtitles. The voice processing device 20 is realized by devices such as one or more servers (e.g., cloud servers), PCs (Personal Computers), etc. In the device, various processes are executed by a program that causes the device to function as the voice processing device 20. The sound processing device 20 performs various processes based on the sound data received from the sound collection device 10. The sound processing device 20 performs, for example, a sound conversion process, a monitoring process, a machine translation process, a subtitle display process, and the like.

[0018] The voice conversion process is a process in which voice data is converted into text data. In the voice conversion process, the voice processing device 20 performs voice recognition and machine translation of the voice data received from the sound collection device 10 using the functions of the simultaneous interpretation engine 21 described later, and obtains text data indicating the translation results. In the voice conversion process, one piece of text data expressed in a first language is machine-translated into multiple pieces of text data expressed in multiple second languages ​​different from the first language.

[0019] The monitoring process is a process executed by the checker to monitor the results of the voice conversion process on the voice data (voice conversion result). In the monitoring process, the voice processing device 20 displays a monitoring screen on the monitoring terminal 30. The monitoring screen is a screen that can accept a selection operation for whether or not to display the text data converted from the voice data, and an operation for correcting the text data.

[0020] The checker can monitor the voice conversion result by checking the voice conversion result displayed on the monitoring screen of the monitoring terminal 30. During monitoring, the checker determines whether or not the text data converted from the voice data in the voice conversion process can be displayed as subtitles (displayability). Criteria for determining whether or not the text data can be displayed include, for example, whether or not there is a recognition error or a conversion error in the voice conversion result, or whether or not the meaning of the text data is understandable. If there is a recognition error or a conversion error in the voice conversion result, or if the meaning of the text data is incomprehensible, the checker determines that the text data cannot be displayed as subtitles. On the other hand, if there is no recognition error or a conversion error in the voice conversion result and the meaning of the text data is understandable, the checker determines that the text data can be displayed as subtitles.

[0021] Depending on the determination result, the checker selects whether or not the text data can be displayed on the monitoring screen, and performs correction operations as necessary. If it is determined that the text data can be displayed, the checker performs an operation to select "displayable" on the monitoring screen. On the other hand, if it is determined that the text data cannot be displayed, the checker performs an operation to select "displayable" on the monitoring screen, and corrects the text data so that it can be displayed. Once the checker has completed the correction operation, the displayability of the corrected text data is automatically changed to "displayable" by the audio processing device 20.

[0022] The monitoring screen may be provided with an automatic selection function for whether or not to display. The automatic selection function is a function that automatically selects whether to display or not display after a predetermined time has elapsed since text data was displayed on the monitoring screen. The predetermined time can be set by any user (e.g., a checker or an administrator) on the monitoring screen to any time (e.g., 3 seconds). Also, whether to automatically select whether to display or not display can be set by any user on the monitoring screen.

[0023] Furthermore, the voice conversion results displayed on the monitoring screen are text data in only one language. The checker only needs to monitor, on the monitoring screen, the text data in one second language among the text data in multiple second languages ​​converted by the voice conversion process. Note that the text data in the second language to be displayed on the monitoring screen among the text data in multiple second languages ​​can be appropriately selected, for example, depending on the languages ​​that the checker can handle.

[0024] The machine translation process is a process in which text data is translated. In the machine translation process, the speech processing device 20 machine-translates the text data that has been determined to be undisplayable in the monitoring process and has been corrected, using the function of the machine translation engine 22 (described later), and acquires text data indicating the translation results. In the machine translation process, one piece of text data expressed in a first language is machine-translated into multiple pieces of text data expressed in multiple second languages ​​different from the first language. Therefore, the speech processing device 20 can acquire, through the machine translation process, text data in which the corrections have been reflected in multiple other languages ​​from one piece of text data corrected in one language.

[0025] The subtitle display process is a process in which subtitles are displayed. In the subtitle display process, the voice processing device 20 causes the display device 40 to display text data indicating the translation result of the machine translation process as subtitles. The audio processing device 20 may cause the display device 40 to display only the subtitles in one designated second language, or may cause the display device 40 to display the subtitles in a plurality of second languages.

[0026] (3) Simultaneous Interpretation Engine 21 The simultaneous interpretation engine 21 is an engine (program) that simultaneously interprets a first language into a second language. The simultaneous interpretation engine 21 converts the voice data of the first language input from the voice processing device 20 into text data of the first language by voice recognition, and machine translates (converts) the text data of the first language into text data of the second language. The simultaneous interpretation engine 21 machine translates one first language into multiple different second languages. In other words, the simultaneous interpretation engine 21 generates text data of multiple second languages ​​from text data of one first language. The function of the simultaneous interpretation engine 21 may be provided by a device or terminal different from the speech processing device 20, or may be provided by the speech processing device 20.

[0027] (4) Machine Translation Engine 22 The machine translation engine 22 is an engine (program) that machine translates a first language into a second language. The machine translation engine 22 machine translates (converts) text data in the first language input from the speech processing device 20 into text data in the second language. The machine translation engine 22 machine translates one first language into multiple different second languages. In other words, the machine translation engine 22 generates text data in multiple second languages ​​from text data in one first language. The function of the machine translation engine 22 may be provided by a device or terminal different from the speech processing device 20, or may be provided by the speech processing device 20 itself.

[0028] (5) Monitoring terminal 30 The monitoring terminal 30 is a terminal used by a checker for monitoring. The monitoring terminal 30 is, for example, a PC, a smartphone, a tablet terminal, etc. Various processes are executed on the terminal by a program that causes the terminal to function as the monitoring terminal 30. The monitoring terminal 30 displays a monitoring screen based on the screen information received from the voice processing device 20, and accepts various operations related to monitoring by the checker.

[0029] For example, a monitoring screen is displayed on the monitoring terminal 30 by an application for utilizing the monitoring function (hereinafter also referred to as a "monitoring app") The checker can monitor the simultaneous interpretation results by operating the monitoring screen displayed on the monitoring terminal 30 by the monitoring app. The functions of the monitoring app may be provided by installing the monitoring app on each device (i.e., a native app), or by a web system (i.e., a web app). In the case of a web app, the monitoring app is managed by a server, and its functions are provided via a web browser.

[0030] (6) Display device 40 The display device 40 is a device that displays subtitles. The display device 40 may be, for example, a display device such as a screen 41, or a device having a display such as a smartphone 42. The display device 40 is communicably connected to the audio processing device 20, and displays subtitles based on screen information received from the audio processing device 20.

[0031] A subtitle display screen is displayed on the display device 40 by, for example, an application for utilizing the subtitle display function (hereinafter also referred to as a "subtitle display application"). Listeners can check the subtitles showing the results of the simultaneous interpretation by operating the subtitle display screen displayed on the display device 40 by the subtitle display application. The functions of the subtitle display app may be provided by installing a subtitle display app on each device (i.e., a native app), or may be provided by a web system (i.e., a web app). In the case of a web app, the subtitle display app is managed by a server, and its functions are provided via a web browser.

[0032] <1-2. Functional configuration of voice processing device> The configuration of the subtitle display system 1 according to the first embodiment has been described above. Next, the functional configuration of the audio processing device 20 according to the first embodiment will be described with reference to Fig. 2 to Fig. 4. Fig. 2 is a block diagram showing an example of the functional configuration of the audio processing device 20 according to the first embodiment. As shown in FIG. 2, the voice processing device 20 includes a communication unit 210, a storage unit 220, a first control unit 230, and a second control unit 240.

[0033] (1) Communications Unit 210 The communication unit 210 has a function of transmitting and receiving various information. The communication unit 210 is communicably connected to the sound collection device 10, the simultaneous interpretation engine 21, the machine translation engine 22, the monitoring terminal 30, and the display device 40, and transmits and receives various information.

[0034] (2) Storage section 220 The storage unit 220 has a function of storing various types of information. The storage unit 220 is configured by a storage medium provided as hardware in the audio processing device 20, such as a hard disk drive (HDD), a solid state drive (SSD), a flash memory, an electrically erasable programmable read-only memory (EEPROM), a random access read / write memory (RAM), a read-only memory (ROM), or any combination of these storage media.

[0035] The storage unit 220 stores, for example, conversion candidate information. The conversion candidate information is information that indicates conversion candidates for text data. The conversion candidate information accumulates text that is a conversion candidate based on the correction history by the checker. The conversion candidates are accumulated, for example, for each lecture. The conversion candidates may be deleted at the end of the lecture, or may remain accumulated. The conversion candidates that remain accumulated may be used for another lecture.

[0036] (3) First control unit 230 The first control unit 230 has a function of controlling processing related to simultaneous interpretation. The first control unit 230 is realized, for example, by causing a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) provided as hardware in the speech processing device 20 to execute a program. As shown in FIG. 2, the first control unit 230 includes a voice data acquisition unit 231 and a voice conversion unit 232.

[0037] (3-1) Audio data acquisition unit 231 The voice data acquisition unit 231 has a function of acquiring voice data. The voice data acquisition unit 231 acquires the voice data that the communication unit 210 receives from the sound collection device 10, and inputs the voice data to the voice conversion unit 232.

[0038] (3-2) Voice conversion unit 232 The speech conversion unit 232 has a function of executing speech conversion processing. The speech conversion unit 232 executes the speech conversion processing using the simultaneous interpretation engine 21. In the speech conversion processing, the speech conversion unit 232 inputs speech data indicating the speech content of the speaker in a first language, which is acquired by the speech data acquisition unit 231, to the simultaneous interpretation engine 21, and acquires text data output from the simultaneous interpretation engine 21 as the translation result. In this way, the speech conversion unit 232 can convert the speech data indicating the speech content of the speaker in the first language into text data indicating the translation result into multiple second languages ​​different from the first language by simultaneous interpretation. The speech conversion unit 232 inputs the text data in the second language obtained by the speech conversion process to the monitoring processing unit 241.

[0039] (4) Second control unit 240 The second control unit 240 has a function of controlling processes related to monitoring and subtitle display. The second control unit 240 is realized, for example, by causing a CPU or a GPU provided as hardware in the audio processing device 20 to execute a program. As shown in FIG. 2, the second control unit 240 includes a monitoring processing unit 241, a machine translation unit 242, and a subtitle processing unit 243.

[0040] (4-1) Monitoring Processing Unit 241 The monitoring processing unit 241 has a function of executing monitoring processing. In the monitoring processing, the monitoring processing unit 241 transmits screen information from the communication unit 210 to the monitoring terminal 30 and displays a monitoring screen on the monitoring terminal 30. The monitoring processing unit 241 receives, from the checker, a selection operation as to whether or not to display the text data, and an operation to correct the text data, via the monitoring screen displayed on the monitoring terminal 30.

[0041] The monitoring processing unit 241 displays on the monitoring screen the text data of one language that has been designated in advance as a target for monitoring, among the text data of the plurality of second languages ​​input from the speech conversion unit 232. For this reason, the monitoring processing unit 241 accepts corrections to the text data for the one second language designated among the plurality of second languages. The one language designated in advance as a target for monitoring is, for example, a language desired by the checker. In the first embodiment, as an example, it is assumed that the checker is Japanese and the language desired by the checker is Japanese.

[0042] The monitoring processing unit 241 displays on the monitoring screen a UI (User Interface) for accepting a selection operation as to whether or not the text data displayed on the monitoring screen should be displayed on the display device 40. The UI is, for example, a button, but may also be a check box, a pull-down menu, or the like. When the monitoring processing unit 241 receives a selection operation for whether to display, the monitoring processing unit 241 controls the display of subtitles of the text data that is the target of the operation. By the control of the monitoring processing unit 241, subtitles of text data for which "not possible" has been selected as the displayable / non-displayable item on the monitoring screen are not displayed on the display device 40, and subtitles of text data for which "possible" has been selected as the displayable / non-displayable item on the monitoring screen are displayed on the display device 40.

[0043] The monitoring processing unit 241 displays a UI for displaying text data on the monitoring screen. The UI is, for example, a text field. The monitoring processing unit 241 may display the text data in the text field on the monitoring screen and accept a selection of whether to display the text data through an operation on the text field. For example, the monitoring processing unit 241 switches whether to display the text data to disabled when the inside of the text field is selected, and switches whether to display the text data to enabled when the outside of the text field is selected after the inside of the text field is selected.

[0044] The monitoring processing unit 241 also accepts correction operations for the text data by operations on the text field in which the text data is displayed. For example, when the inside of a text field is selected by a checker, the monitoring processing unit 241 accepts corrections for the text data displayed in the selected text field. The checker can correct the text data by, for example, manual input.

[0045] When a portion of text data to be corrected is selected on the monitoring screen, the monitoring processing unit 241 may display conversion candidates near the portion to be corrected. The monitoring processing unit 241 inserts text selected by the checker from the conversion candidates into the portion to be corrected. In this way, the checker can correct text data not only by manual input but also by selecting conversion candidates.

[0046] The monitoring processing unit 241 compares the pre-correction text data with the corrected text data and adds, to the conversion candidates, any text that is not included in the conversion candidates and is detected as a difference. In this case, the monitoring processing unit 241 adds the new conversion candidate to the conversion candidate information stored in the storage unit 220. As a result, text that becomes a conversion candidate is accumulated in the conversion candidate information based on the correction history by the checker.

[0047] If correction of the text data displayed in the text field is not required and "possible" is selected for "displayable", the monitoring processing unit 241 inputs the text data in the multiple second languages ​​obtained by the voice conversion process by the voice conversion unit 232 to the subtitle processing unit 243. On the other hand, if correction of the text data displayed in the text field is required and "impossible" is selected for "displayable", the monitoring processing unit 241 inputs the corrected text data to the machine translation unit 242.

[0048] In addition, if "No" is selected as the display option for text data for which subtitles are already displayed on the display device 40, the monitoring processing unit 241 will hide the subtitles for the text data for which "No" was selected, and will be able to accept correction operations.

[0049] Here, the monitoring screen according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the monitoring screen according to the first embodiment.

[0050] The buttons B1, B2, and pull-down PD on the monitoring screen G1 shown in Fig. 3 are UIs for making settings related to the automatic selection of whether or not to display. The button B1 is a button for setting whether or not to display is automatically selected as possible (○). The button B2 is a button for setting whether or not to display is automatically selected as not possible (×). The pull-down PD is a pull-down for setting a predetermined time from when the text data is displayed on the monitoring screen G1 until whether or not to display is automatically selected. 3, as an example, button B1 is on, button B2 is off, and the predetermined time is set to three seconds. In this case, "Yes" is automatically selected as the display possibility three seconds after the text data is displayed on the monitoring screen. On the other hand, if button B1 is off and button B2 is on, "No" is automatically selected as the display possibility three seconds after the text data is displayed on the monitoring screen.

[0051] 3 are UIs for the checker to correct the displayed text data. The monitoring processing unit 241 displays, in the text fields, text data in one language that has been designated in advance as a target for monitoring, from among the text data in multiple second languages ​​input from the speech conversion unit 232. In the example shown in FIG. 3, four pieces of text data converted from four pieces of voice data are displayed in chronological order in text fields F1 to F4, respectively.

[0052] Buttons B3 and B4 on the monitoring screen G1 in FIG. 3 are UIs that allow the checker to select whether or not to display the text data. Button B3 is a button for selecting "not possible" (x) as whether or not to display. Button B4 is a button for selecting "possible" (o) as whether or not to display. Buttons B3 and B4 are displayed for each text field F. Note that the selection of buttons B3 and B4 may not only be made by the checker, but may also be switched under the control of the monitoring processing unit 241 after a predetermined time has elapsed, when the text data begins to be modified, or after the text data has been modified.

[0053] In the example shown in FIG. 3, for example, buttons B3-1 to B3-4 and buttons B4-1 to B4-4 are displayed in text fields F1 to F4, respectively. In the example of buttons B3-1 and B4-1 in text field F1, the checker determines that the text data displayed in text field F1 is correct and can be displayed, and button B4-1 is selected. In the example of buttons B3-2 and B4-2 in text field F2, the checker determines that the meaning of the text data displayed in text field F2 is incomprehensible, and button B3-2 is selected. In the example of buttons B3-3 and B4-3 in text field F3, the checker determines that the text data displayed in text field F3 is correct and can be displayed, and button B4-3 is selected. In the example of buttons B3-4 and B4-4 in text field F4, the checker determined that the text data displayed in text field F4 needed to be corrected and could not be displayed, and button F3-4 was selected, but after the text data was corrected, button B4-4 was selected under the control of the monitoring processing unit 241.

[0054] Here, the correction operation procedure according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the correction operation procedure according to the first embodiment. Fig. 4 shows an example in which a checker corrects text data displayed in text field F5. It should be noted that, between buttons B3-5 and B4-5, it is assumed that button B4-5 is selected as the initial selection.

[0055] As shown in Fig. 4, first, the checker checks the text data displayed in the text field F5, and selects (touches) any position in the text field F5 that needs to be corrected (step S1). In the example shown in Fig. 4, it is assumed that the correction location (position P) is selected.

[0056] After the checker selects text field F5, the monitoring processing unit 241 selects text field F5 (for example, displays a bold frame), switches the display / non-display selection from button B4-5 to button B3-5, displays cursor K inside text field F5, and displays window W1 showing conversion candidates near the correction point (step S2). If the position of cursor K is off from the correction point, the checker moves the position of cursor K to the correction point. If the subtitles of the text data displayed in text field F5 are already displayed on display device 40, the subtitles of the text data displayed in text field F5 will be hidden (deleted) when the correction part is selected by the checker.

[0057] The checker deletes the text to be corrected (step S3). In the example shown in Fig. 4, the checker deletes the text indicating "connection."

[0058] The checker selects the correct text from among the conversion candidates shown in window W1 (step S4). In the example shown in Fig. 4, the checker selects the text indicating "linkage."

[0059] After the checker selects the correct text, the monitoring processing unit 241 inserts the selected text into the correction portion in the text field F5 (step S5). If the correct text is not among the conversion candidates displayed in the window W1, the checker can manually input the correct text into the correction portion.

[0060] Since the correction is complete, the checker selects outside of text field F5 (step S6). As a result, the monitoring processing unit 241 returns text field F5 to the unselected state (cancels the bold frame display) and switches the display / non-display selection from button B3-5 to button B4-5. Furthermore, the monitoring processing unit 241 inputs the corrected text data to the machine translation unit 242.

[0061] (4-2) Machine Translation Department 242 The machine translation unit 242 has a function of executing machine translation processing. The machine translation unit 242 executes the machine translation processing using the machine translation engine 22. In the machine translation processing, the machine translation unit 242 inputs the text data corrected by the monitoring processing unit 241 to the machine translation engine 22, and acquires the text data output from the machine translation engine 22 as the translation result. As a result, the machine translation unit 242 can translate the corrected text data into a second language that was not specified at the time of correction, and acquire text data indicating the translation result in which the correction is reflected for each of the multiple second languages. The machine translation unit 242 inputs the corrected text data obtained by the monitoring process and the text data in the second language obtained by the machine translation process to the subtitle processing unit 243.

[0062] (4-3) Subtitle processing unit 243 The subtitle processing unit 243 has a function of executing subtitle display processing. In the subtitle display processing, the subtitle processing unit 243 displays subtitles on the display device 40 using text data in a second language converted by the speech conversion unit 232 and input from the monitoring processing unit 241, or text data in a second language input from the machine translation unit 242. When using text data in a plurality of second languages ​​input from the monitoring processing unit 241, the subtitle processing unit 243 can display, as subtitles, text data indicating translation results for which the monitoring process has determined that no correction is required on the display device 40. On the other hand, when using text data in a plurality of second languages ​​input from the machine translation unit 242, the subtitle processing unit 243 can display, as subtitles, text data indicating translation results for which the monitoring process has determined that correction is required and in which the correction is reflected on the display device 40.

[0063] <1-3. Processing flow> The functional configuration of the audio processing device 20 according to the first embodiment has been described above. Next, the processing flow in the subtitle display system 1 according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a sequence diagram showing an example of the processing flow in the subtitle display system 1 according to the first embodiment.

[0064] 5, first, the sound collection device 10 transmits sound data of collected sound to the sound processing device 20 (step S101). The sound data acquisition unit 231 of the sound processing device 20 acquires the sound data received by the communication unit 210 from the sound collection device 10.

[0065] Next, the voice conversion unit 232 of the voice processing device 20 requests the simultaneous interpretation engine 21 to perform simultaneous interpretation of the voice data acquired by the voice data acquisition unit 231 (step S102). The voice conversion unit 232 transmits the voice data to the simultaneous interpretation engine 21 via the communication unit 210.

[0066] Next, the simultaneous interpretation engine 21 simultaneously interprets (speech recognition and machine translation) the voice data received from the voice processing device 20, and transmits the result of the simultaneous interpretation to the voice processing device 20 (step S103). The result of the simultaneous interpretation is text data in a plurality of second languages.

[0067] Next, the monitoring processing unit 241 of the voice processing device 20 performs a process of displaying a monitoring screen (step S104). The monitoring processing unit 241 transmits screen information to the monitoring terminal 30 via the communication unit 210, and causes the monitoring terminal 30 to display the monitoring screen.

[0068] The monitoring terminal 30 displays a monitoring screen based on the screen information received from the voice processing device 20 (step S105). After displaying the monitoring screen, the monitoring terminal 30 accepts corrections to the text data made by the checker on the monitoring screen, and transmits correction information indicating the correction content to the voice processing device 20 (step S106).

[0069] The monitoring processing unit 241 determines whether or not there is any correction to the text data depending on whether or not the communication unit 210 receives correction information from the monitoring terminal 30 (step S107). If there is any correction (step S107 / YES), the process proceeds to step S108. On the other hand, if there is no correction (step S107 / NO), the process proceeds to step S110.

[0070] When the process proceeds to step S108, the machine translation unit 242 of the speech processing device 20 requests the machine translation engine 22 to perform machine translation of the text data corrected by the checker (step S108). The machine translation unit 242 transmits the text data to the machine translation engine 22 via the communication unit 210.

[0071] Next, the machine translation engine 22 performs machine translation on the text data received from the speech processing device 20 and transmits the results of the machine translation to the speech processing device 20 (step S109). The results of the machine translation are text data in a plurality of second languages. After transmission, the process proceeds to step S111.

[0072] If the process proceeds to step S110, the monitoring processing unit 241 determines whether the text data can be displayed (step S110). If the displayability is determined to be displayable (step S110 / YES), the process proceeds to step S111. On the other hand, if the displayability is determined to be undisplayable (step S110 / NO), the process ends.

[0073] When the process proceeds to step S111, the subtitle processing unit 243 of the audio processing device 20 executes a subtitle display process (step S111). The subtitle processing unit 243 transmits text data in a plurality of second languages ​​to the display device 40 via the communication unit 210. The display device 40 displays the text data in the second language received from the audio processing device 20 as subtitles (step S112).

[0074] The processing flow according to the first embodiment has been described above. As described above, the voice processing device 20 according to the first embodiment includes a voice conversion unit 232 that converts voice data indicating the content of a user's utterance in a first language into text data indicating the translation results into a plurality of second languages ​​different from the first language through simultaneous interpretation, a monitoring processing unit 241 that accepts corrections to the text data for one specified second language from among the plurality of second languages, and a machine translation unit 242 that translates the corrected text data into a second language that was not specified at the time of correction and acquires text data indicating the translation results in which the corrections are reflected for each of the plurality of second languages.

[0075] With this configuration, in simultaneous interpretation of multiple languages, if the simultaneous interpretation result contains a misrecognition or mistranslation, the checker can correct the simultaneous interpretation result of only one of the multiple languages ​​of the translated word, and can also correct the simultaneous interpretation results of the other languages. Therefore, the speech processing device 20 according to the first embodiment makes it possible to easily correct the results of simultaneous interpretation in simultaneous interpretation of multiple languages.

[0076] Furthermore, if the simultaneous interpretation of the speaker's speech contains a misrecognition or mistranslation, the checker can correct the misrecognition or mistranslation before the simultaneous interpretation result (subtitles) is presented to the audience. This ensures that the audience will not be presented with the simultaneous interpretation result that contains the misrecognition or mistranslation, but only with the corrected simultaneous interpretation result that does not contain the misrecognition or mistranslation. Therefore, the speech processing device 20 according to the first embodiment enables the user to correctly understand the content of the simultaneous translation.

[0077] <<2. Second Embodiment>> The first embodiment has been described above. Next, a second embodiment will be described with reference to Figs. 6 to 12. In the second embodiment, a method of correcting text data on a monitoring screen that is different from that in the first embodiment will be described. In the following, descriptions that overlap with those in the first embodiment will be omitted as appropriate. In the second embodiment, the language spoken by the speaker (first language) is any one language (for example, Japanese), and the language of the multilingually translated subtitles (second language) is any multiple languages ​​(for example, languages ​​other than Japanese).

[0078] <2-1. Configuration of the subtitle display system> The configuration of a subtitle display system 1a according to the second embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the configuration of a subtitle display system 1a according to the second embodiment. As shown in FIG. 6, the subtitle display system 1a includes a sound collection device 10, a sound processing device 20a, a machine translation engine 22, a voice recognition engine 23, a conversion candidate API (Application Programming Interface) 24, a monitoring terminal 30, and a display device 40.

[0079] (1) Sound collection device 10 The sound collector 10 according to the second embodiment is similar to the sound collector 10 according to the first embodiment, and therefore a description thereof will be omitted.

[0080] (2) Audio processing device 20a The audio processing device 20a according to the second embodiment differs from the audio processing device 20 according to the first embodiment in the process it performs due to a different method of correction on the monitoring screen. Based on the audio data received from the sound collection device 10, the audio processing device 20a performs not only audio conversion processing, monitoring processing, machine translation processing, and subtitle display processing, but also morphological analysis processing, conversion priority processing, and an automatic conversion processing unit.

[0081] (3) Machine Translation Engine 22 The machine translation engine 22 according to the second embodiment is similar to the machine translation engine 22 according to the first embodiment, and therefore a description thereof will be omitted.

[0082] (4) Voice Recognition Engine 23 The voice recognition engine 23 is an engine (program) that performs voice recognition on voice data. The voice recognition engine 23 converts the voice data in the first language input from the voice processing device 20a into text data in the first language by voice recognition. The function of the voice recognition engine 23 may be provided by a device or terminal different from the voice processing device 20a, or may be provided by the voice processing device 20a.

[0083] (5) Conversion candidate API24 The conversion candidate API 24 is an API that acquires conversion candidates corresponding to the reading of text data. The conversion candidate API 24 acquires conversion candidates corresponding to the reading of text data in the first language based on information indicating the reading of text data in the first language input from the speech processing device 20a. The function of the conversion candidate API 24 may be provided by a device or terminal different from the voice processing device 20a, or may be provided by the voice processing device 20a.

[0084] (6) Monitoring terminal 30 The monitoring terminal 30 according to the second embodiment is similar to the monitoring terminal 30 according to the first embodiment, and therefore a description thereof will be omitted.

[0085] (7) Display device 40 The display device 40 according to the second embodiment is similar to the display device 40 according to the first embodiment, and therefore a description thereof will be omitted.

[0086] <2-2. Functional configuration of voice processing device> The configuration of the subtitle display system 1a according to the second embodiment has been described above. Next, the functional configuration of the audio processing device 20a according to the second embodiment will be described with reference to Fig. 7 to Fig. 10. Fig. 7 is a block diagram showing an example of the functional configuration of the audio processing device 20a according to the second embodiment. As shown in FIG. 7, the voice processing device 20a includes a communication unit 210a, a storage unit 220a, a first control unit 230a, and a second control unit 240a.

[0087] (1) Communication unit 210a The communication unit 210a is communicably connected to the sound collection device 10, the machine translation engine 22, the voice recognition engine 23, the conversion candidate API 24, the monitoring terminal 30, and the display device 40, and transmits and receives various information.

[0088] (2) Storage unit 220a The storage unit 220a differs from the storage unit 220 according to the first embodiment in that it also stores priority information. The priority information indicates the priority of automatic conversion for each morpheme into which text data is divided by morphological analysis. The priority is calculated based on the past correction history during monitoring by the checker.

[0089] (3) First control unit 230a As shown in FIG. 7, the first control unit 230a includes a voice data acquisition unit 231a and a voice conversion unit 232a.

[0090] (3-1) Voice data acquisition unit 231a The voice data acquisition unit 231a according to the second embodiment is similar to the voice data acquisition unit 231 according to the first embodiment, and therefore a description thereof will be omitted.

[0091] (3-2) Voice conversion unit 232a The speech conversion unit 232a differs from the speech conversion unit 232 according to the first embodiment in that it only performs speech recognition rather than simultaneous interpretation. The voice conversion unit 232a executes voice recognition processing as the voice conversion processing using the voice recognition engine 23. In the voice recognition processing, the voice conversion unit 232a inputs voice data indicating the contents of the speaker's utterance in the first language, which is acquired by the voice data acquisition unit 231a, to the voice recognition engine 23, and acquires text data output from the voice recognition engine 23 as a voice recognition result. In this way, the voice conversion unit 232a can convert the voice data indicating the contents of the user's utterance in the first language into text data indicating the contents of the user's utterance in the first language by voice recognition. The speech conversion unit 232a inputs the text data in the first language obtained by the speech recognition process to the monitoring processing unit 241a.

[0092] (4) Second control unit 240a As shown in FIG. 7, the second control unit 240a includes a monitoring processing unit 241a, a machine translation unit 242a, a subtitle processing unit 243a, a morphological analysis unit 244, a conversion priority processing unit 245, and an automatic conversion processing unit 246.

[0093] (4-1) Monitoring processing unit 241a During the monitoring process, the monitoring processing unit 241a accepts a correction operation for the text data from the checker via the monitoring screen displayed on the monitoring terminal 30, but does not accept a selection operation for whether or not to display the text data. Furthermore, the monitoring processing unit 241a displays the text data on the monitoring screen by dividing it into multiple texts, and accepts corrections to the text data in units of divided texts. For this reason, the first embodiment and the second embodiment differ in the method of correcting text data on the monitoring screen.

[0094] The monitoring processing unit 241a displays a UI for displaying text data on the monitoring screen. The UI is, for example, a text field. The monitoring processing unit 241a displays multiple text fields for one piece of text data on the monitoring screen. The monitoring processing unit 241a divides the text data into text in morpheme units based on the result of morpheme analysis by the morpheme analysis unit 244, which will be described later. The morpheme units are, for example, parts of speech units. In this case, the monitoring processing unit 241a divides the text data into parts of speech units based on information (hereinafter also referred to as "part of speech information") indicating the result of dividing the text data into parts of speech units by the morpheme analysis unit 244. The part of speech information includes, for example, the divided text and information indicating the reading of each text. After the division, the monitoring processing unit 241a displays the divided texts in their respective text fields. The monitoring processing unit 241a accepts correction operations for each of the multiple text fields for one piece of text data. This allows the checker to correct the text data on a morpheme-by-morpheme basis.

[0095] When text to be corrected as text data is selected on the monitoring screen, the monitoring processing unit 241a displays conversion candidates near the selected text. The monitoring processing unit 241a acquires conversion candidates using the conversion candidate API 24. The monitoring processing unit 241a inputs part-of-speech information acquired by the morphological analysis unit 244 to the conversion candidate API 24, and acquires information indicating the conversion candidates output from the conversion candidate API 24 (hereinafter also referred to as "conversion candidate information"). The conversion candidate API 24 refers to the reading of the text included in the part-of-speech information, and acquires homonyms as conversion candidates. This allows the monitoring processing unit 241a to display conversion candidates on the monitoring screen based on the acquired conversion candidate information. When a conversion candidate displayed on the monitoring screen is selected, the monitoring processing unit 241a replaces the text to be corrected with the text selected from the conversion candidates. In this way, the checker can correct the text data on a morpheme-by-morpheme basis by selecting a conversion candidate.

[0096] The monitoring processing unit 241a displays a text input field along with the display of the conversion candidates, and accepts corrections by inputting text into the input field. This allows the checker to appropriately correct the text data by inputting appropriate text into the input field, for example, when there is no appropriate conversion candidate among the conversion candidates.

[0097] When displaying text data on the monitoring screen, the monitoring processing unit 241a may convert the text data based on past correction records before displaying it. For example, suppose that the text data to be displayed contains text corresponding to text previously corrected in other text data. The corresponding text is, for example, a homonym. In this case, the monitoring processing unit 241a converts the corresponding text in the text data to be displayed into the previously corrected text before displaying it. That is, when displaying second text data obtained by speech recognition performed after corrections are made to the first text data, if the second text data contains second text corresponding to the first text corrected in the first text data, the monitoring processing unit 241a converts the second text into the corrected first text and displays it. As a result, text that has been corrected in the past is automatically converted by the monitoring processing unit 241a, which reduces the load on the checker in monitoring.

[0098] The monitoring processing unit 241a may display not only the text data converted from the voice data, but also the text data obtained by translating the text data into a different language on the monitoring screen. When the text data converted from the voice data is corrected on the monitoring screen, the monitoring processing unit 241a acquires text data indicating the translation result reflecting the correction using the machine translation engine 22, and updates the display of the translation with the acquired text data. The translation displayed on the monitoring screen by the monitoring processor 241a may be a translation of one of the plurality of second languages, or may be a translation of a plurality of languages.

[0099] Here, a monitoring screen according to the second embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of a monitoring screen according to the second embodiment.

[0100] As an example, the monitoring screen G2 shown in FIG. 8 displays text data obtained by executing a speech recognition process for two pieces of audio data, and the translation results obtained by machine translation for each piece of text data.

[0101] The first voice data is a voice indicating the spoken content, for example, "The company will assign a conductor." The voice data is converted into text data, "The company will assign a conductor," by speech recognition processing by the voice conversion unit 232a. The monitoring processing unit 241a divides the text data into parts of speech based on the results of morphological analysis and then displays them in the text field F11. In the example shown in FIG. 8, the text data is divided into seven parts of speech (words), "company," "at," "is," "conductor," "to," "attach," and "masu," and is displayed in the text field F11. Furthermore, the monitoring processing unit 241a displays text data indicating the result of the text data being translated by the machine translation unit 242a in the text field F12. In the example shown in Fig. 8, the text data is translated into "There is a conductor in the company." and displayed in the text field F12. The text data displayed in the text field F11 is text data obtained by speech recognition of speech data indicating the content of a speech by a speaker between certain times T1 and T2, and is the first text data. Time T2 is assumed to be a time later than time T1. The text data displayed in the text field F12 is text data showing the translation result into a predetermined arbitrary language among a plurality of languages ​​(a plurality of second languages) that are the target of multilingual translation. In the second embodiment, as an example, it is assumed that the translation result from Japanese to English is set to be displayed.

[0102] The second voice data is a voice indicating the utterance content, for example, "We wear our company emblem mainly in important expressions." The voice data is converted into text data, "We wear our company emblem mainly in important expressions," by speech recognition processing by the voice conversion unit 232a. The monitoring processing unit 241a divides the text data into parts of speech based on the results of morphological analysis and then displays them in the text field F13. In the example shown in FIG. 8, the text data is divided into nine parts of speech (words): "company emblem," "is," "mainly," "important," "na," "expression," "de," "wear," and "masu," and is displayed in the text field F13. Furthermore, the monitoring processing unit 241a displays text data indicating the result of translation of the text data by the machine translation unit 242a in the text field F14. In the example shown in Fig. 8, the text data is translated into "We put the company emblem mainly at important ceremonies." and displayed in the text field F14. The text data displayed in the text field F13 is text data obtained by speech recognition of speech data indicating the content of a speech by a speaker between certain times T3 and T4, and is the second text data. Time T3 is assumed to be later than time T2, and time T4 is assumed to be later than time T3.

[0103] Here, a correction operation procedure according to the second embodiment will be described with reference to Fig. 9. Fig. 9 is a diagram showing an example of a correction operation procedure according to the second embodiment. Fig. 9 shows an example of correcting text data in the text field F11 in Fig. 8.

[0104] 9, first, the checker checks the text data displayed in the text field F11 and selects (touches) the text that needs to be corrected (step S11). In the example shown in FIG. 9, it is assumed that "conductor" has been selected as the text to be corrected.

[0105] After the checker selects text, the monitoring processing unit 241a displays a conversion candidate window W11, a text field F15, an add button B5, and a delete button B6 near the selected text (step S12). The checker selects appropriate text for correction from the conversion candidates displayed in the conversion candidate window W11. If there is no appropriate text among the conversion candidates, the checker enters appropriate text in the text field F15 (input field) and presses the add button B5. If the checker wants to delete the text selected in the text field F11, the checker presses the delete button B6. If the text selected by the checker from text field F11 is automatically converted, the text before the automatic conversion is displayed in text field F15. In this case, the checker can return the text selected in text field F11 to the text before the automatic conversion by pressing the Add button B5.

[0106] After the checker makes the correction, the monitoring processing unit 241 displays the text selected in the text field F11 as the corrected text (step S13). Furthermore, the monitoring processing unit 241a displays, in the text field F12, the text data obtained by translating the corrected text data shown in the text field F11. In the example shown in Figure 9, "conductor" has been corrected to "company emblem." As a result, in the example shown in Figure 8, the text recognized as "conductor" in the speech recognition of the second voice data is automatically converted to "company emblem" and then displayed in text field F13.

[0107] (4-2) Machine Translation Unit 242a The machine translation unit 242a translates the text data converted from the voice data in the first language by the voice recognition process performed by the voice conversion unit 232a into a second language different from the first language, and obtains text data indicating the translation result. This allows the machine translation unit 242a to obtain text data indicating the translation result, which is displayed on the monitoring screen together with the text data converted from the voice data. Furthermore, when text data displayed on the monitoring screen is corrected, the machine translation unit 242a translates the corrected text data and acquires text data indicating the translation result reflecting the corrections. This allows the checker to check the corrected translation on the monitoring screen, and also allows the machine translation unit 242a to acquire text data reflecting the corrections that can be displayed as is on the display device 40 as subtitles.

[0108] (4-3) Subtitle processing unit 243a The subtitle processing unit 243a according to the second embodiment is similar to the subtitle processing unit 243 according to the first embodiment, and therefore a description thereof will be omitted.

[0109] (4-4) Morphological analysis section 244 The morphological analysis unit 244 has a function of performing morphological analysis. The morphological analysis unit 244 performs morphological analysis on text data converted from voice data. The morphological analysis unit 244 divides the text data into multiple texts, for example, by part of speech.

[0110] (4-5) Conversion priority processing unit 245 The conversion priority processing unit 245 has a function of performing processing related to priority information. Priority information is information indicating the priority of text used for conversion in automatic text conversion. The priority is determined based on the history of past corrections. When text data displayed on the monitoring screen is corrected, the conversion priority processing unit 245 changes the priority (weight) according to the content of the correction.

[0111] Here, a change in conversion priority according to the second embodiment will be described with reference to Fig. 10. Fig. 10 is a diagram showing an example of a change in conversion priority according to the second embodiment. The table on the left side of Fig. 10 shows an example of the priority before the change, and the table on the right side shows an example of the priority after the change.

[0112] 10 shows the pre-conversion text, the post-conversion text, and a priority (weight) as priority information. The priority information indicates that when pre-conversion text is included in the speech-recognized text data, it is converted into post-conversion text with a higher priority. The table on the left shows that the priority for converting "park" to "lecture" is "1.05," the priority for converting "park" to "performance" is "1.00," the priority for converting "park" to "oral speech" is "0.95," and the priority for converting "lecture" to "performance" is "1.00." Multiple priority information is shown for "park." If "park" is included in the speech-recognized text data, it will be automatically converted to "lecture," which has the highest priority. When the checker corrects the speech-recognized text data, the conversion priority processing unit 245 changes the priority information. For example, suppose the speech-recognized text data includes "lecture" and the checker corrects "lecture" to "performance." In this case, the conversion priority processing unit 245 changes the table shown on the left side of FIG. 10 to the table shown on the right side. Because "lecture" has been corrected to "performance," the conversion priority processing unit 245 increases the priority of conversion to "performance" and decreases the priority of conversion to anything other than "performance."

[0113] (4-6) Automatic conversion processing unit 246 The automatic conversion processing unit 246 has a function of automatically converting text. Before displaying text data divided into multiple texts, the automatic conversion processing unit 246 refers to priority information. When multiple texts include text to be automatically converted, the automatic conversion processing unit 246 selects text according to priority from text used for past corrections, and converts the target text with the selected text. The text used for past corrections is the corrected text when the checker previously corrected text in the text data, and is the converted text in the priority information shown in FIG. 10. The text selected by the automatic conversion processing unit 246 according to the priority is, for example, the text with the highest priority. When only one piece of priority information exists for one text, the automatic conversion processing unit 246 selects the text indicated by the priority information and performs automatic conversion. When multiple pieces of priority information exist for one text, the automatic conversion processing unit 246 selects the text indicated by the priority information with the highest priority and performs automatic conversion.

[0114] <2-3. Processing flow> The functional configuration of the voice processing device 20a according to the second embodiment has been described above. Next, the flow of processing according to the second embodiment will be described with reference to Figs.

[0115] (1) Processing flow in the subtitle display system 1a The flow of processing in the subtitle display system 1a according to the second embodiment will be described with reference to Fig. 11. Fig. 11 is a sequence diagram showing an example of the flow of processing in the subtitle display system 1a according to the second embodiment.

[0116] 11, first, the sound collection device 10 transmits sound data of collected sound to the sound processing device 20a (step S201). The sound data acquisition unit 231a of the sound processing device 20a acquires the sound data received by the communication unit 210a from the sound collection device 10.

[0117] Next, the voice conversion unit 232a of the voice processing device 20a performs voice conversion processing (step S202). The voice conversion unit 232a transmits the voice data acquired by the voice data acquisition unit 231a to the voice recognition engine 23 via the communication unit 210a and requests voice recognition. The voice recognition engine 23 performs voice recognition on the voice data received from the voice processing device 20a and transmits the result of the voice recognition to the voice processing device 20a.

[0118] Next, the machine translation unit 242a of the speech processing device 20a performs machine translation processing (step S203). The machine translation unit 242a transmits the text data acquired by the speech conversion unit 232a to the machine translation engine 22 via the communication unit 210a and requests machine translation. The machine translation engine 22 performs machine translation of the text data received from the speech processing device 20a and transmits the result of the machine translation to the speech processing device 20a.

[0119] Next, the audio processing device 20a performs a display preparation process (step S204). The display preparation process is a preparation process for displaying a monitoring screen. The display preparation process will be described in detail later.

[0120] Next, the monitoring processing unit 241a of the voice processing device 20a performs a process of displaying a monitoring screen (step S205). The monitoring processing unit 241a transmits screen information to the monitoring terminal 30 via the communication unit 210a, and causes the monitoring terminal 30 to display the monitoring screen.

[0121] The monitoring terminal 30 displays a monitoring screen based on the screen information received from the voice processing device 20a (step S206). After displaying the monitoring screen, the monitoring terminal 30 accepts corrections to the text data made by the checker on the monitoring screen, and transmits correction information indicating the correction content to the voice processing device 20a (step S207).

[0122] The monitoring processing unit 241a determines whether or not there is any correction to the text data depending on whether or not the communication unit 210a receives correction information from the monitoring terminal 30 (step S208). If there is any correction (step S208 / YES), the process proceeds to step S209. On the other hand, if there is no correction (step S208 / NO), the process proceeds to step S212.

[0123] If the process proceeds to step S209, the machine translation unit 242a of the speech processing device 20a performs machine translation processing (step S209). The machine translation unit 242a transmits the text data corrected by the checker to the machine translation engine 22 via the communication unit 210a and requests machine translation. The machine translation engine 22 performs machine translation of the text data received from the speech processing device 20a and transmits the results of the machine translation to the speech processing device 20a.

[0124] Next, the monitoring processing unit 241a performs an update process of the monitoring screen (step S210). The monitoring processing unit 241a updates the translation display on the monitoring screen based on the result of the machine translation acquired by the machine translation unit 242a.

[0125] Next, the conversion priority processing unit 245 of the voice processing device 20a updates the priority (step S211). Based on the correction information received by the communication unit 210a from the monitoring terminal 30, the conversion priority processing unit 245 updates the priority of the priority information corresponding to the correction information.

[0126] When the process proceeds to step S212, the subtitle processing unit 243a of the audio processing device 20a executes a subtitle display process (step S212). The subtitle processing unit 243a transmits text data in a plurality of second languages ​​to the display device 40 via the communication unit 210a. The display device 40 displays the text data in the second language received from the audio processing device 20a as subtitles (step S213).

[0127] (2) Display preparation process flow The flow of the display preparation process according to the second embodiment will be described with reference to Fig. 12. Fig. 12 is a sequence diagram showing an example of the flow of the display preparation process according to the second embodiment.

[0128] As shown in FIG. 12, first, the morphological analysis unit 244 of the speech processing device 20a performs morphological analysis on the text data acquired by the speech conversion unit 232a (step S301).

[0129] Next, the automatic conversion processing unit 246 of the voice processing device 20a acquires the priority information stored in the storage unit 220a (step S302).

[0130] The automatic conversion processing unit 246 refers to the priority information and checks whether or not there is a part of speech (text) with a high priority for automatic conversion in the morphologically analyzed text data (step S303). If there is a part of speech with a high priority for automatic conversion (step S303 / YES), the process proceeds to step S304. On the other hand, if there is no part of speech with a high priority for automatic conversion (step S303 / NO), the process proceeds to step S305.

[0131] If the process proceeds to step S304, the automatic conversion processing unit 246 performs automatic conversion (step S304). After the automatic conversion, the process proceeds to step S305.

[0132] When the process proceeds to step S305, the monitoring processing unit 241a of the speech processing device 20a transmits the part-of-speech information acquired by the morphological analysis unit 244 to the conversion candidate API 24 via the communication unit 210a (step S305).

[0133] The conversion candidate API 24 acquires conversion candidate information based on the part-of-speech information received from the speech processing device 20a, and transmits the information to the speech processing device 20a (step S306).

[0134] The processes from step S302 to step S306 are performed for each part of speech (text unit) based on the results of the morphological analysis in step S301. Therefore, the processes from step S302 to step S306 are repeated until conversion candidate information for all parts of speech is acquired. When conversion candidate information for all parts of speech has been acquired, the display preparation process ends.

[0135] The processing flow according to the second embodiment has been described above. As described above, the voice processing device 20a according to the second embodiment includes a voice conversion unit 232a that converts voice data indicating the content of a user's utterance into text data by voice recognition, and a monitoring processing unit 241a that divides the text data into multiple texts and displays them, and accepts corrections to the text data in units of divided texts. When displaying second text data obtained by voice recognition performed after corrections are made to the first text data, if the second text data contains second text corresponding to the first text corrected in the first text data, the monitoring processing unit 241a converts the second text into the corrected first text and displays it.

[0136] With this configuration, when the checker corrects a recognition error that occurred in the speech recognition of certain speech data, even if the same recognition error occurs in the speech recognition of other speech data after the correction, it will be automatically converted into correct text. This means that the checker does not need to make the same correction every time the same recognition error occurs. Therefore, the voice processing device 20a according to the second embodiment can reduce the load of correcting erroneous voice recognition.

[0137] Furthermore, if the simultaneous interpretation of the speaker's speech contains a misrecognition or mistranslation, the checker can correct the misrecognition or mistranslation before the simultaneous interpretation result (subtitles) is presented to the audience. This ensures that the audience will not be presented with the simultaneous interpretation result that contains the misrecognition or mistranslation, but only with the corrected simultaneous interpretation result that does not contain the misrecognition or mistranslation. Therefore, the speech processing device 20a according to the second embodiment enables the user to correctly understand the content of the simultaneous translation.

[0138] <<3. Modifications>> The above describes the embodiments. Next, modifications of the above-described embodiments will be described. The modifications described below may be applied to the embodiments alone or in combination with each other. Furthermore, the modifications may be applied in place of the configurations described in the embodiments, or may be applied in addition to the configurations described in the embodiments.

[0139] The first and second embodiments described above may be implemented in combination. For example, the monitoring screen of the first embodiment may display a translation in the same manner as the monitoring screen of the second embodiment. Furthermore, the monitoring screen of the second embodiment may display a UI that can accept a selection operation from the checker as to whether or not to display text data, similar to the monitoring screen of the first embodiment.

[0140] In the first embodiment described above, an example has been described in which the first language is a language other than Japanese and the second language is a plurality of languages ​​including Japanese, but the present invention is not limited to such an example. For example, the first language may be Japanese and the second language may be a language other than Japanese. In the second embodiment described above, an example has been described in which the first language is Japanese and the second language is a plurality of languages ​​other than Japanese, but the present invention is not limited to such an example. For example, the first language may be a language other than Japanese, and the second language may be a plurality of languages ​​including Japanese. In addition, in each of the above-described embodiments, the display device 40 may be capable of displaying subtitles in a first language.

[0141] The above describes the modified examples of the embodiment. Note that some or all of the functions of the subtitle display system 1, 1a and the audio processing device 20, 20a in the above-described embodiments may be implemented by a computer. In this case, a program for implementing the functions may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into and executed by a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices. Additionally, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, CD-ROMs, etc., and storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and devices that store programs for a certain period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in a computer system, or may be one that is realized using a programmable logic device such as an FPGA (Field Programmable Gate Array).

[0142] The embodiments of the present invention have been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes can be made within the scope of the gist of the present invention. [Explanation of symbols]

[0143] 1,1a...subtitle display system, 10...sound collection device, 20,20a...speech processing device, 21...simultaneous interpretation engine, 22...machine translation engine, 23...speech recognition engine, 24...conversion candidate API, 30...monitoring terminal, 40...display device, 41...screen, 42...smartphone, 210,210a...communication unit, 220,220a...storage unit, 230,230a...first control unit, 231,231a...speech data acquisition unit, 232,232a...speech conversion unit, 240,240a...second control unit, 241,241a...monitoring processing unit, 242,242a...machine translation unit, 243,243a...subtitle processing unit, 244...morphological analysis unit, 245...conversion priority processing unit, 246...automatic conversion processing unit

Claims

1. a speech conversion unit that converts speech data representing the contents of a user's speech in a first language into text data representing translation results into a plurality of second languages ​​different from the first language by simultaneous interpretation; a monitoring processing unit that accepts corrections to the text data for one designated second language among the plurality of second languages; a translation unit that translates the corrected text data into the second language that was not specified at the time of correction, and obtains text data indicating translation results in which the correction is reflected for each of the second languages; An audio processing device comprising:

2. a subtitle processing unit that displays the text data indicating the translation result reflecting the correction as subtitles on a display device; The audio processing device of claim 1 further comprising:

3. the monitoring processing unit displays, on the monitoring terminal, a monitoring screen capable of accepting a correction operation for the text data converted from the voice data and a selection operation for whether or not to display the text data. The audio processing device according to claim 1 .

4. Subtitles of the text data for which "displayable" is selected as "unavailable" on the monitoring screen are not displayed on the display device, and subtitles of the text data for which "displayable" is selected as "available" on the monitoring screen are displayed on the display device. The audio processing device according to claim 3 .

5. the monitoring processing unit displays the text data in a text field on the monitoring screen, and when an inside of the text field is selected, switches the display enable / disable state to disabled, and when an outside of the text field is selected after the inside of the text field is selected, switches the display enable / disable state to enabled. The audio processing device according to claim 3 .

6. When a portion of the text data to be corrected is selected on the monitoring screen, the monitoring processing unit displays conversion candidates near the portion of the text data to be corrected, and inserts text selected from the conversion candidates into the portion of the text data to be corrected. The audio processing device according to claim 3 .

7. the monitoring processing unit compares the text data before correction with the text data after correction and detects a difference between the text data, and adds text that is not included in the conversion candidates to the conversion candidates. The audio processing device according to claim 6 .

8. a speech conversion process for converting speech data representing the contents of a user's speech in a first language into text data representing translation results into a plurality of second languages ​​different from the first language by simultaneous interpretation; a monitoring process for receiving corrections to the text data for one designated second language among the plurality of second languages; a translation process of translating the corrected text data into the second language that was not specified at the time of correction, and obtaining text data indicating translation results in which the correction is reflected for each of the second languages; 1. A computer-implemented method for audio processing, comprising:

9. Computer, a speech conversion means for converting speech data representing the contents of a user's speech in a first language into text data representing the results of translation into a plurality of second languages ​​different from the first language by simultaneous interpretation; a monitoring processing means for receiving corrections to the text data for one designated second language among the plurality of second languages; a translation means for translating the corrected text data into the second language that was not specified at the time of correction, and obtaining text data indicating the translation result in which the correction is reflected for each of the plurality of second languages; A program to function as a

Citation Information

Patent Citations

  • Translation apparatus

    JP2009122989A

  • Content participation translation apparatus and content participation translation method using the same

    JP2016100023A

  • Multilanguage voice translation system for TV conference system

    JP2017191959A

  • Systems, methods, and apparatus for determining an official transcription and speaker language from a plurality of transcripts of text in different languages

    US20220414349A1

  • Text correction device, text correction method and text correction program

    JP2019148681A