Subtitle creation apparatus
By designing a subtitle generation device that includes character information acquisition, display and generation units, the problem of cumbersome proofreading tasks still exist in the prior art when generating subtitle in the art is solved, and efficient proofreading and subtitle generation of subtitle generation devices are realized.
Patent Information
- Application Number
- JP2023181050
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2025-05-02
AI Technical Summary
In the generation of subtitles, although the burden of modifying subtitles to make them more readable is reduced, there are still a large number of cumbersome proofreading tasks, resulting in the subtitle generation device still poses a significant work burden to the proofreading personnel when generating subtitles from speech recognition results in real time.
A subtitle generation device is designed, which includes a character information acquisition unit, a proofreading information display unit, and a character information generation unit. The device obtains the speech recognition results and divides them into fixed-sized blocks, displays them to the user for proofreading, and modifies the character information according to the user's proofreading instructions, and finally generates the corrected character information.
Through the use of this device, the work burden of proofreading personnel during the subtitle generation process is significantly reduced, and the efficiency and quality of subtitle generation are improved.
Smart Images

Figure 2025070599000001_ABST
Abstract
Description
[Technical field]
[0001] This relates to technology for a device that generates subtitles to be superimposed on broadcast video and transfers them to a subtitle transmission system (a system that broadcasts broadcast video generated by synthesizing text information with original broadcast video). [Background technology]
[0002] There are systems for creating subtitles to be superimposed on broadcast images. In recent years, computer-based natural language processing technology has continued to evolve, and there are cases where broadcast audio is converted into text information through computer-based speech recognition processing, and the text information is used as subtitles.
[0003] Against this background, for example, Patent Document 1 proposes a technique for a caption generation device that reduces the burden of correcting captions to make them easier to read when generating captions in real time from speech recognition results. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2023-067353 A Summary of the Invention [Problem to be solved by the invention]
[0005] However, while the above-mentioned conventional technology is said to reduce the burden of correcting subtitles to make them easier to read when generating subtitles in real time from speech recognition results, there is still a problem in that there is still a lot of cumbersome work left for proofreaders.
[0006] In view of the above problems, the present invention has an object to provide a subtitling device that reduces the workload of a person who proofreads subtitles for broadcasting when the subtitles are produced from the results of computer-generated speech recognition. [Means for solving the problem]
[0007] One form of the disclosed subtitle production device comprises a character information acquisition means for acquiring the results of voice recognition of audio data related to broadcast audio as time-series character information, a proofreading information display means for displaying the character information on a proofreading display device in chunks of a predetermined size for proofreading by a user, and a character information generation means for correcting the character information in accordance with a proofreading instruction from the user and generating corrected character information, wherein when the proofreading information display means accepts a selection by the user for a predetermined chunk of the character information, which is time-series data divided into chunks of the predetermined size, to be displayed on the proofreading display device, the character information prior to the predetermined chunk is excluded from the objects to be displayed on the proofreading display device. Effect of the Invention
[0008] The disclosed subtitling device reduces the workload of a person who proofreads subtitles for broadcast when the subtitles are produced from the results of computer-generated speech recognition. [Brief description of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an overview of a subtitling device according to an embodiment of the present invention. [Diagram 2] 1 is a functional block diagram of a subtitling device according to an embodiment of the present invention. [Diagram 3] 10A and 10B are diagrams for explaining the processing of the calibration information display device according to the present embodiment. [Figure 4] 10A and 10B are diagrams for explaining the processing of the calibration information display device according to the present embodiment. [Diagram 5] 1 is a diagram illustrating an example of a hardware configuration of a subtitling device according to an embodiment of the present invention. [Figure 6] 10 is a flowchart showing a flow of an example of processing by the subtitling device according to the present embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS With reference to the drawings, an embodiment of the present invention will be described. (Operation principle of the subtitling device according to the present embodiment)
[0011] The operation principle of a subtitling device 100 according to this embodiment (hereinafter simply referred to as "this device") will be described with reference to Figures 1 to 4. Figure 1 is a diagram showing the connection relationship between this device 100 and external devices, and Figure 2 is a functional block diagram of this device 100. Figures 3 and 4 are diagrams for explaining the processing by the proofreading information display means 140.
[0012] 1, the present device 100 is connected to an external device 340, a voice recognition server 290, and a subtitle sending system 320. The external device 340 is a device that provides the present device 100 with audio data 220 related to the broadcast audio 210.
[0013] The device 100 and the voice recognition server 290 may be connected via a public communication network 330 such as the Internet. The voice recognition server 290 receives voice data 220, performs voice recognition processing on the voice data 220, and notifies the device 100 of character information 240 generated as the voice recognition result 230.
[0014] The subtitle transmission system 320 is a system that broadcasts a broadcast image 310 generated by synthesizing the text information 260 proofread by the present device 100 with an original broadcast image 300.
[0015] As shown in FIG. 2, the device 100 has an audio data acquisition means 110, an audio data notification means 120, a text information acquisition means 130, a proofreading information display means 140, a proofreading instruction receiving means 150, a text information generation means 160, and a broadcast text information notification means 170.
[0016] The audio data acquiring means 110 acquires audio data 220 relating to the broadcast audio 210 from an external device 340. The audio data 220 is, for example, audio data relating to the voice uttered by an announcer, a narrator, or the like.
[0017] The voice data notification means 120 notifies the voice data 220 acquired by the voice data acquisition means 110 to a voice recognition server 290 that uses AI (Artificial Intelligence). The voice recognition server 290 performs voice recognition processing on the acquired voice data 220 and generates text information 240.
[0018] The character information acquisition means 130 acquires the voice recognition result 230 of the voice data 220 from the voice recognition server 290 as time-series character information 240. The voice data 220 is related to the broadcast voice 210, and therefore occurs continuously one after another, and each of the data has a concept of coming before and coming after. In other words, the voice data 220 is an arrangement of the broadcast voice 210 in chronological order, and therefore the character information 240 is also data having a time-series nature.
[0019] As shown in FIG. 3, the character information 240 is, for example, in the form of character information 1, character information 2, . . . , character information 6 in chronological order, and is data that is generated successively over time.
[0020] The proofreading information display means 140 displays the character information 240 on the proofreading display device 280 in chunks 250 of a predetermined size for proofreading by a user (proofreader) 270. As shown in Fig. 3, the character information 240 is divided into chunks 250 of a predetermined size, such as character information 1, character information 2, ..., character information 6. This is because the subtitles for broadcast are output (broadcast) by dividing them into sizes (units) that are easy for the viewer to see. The calibration display device 280 may be the display device 570 included in the present device 100 , or may be a display device connected to the present device 100 .
[0021] When a user 270 selects one block 250 of character information (time series data) 240 divided into blocks 250 of a predetermined size to be displayed on the proofreading display device 280, the proofreading information display means 140 excludes the character information 240 prior to that block 250 from being displayed (proofread) on the proofreading display device 280.
[0022] When the proofreading information display means 140 accepts a selection by a user 270 of one block 250 of character information 240, which is time-series data divided into blocks 250 of a predetermined size, to be displayed on the proofreading display device 280, the proofreading information display means 140 excludes the character information 240 prior to the one block 250 from the objects to be displayed (proofread) on the proofreading display device 280.
[0023] Even when commercials are inserted during a broadcast program or when comment follow captions appear, processing by audio data acquisition means 110, audio data notification means 120, and text information acquisition means 130 continues, so the amount of text information 240 that does not require work by user 270 continues to increase in device 100. And, if there is no processing by proofreading information display means 140, user 270 has to delete text information 240 that does not require proofreading by himself.
[0024] Therefore, when user 270 performs an operation to select character information 240 (a block 250) to be proofread, the proofreading information display means 140 deletes character information 240 prior to the character information 240 (a block 250) selected by user 270 as character information 240 that does not require proofreading, in other words, excludes it from being displayed on the proofreading display device 280.
[0025] 3, character information 240 is time-series data divided into chunks 250 of a predetermined size, for example, character information 1, character information 2, ..., character information 6. In this case, when user 270 performs a selection operation for chunk 250 identified by character information 4 to be displayed on proofreading display device 280, proofreading information display means 140 excludes character information 1, character information 2, and character information 3 that precede character information 4 selected by user 270 from the objects to be displayed on proofreading display device 280.
[0026] In addition, excluding specific character information 240 from being displayed on proofreading display device 280 may be in the form of deleting specific character information 240 from being displayed on display device 570 included in device 100.
[0027] 4, in the above case, proofreading information display means 140 displays character information 4, selected by user 270, on proofreading display device 280, and makes it the subject of proofreading by user 270. Subsequently, when proofreading information display means 140 receives a selection operation by user 270 to make character information 5 and character information 6 the subject of proofreading, it also displays character information 5 and character information 6 on proofreading display device 280.
[0028] By the processing by the proofreading information display means 140 as described above, the present device 100 can reduce the workload of the person in charge of proofreading 270 when producing subtitles for broadcast 260 from computer-generated speech recognition results 230. The proofreading instruction receiving means 150 receives a proofreading instruction from the user 270 regarding the text information 240 . The character information generating means 160 corrects the character information 240 in accordance with the proofreading instructions received by the proofreading instruction receiving means 150 , and generates corrected character information 260 .
[0029] The broadcast text information notification means 170 notifies the subtitle sending system 320 of the corrected text information 260 generated by the text information generation means 160. The subtitle sending system 320 broadcasts the broadcast video 310 generated by synthesizing the acquired corrected text information 260 (broadcast subtitles) with the broadcast original video 300.
[0030] Based on the above-mentioned operating principle, this device 100 reduces the workload of a person in charge of proofreading 270 when producing subtitles 260 for broadcast from a speech recognition result 230 by a computer. (Hardware configuration of the subtitling device according to this embodiment)
[0031] An example of the hardware configuration of the present device 100 will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the hardware configuration of the present device 100. As shown in Fig. 5, the present device 100 has a CPU (Central Processing Unit) 510, a ROM (Read-Only Memory) 520, a RAM (Random Access Memory) 530, an auxiliary storage device 540, a communication I / F 550, an input device 560, an output device 570, and a storage medium I / F 580.
[0032] CPU 510 is a device that executes a program stored in ROM 520, performs arithmetic processing on data expanded (loaded) into RAM 530 in accordance with instructions from the program, and controls the entire device 100. ROM 520 stores programs and data to be executed by CPU 510. When CPU 510 executes a program stored in ROM 520, the program and data to be executed are expanded (loaded) into RAM 530, and RAM 530 temporarily holds the arithmetic data during calculation.
[0033] The auxiliary storage device 540 is a device that stores the OS (Operating System) which is basic software, the application program according to the present embodiment, and the like together with related data. The auxiliary storage device 540 is, for example, a HDD (Hard Disk Drive) or a flash memory.
[0034] The communication I / F 550 is an interface for connecting to a communication network 330 such as a wired or wireless LAN (Local Area Network) or the Internet, and for transmitting and receiving data to and from another device 340 that provides a communication function.
[0035] The input device 560 is a device such as a keyboard for inputting data to the present device 100. The output device 570 also includes a device configured with an LCD (Liquid Crystal Display) or the like, and functions as a user interface when the user utilizes the functions of the present device 100 or when various settings are made. The storage medium I / F 580 is an interface for transmitting and receiving data to and from a storage medium 590 such as a CD-ROM, a DVD-ROM, or a USB memory.
[0036] Each of the means of the present device 100 may be realized by CPU 510 executing a program corresponding to each of the means stored in ROM 520 or auxiliary storage device 540. Each of the means of the present device 100 may be realized by hardware that performs processing related to each of the means. Alternatively, the present device 100 may be caused to read the program according to the present invention from an external server device via communication I / F 550, or to read the program according to the present invention from storage medium 590 via storage medium I / F 580, and cause the present device 100 to execute the program. (Processing example by the subtitling device according to the present embodiment) A flow of an example of processing by the present device 100 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing a flow of an example of processing by the present device 100.
[0037] In S10, the audio data acquisition means 110 acquires audio data 220 relating to the broadcast audio 210 from the external device 340. The audio data 220 is, for example, audio data relating to the voice uttered by an announcer, a narrator, or the like.
[0038] In S20, the voice data notification means 120 notifies the voice data 220 acquired in S10 to the voice recognition server 290. The voice recognition server 290 performs voice recognition processing on the acquired voice data 220 and generates text information 240.
[0039] In S30, the character information acquisition means 130 acquires the voice recognition result 230 of the voice data 220 from the voice recognition server 290 as time-series character information 240. As shown in Fig. 3, the character information 240 has a form of, for example, character information 1, character information 2, ..., character information 6 in chronological order, and is data that is generated successively over time.
[0040] In S40, the proofreading information display means 140 displays the character information 240 on the proofreading display device 280 in chunks 250 of a predetermined size for proofreading by the user 270. As shown in Fig. 3, the character information 240 is divided into chunks 250 of a predetermined size, such as character information 1, character information 2, ..., character information 6. This is because the subtitles for broadcast are output (broadcast) by dividing them into sizes (units) that are easy for the viewer to see.
[0041] Furthermore, in S40, when the user 270 selects one block 250 of character information (time series data) 240 divided into blocks 250 of a predetermined size to be displayed on the proofreading display device 280, the proofreading information display means 140 excludes the character information 240 prior to that block 250 from being displayed (proofread) on the proofreading display device 280.
[0042] In other words, when the proofreading information display means 140 accepts a selection by the user 270 of one block 250 of character information 240, which is time-series data divided into blocks 250 of a predetermined size, to be displayed on the proofreading display device 280, the proofreading information display means 140 excludes the character information 240 prior to the one block 250 from the objects to be displayed (proofread) on the proofreading display device 280.
[0043] Even when commercials are inserted during a broadcast program or when comment follow captions appear, processing by audio data acquisition means 110, audio data notification means 120, and text information acquisition means 130 continues, so the amount of text information 240 that does not require work by user 270 continues to increase in device 100. And, if there is no processing by proofreading information display means 140, user 270 has to delete text information 240 that does not require proofreading by himself.
[0044] Therefore, when user 270 performs an operation to select character information 240 (first block 250) to be proofread, the proofreading information display means 140 deletes character information 240 prior to the character information 240 (first block 250) selected by user 270 as character information 240 that does not require proofreading, in other words, excludes it from being displayed on the proofreading display device 280.
[0045] 3, character information 240 is time-series data divided into chunks 250 of a predetermined size, for example, character information 1, character information 2, ..., character information 6. In this case, when user 270 performs a selection operation for displaying chunk 250 identified by character information 4 on proofreading display device 280, proofreading information display means 140 excludes character information 1, character information 2, and character information 3 that precede character information 4 selected by user 270 from the objects to be displayed on proofreading display device 280.
[0046] In addition, excluding specific character information 240 from being displayed on proofreading display device 280 may be in the form of deleting specific character information 240 from being displayed on display device 570 included in device 100.
[0047] 4, in the above case, proofreading information display means 140 displays character information 4, selected by user 270, on proofreading display device 280, and makes it the subject of proofreading by user 270. Subsequently, when proofreading information display means 140 receives a selection operation from user 270 to make character information 5 and character information 6 the subject of proofreading, it also displays character information 5 and character information 6 on proofreading display device 280.
[0048] By the processing by the proofreading information display means 140 as described above, the present device 100 can reduce the workload of the person in charge of proofreading 270 when producing subtitles for broadcast 260 from computer-generated speech recognition results 230. In S50, the proofreading instruction receiving means 150 receives a proofreading instruction from the user 270 regarding the text information 240. In S60, the character information generating means 160 corrects the character information 240 in accordance with the proofreading instruction received in S50, and generates corrected character information 260.
[0049] In S70, the broadcast text information notification means 170 notifies the subtitle sending system 320 of the corrected text information 260 generated in S60. The subtitle sending system 320 broadcasts the broadcast video 310 generated by synthesizing the acquired corrected text information 260 (broadcast subtitles) with the broadcast original video 300.
[0050] By carrying out the above-mentioned processing, the present device 100 reduces the workload of a person in charge of proofreading 270 when producing subtitles 260 for broadcast from a speech recognition result 230 by a computer.
[0051] Although the embodiment of the present invention has been described in detail above, the present invention is not limited to such specific embodiment, and various modifications and variations are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]
[0052] 100 Subtitle production equipment 110 Voice data acquisition means 120 Voice data notification means 130 Character information acquisition means 140 Calibration information display means 150 Proofreading instruction receiving means 160 Character information generation means 170 Broadcast text information notification means 210 Broadcast Audio 220 Audio data related to broadcast audio 230 Voice recognition results for voice data 240 Text information 250 Textual information chunks 260 Corrected character information 270 users (proofreaders) 280 Calibration display device 290 Voice Recognition Server 300 Original broadcast footage 310 Broadcast footage 320 Subtitle Transmission System 330 Communication Network 340 External device 510 CPU 520 ROM 530 RAM 540 Auxiliary storage 550 Communication Interface 560 Input Device 570 Output Device 580 Storage Media Interface 590 Storage medium
Claims
1. a character information acquiring means for acquiring a voice recognition result of voice data relating to a broadcast voice as time-series character information; a proofreading information display means for displaying the character information on a proofreading display device for proofreading by a user for each block of a predetermined size; a character information generating means for correcting the character information in accordance with a proofreading instruction from the user and generating corrected character information, A subtitling production device characterized in that when the proofreading information display means accepts a selection by the user to display a specified block of the character information, which is time-series data divided into blocks of the specified size, on the proofreading display device, the proofreading information display means excludes the character information prior to the specified block from being displayed on the proofreading display device.
2. 2. The subtitling device according to claim 1, wherein said proofreading information display means deletes said character information preceding said predetermined block from candidates for selection by said user.
3. a voice data acquisition means for acquiring the voice data from an external device; a voice data notifying means for notifying an external voice recognition server of the voice data; 2. The subtitling device according to claim 1, wherein said character information acquiring means acquires, from said voice recognition server, said time-series character information generated by said voice recognition server.
4. 4. The subtitling production device according to claim 1, further comprising a broadcast character information notifying means for notifying the corrected character information to a subtitle sending system which broadcasts a broadcast image in which subtitle information is synthesized with an original broadcast image.
5. A step in which a character information acquisition means acquires a voice recognition result of voice data related to a broadcast voice as time-series character information; a step in which a proofreading information display means displays the character information on a proofreading display device for proofreading by a user for each block of a predetermined size; A character information generating means corrects the character information in accordance with a proofreading instruction from the user to generate corrected character information; A subtitle production method characterized in that, when the proofreading information display means accepts a selection by the user for a specified block of the character information, which is time-series data divided into blocks of the specified size, to be displayed on the proofreading display device, the proofreading information display means excludes the character information prior to the specified block from being displayed on the proofreading display device.
6. 6. The subtitling method according to claim 5, wherein said proofreading information display means deletes said character information preceding said predetermined block from candidates for selection by said user.
7. A step of acquiring the voice data from an external device by a voice data acquisition means; a step of notifying an external voice recognition server of the voice data by a voice data notifying means; 6. The subtitling method according to claim 5, wherein said character information acquiring means acquires, from said voice recognition server, said time-series character information generated by said voice recognition server.
8. 8. A subtitle production method according to claim 5, further comprising a step of notifying a subtitle sending system which broadcasts a broadcast image in which subtitle information is mixed with an original broadcast image, said modified character information being notified by a broadcast character information notifying means.
9. A subtitling program for causing a computer to execute the method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Subtitle generator, method and program
JP2023067353A