System for broadcasting and subtitle creation method

The broadcasting system addresses the limitations of existing subtitle systems by employing multiple end-to-end voice recognition models to adapt to different program types, ensuring accurate subtitle generation for both live and recorded content.

JP2025147254APending Publication Date: 2025-10-07NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024047437
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Existing subtitle broadcasting systems are not suitable for recorded broadcasts and struggle to generate appropriate subtitles for different program genres due to the use of a single acoustic model.

Method used

A broadcasting system equipped with multiple end-to-end voice recognition models tailored to program characteristics, allowing for real-time correction and learning using corrected text data to generate accurate subtitles for both live and recorded broadcasts.

Benefits of technology

The system can generate appropriate subtitles for both live and recorded broadcasts, improving accuracy by using genre-specific voice recognition models and continuous learning with corrected data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025147254000001_ABST
    Figure 2025147254000001_ABST
Patent Text Reader

Abstract

To provide a system for broadcasting and a subtitle creation method that are adaptive to both a live broadcast and recorded broadcast, and can create proper subtitles according to the kind of a program.SOLUTION: A system 100 for broadcasting comprises: speech recognition means 101 which has a plurality of end-to-end speech recognition models corresponding to features of programs, and uses one of the end-to-end speech recognition models to recognize speeches of speakers of a program; and feedback means 102 which outputs text data, generated by proofreading text data based upon recognition results of the speech recognition means 101, as subtitles, and also outputs the data to the speech recognition means as correct answer data of the subtitles, wherein the end-to-end speech recognition models learn the text data as the correct answer data.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a broadcasting system and a subtitling method. [Background technology]

[0002] Subtitle broadcasting, which superimposes subtitles corresponding to the audio of speakers in a television (hereafter referred to as TV) program onto the video of the program, has been realized. Subtitle broadcasting is useful for elderly people and people with hearing impairments who have difficulty hearing the audio of television programs. Subtitle broadcasting, which adds subtitles to live programs in real time, has been realized.

[0003] Patent Document 1 describes generating subtitles using speech recognition. Patent Document 1 also describes placing one or more operators who correct errors in real time after the speech recognition device in order to correct recognition errors in speech recognition. Patent Document 1 also describes transmitting manually corrected text as subtitled broadcast.

[0004] Patent Document 1 also describes generating a word lattice consisting of a network of candidate words using a preset acoustic model, language model, and pronunciation dictionary. Patent Document 1 further describes feeding back error-corrected text to a speech recognition device.

[0005] Patent Document 2 describes a technique in which the speech recognition decoder that performs speech recognition is switched to the latest model, and highly accurate speech recognition is continuously realized using the latest model at all times.

[0006] Patent document 3 describes a device that has speech language corpora for each genre, such as news programs, sports programs, and information programs (each of which is a speech language corpus for a specific program), and selects the speech language corpus for a specific program depending on the program and trains an acoustic model. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-197410 [Patent Document 2] Japanese Patent Application Laid-Open No. 2010-54685 [Patent Document 3] Japanese Patent Application Laid-Open No. 2017-45027 Summary of the Invention [Problem to be solved by the invention]

[0008] The systems including the speech recognition devices described in Patent Documents 1 and 2 aim to display subtitles as accurately as possible during subtitled broadcasts, and therefore are not suitable for use in recorded broadcasts that use pre-recorded material.

[0009] Furthermore, the device described in Patent Document 3 uses different speech language corpora, but basically uses one acoustic model, which makes it difficult to say that it is adequately suited to generating appropriate subtitles for each genre.

[0010] An object of the present invention is to provide a broadcasting system and a subtitling method that can handle both live broadcasts and recorded broadcasts and generate appropriate subtitles depending on the type of program. [Means for solving the problem]

[0011] A broadcasting system based on the present disclosure is a broadcasting system having the function of generating subtitles corresponding to the voice of a speaker in a program, and is equipped with a voice recognition means having a plurality of end-to-end (E2E) voice recognition models according to the characteristics of the program and recognizing the voice of a speaker in the program using one of the end-to-end voice recognition models, and a feedback means for outputting text data that has been corrected based on the recognition results of the voice recognition means as subtitles and outputting it to the voice recognition means as correct data for the subtitles, and the end-to-end voice recognition model is trained using the text data as correct data.

[0012] A subtitle creation method based on the present disclosure is a subtitle creation method that generates subtitles corresponding to the speech of a speaker in a program, recognizes the speech of the speaker in the program using one of a plurality of end-to-end speech recognition models according to the characteristics of the program, corrects text data based on the recognition results, outputs the corrected text data as subtitles, and feeds it back to the speech recognition model as correct data for the subtitles, and the end-to-end speech recognition model trains using the text data as correct data.

[0013] A subtitle creation program based on the present disclosure causes a computer to recognize the speech of a speaker in a program using one of a plurality of end-to-end speech recognition models according to the characteristics of the program, outputs text data that has been corrected based on the recognition results as subtitles, and feeds the corrected text data back into the speech recognition model as correct data for the subtitles, causing the end-to-end speech recognition model to learn the text data as correct data. [Effects of the Invention]

[0014] According to the present invention, a broadcasting system and a subtitling method are provided that can handle both live broadcasts and recorded broadcasts and generate appropriate subtitles depending on the type of program. [Brief explanation of the drawings]

[0015] [Figure 1]1 is a block diagram showing an example of the configuration of an embodiment of a broadcasting system. [Figure 2] FIG. 10 is an explanatory diagram showing an example of voice recognition and proofreading. [Figure 3] FIG. 10 is a block diagram showing an example of the configuration of another embodiment of a broadcasting system. [Figure 4] FIG. 1 is an explanatory diagram for explaining an example of a method for voice recognition, proofreading, and learning. [Figure 5] FIG. 10 is an explanatory diagram for explaining the processing of a voice recognition device and the like during live broadcasting. [Figure 6] FIG. 10 is an explanatory diagram for explaining the processing of a voice recognition device and the like during recorded broadcasting. [Figure 7] FIG. 10 is an explanatory diagram showing an example of how video data is used. [Figure 8] 10 is a flowchart showing the operation of the speech recognition device and the proofreading terminal when creating subtitles. [Figure 9] 10 is a flowchart showing the operation of the voice recognition device regarding re-learning. [Figure 10] FIG. 1 is a block diagram illustrating an example of the configuration of an information processing system. [Figure 11] FIG. 1 is a block diagram showing the main parts of a broadcasting system. DETAILED DESCRIPTION OF THE INVENTION

[0016] Hereinafter, an embodiment will be described with reference to the drawings.

[0017] Fig. 1 is a block diagram showing an example of the configuration of one embodiment of a broadcasting system. The unidirectional arrows in Fig. 1 simply indicate the direction of signal (data) flow, but do not exclude bidirectionality. This also applies to other block diagrams. The broadcasting system 1 illustrated in Fig. 1 may be referred to as the broadcasting system 1 of the first embodiment.

[0018] The broadcasting system 1 shown in FIG. 1 includes a speech recognition device 10, a proofreading terminal 11, a subtitle sending device 12, a video and audio processing device (hereinafter referred to as a video and audio device) 13, a subtitle insertion device 14, and a compression / multiplexing device 15.

[0019] The voice recognition device 10 receives as input the audio and video signals from camera material (live broadcast video and audio: live broadcast material) or recorded material, and performs voice recognition processing.

[0020] The proofreading terminal 11 is, for example, a personal computer. The proofreading terminal 11 inputs the recognition result of the speech recognition device 10 as text data. Several operators (for example, 1 to 5 operators) proofread the text displayed on the display unit of the proofreading terminal 11. Although one proofreading terminal 11 is shown in FIG. 1, in reality, there are several proofreading terminals 11 that can be operated by each operator.

[0021] Furthermore, the proofreading terminal 11 feeds back the proofread text data to the speech recognition device 10.

[0022] The subtitle sending device 12 receives the proofread text data from the proofreading terminal 11. The subtitle sending device 12 outputs the received text data as subtitle data to the subtitle insertion device 14 at a predetermined timing.

[0023] The audio-visual device 13 receives video and audio signals from camera footage or recorded footage and adds required signals (e.g., signals related to audio-visual effects). Hereinafter, the signals (data) output by the audio-visual device 13 will be referred to as audio-visual data. The audio-visual device 13 also outputs audio-visual signals obtained by performing predetermined signal processing on the audio-visual data to the subtitle insertion device 14. The subtitle insertion device 14 superimposes subtitle data on the audio-visual signals and outputs the audio-visual signals on which the subtitle data has been superimposed to the compression / multiplexing device 15.

[0024] The compression / multiplexing device 15 performs data compression on the video and audio signals, and also multiplexes necessary data onto the data-compressed video and audio signals.

[0025] FIG. 2 is an explanatory diagram showing an example of voice recognition and proofreading in the broadcasting system 1. As shown in FIG.

[0026] In the broadcasting system 1 of the first embodiment, the voice recognition device 10 has voice recognition functions (voice recognition engines) for different program types (hereinafter referred to as genres), such as news programs, sports programs, and information programs. In Fig. 2, these are shown as an engine for program A, an engine for program B, and an engine for program C. The voice recognition device 10 uses one of the voice recognition engines in response to an instruction from outside the broadcasting system 1.

[0027] Note that preparing a voice recognition engine for each genre is just an example, and a voice recognition engine may be prepared according to some feature of the program. For example, it is conceivable to prepare a voice recognition engine according to the nature of the program or according to the speaking style of the speaker in the program. Note that the following explanation will be given taking as an example a case where a voice recognition engine for each genre is prepared.

[0028] Each speech recognition engine includes a learning model for speech recognition. As will be described later, the learning model is a model (E2E learning model) based on end-to-end (E2E) deep learning. As noted in FIG. 2, the speech recognition device 10 may have a speaker identification function in addition to the speech recognition function of each speech recognition engine (learning model, specifically, E2E learning model). The speech recognition devices described in Patent Documents 1 and 2 include an acoustic model, a language model, and a pronunciation dictionary. That is, these speech recognition devices perform training by dividing a neural network into multiple subtasks.

[0029] Figure 2 shows an example in which the speech recognition device 10 outputs text data of "Thank you very much. Goodbye" as a speech by person C, and the operator at the proofreading terminal 11 proofreads it to "Thank you very much. Goodbye" as a speech by person B.

[0030] The speaker is identified, for example, by the operator of the proofreading terminal 11. The operator can identify the speaker by, for example, visually checking the video data.

[0031] If the speech recognition device 10 also has a speaker identification function, the speech recognition device 10 identifies the speaker. Then, the speech recognition device 10 supplies information indicating the speaker (speaker data) to the proofreading terminal 11. In this case, the operator of the proofreading terminal 11 also targets the speaker data for proofreading.

[0032] The proofread text data and information indicating the speaker (speaker data) are fed back to the speech recognition device 10 together with the text data before proofreading (spoken voice data) as appropriate. In the speech recognition device 10, the speech recognition engine uses the proofread text data as correct answer data. Furthermore, if the speech recognition engine also has a speaker identification function, the speaker identification function uses the speaker data as correct answer data for speaker identification.

[0033] Fig. 3 is a block diagram showing an example of the configuration of another embodiment of a broadcasting system. The broadcasting system 2 shown in Fig. 3 may be referred to as the broadcasting system 2 of the second embodiment.

[0034] The broadcasting system 2 of the second embodiment includes a data server (DS) and an automatic program control system (APS) in addition to the configuration of the broadcasting system 1 of the first embodiment. The DS is a facility that creates program progress data and centrally manages broadcast data. The APS is a facility that controls devices in the broadcasting system 2 in real time according to the program progress data (including the program schedule) received from the DS. Figure 3 shows the DS and APS collectively represented as DS / APS 20.

[0035] In the second embodiment, the speech recognition device 10 receives program metadata as program information from the DS / APS 20, and selects a speech recognition engine corresponding to the genre identified by the program metadata under the control of the APS. That is, the speech recognition device 10 can use a speech recognition engine suited to the characteristics of the program in cooperation with the DS / APS 20. The program metadata includes information that can identify the program and the program schedule.

[0036] FIG. 4 is an explanatory diagram showing an example of voice recognition and calibration in the broadcasting system 2 of the second embodiment.

[0037] An example of customizing a speech recognition engine (generating a speech recognition engine for each genre) is also illustrated in Figure 4. In the example shown in Figure 4, a speech recognition engine for each genre is generated based on a large-scale trained general-purpose engine (a training model for general-purpose speech recognition).

[0038] Figure 4 shows examples of customized voice recognition engines (learning models, i.e., voice recognition models), including a voice recognition engine specialized for news programs, a voice recognition engine specialized for animation programs, and a voice recognition engine specialized for sports programs.

[0039] As an example, the general-purpose engine is made to perform learning related to voice recognition in news programs. For example, if there is an error in the text data resulting from voice recognition in a news program, the general-purpose engine is made to perform learning using the text data in which the error has been corrected (proofread) as the correct answer data. By repeating such learning, a voice recognition engine specialized for news programs can be obtained based on the general-purpose engine. Note that in this embodiment, the text data proofread by the operator of the proofreading terminal 11 is used as the correct answer data.

[0040] Although a voice recognition engine specialized for news programs has been used as an example here, voice recognition engines can be similarly obtained for other genres.

[0041] FIG. 5 is an explanatory diagram for explaining the processing of the voice recognition device and the like during live broadcasting.

[0042] The APS 202 controls broadcasting in accordance with the program progress data received from the DS 201. The DS 201 and the APS 202 correspond to the DS / APS 20 shown in Figs. 3 and 4. Hereinafter, the signal from the APS 202 will be referred to as broadcast control data. The broadcast control data is supplied to the speech recognition device 10. The broadcast control data (including the program progress data) includes metadata of the program.

[0043] The voice recognition device 10 identifies the voice recognition engine to be used, i.e., the voice recognition engine corresponding to the genre to which the program belongs, based on the broadcast control data, specifically the program progress data. The voice recognition device 10 also starts storing the program data in the voice recognition device 10.

[0044] The speech recognition device 10 executes speech recognition processing using the identified speech recognition engine. Although not explicitly shown in FIG. 5, if the speech recognition device 10 has a speaker recognition function, it includes, for example, a speaker recognition function (speaker recognition model). Then, the speech recognition device 10 executes speaker recognition processing using the speaker recognition function. Then, the speaker recognition result that has been proofread by an operator is fed back to the speech recognition device 10 as correct answer data for the speaker.

[0045] The speech recognition device 10 having a speaker recognition function performs re-learning using correct answer data of the speaker.

[0046] The speech recognition device 10 also accumulates program data. The program data includes, for example, audio data and video data. The speech recognition device 10 adds information (speaker data) indicating a speaker identified by speaker recognition processing to the speech data.

[0047] 5 shows a broadcasting system 2 according to the second embodiment. When the broadcasting system 1 according to the first embodiment is used, the voice recognition device 10 executes voice recognition processing and speaker recognition processing using a voice recognition engine corresponding to the genre to which the program belongs, in response to an instruction from outside the broadcasting system 1. The voice recognition device 10 also stores program data in response to an instruction from outside the broadcasting system 1. The instruction from outside is given, for example, by an operator of the APS 202.

[0048] FIG. 5 also illustrates that correct text and correct speaker data are stored together with the audio data and video data.

[0049] As described above, in the first and second embodiments, each genre-specific speech recognition model is configured as an E2E learning model. That is, the speech recognition model converts speech data into text data using a single neural network. In other words, the speech recognition engine executes speech recognition processing using a single model.

[0050] 5, the speech recognition engine performs the above-mentioned E2E speech recognition through the processes of data acquisition (acquisition of accumulated speech data), feature extraction, and data conversion. Note that, although it is said that an E2E learning model that uses one model (E2E learning model) is difficult to customize, in the first and second embodiments, a speech recognition model is prepared for each genre, thereby eliminating the drawback of difficult customization.

[0051] As described above, the text data that is the result of the speech recognition processing (the output of the E2E learning model) is proofread by an operator and then used as subtitle data, and is also fed back to the speech recognition engine as correct answer data (correct answer text for speech recognition).

[0052] A voice recognition engine having a speaker recognition function may re-learn using the correct voice recognition text and the correct speaker data as training data each time a program ends, or may re-learn after a certain amount of correct voice recognition text has been generated.

[0053] Note that although FIG. 5 shows the part that stores data and the E2E learning model as existing outside the speech recognition device 10, in reality, at least the E2E learning model is built into the speech recognition device 10.

[0054] 5 correspond in time to the audio data and video data. That is, they are time-synchronized. The correct answer text and the correct answer data of the speaker are time-synchronized with the audio data and video data.

[0055] FIG. 6 is an explanatory diagram for explaining the processing of the voice recognition device and the like during recording and broadcasting.

[0056] The APS 202 outputs broadcast control data to the voice recognition device 10 in accordance with the program progress data received from the DS 201. The voice recognition device 10 stores program data.

[0057] FIG. 6 illustrates that speaker subtitle data and program metadata are stored along with the subtitle data, audio data, and video data contained in the recorded material.

[0058] When a recorded broadcast is performed, the program data generally includes subtitle data. If speaker information is also attached to the subtitle data, the speaker data accompanying the subtitle data is also included in the program data. The speech recognition device 10 executes speech recognition processing and speaker identification processing at an appropriate time after the program is recorded, and can edit the subtitles via the proofreading terminal 11. The speech recognition processing and speaker identification processing at this time are the same as those performed during live broadcasting described above.

[0059] When a recorded broadcast is made, the correct answer data of the correct text and the speaker has already been collected, so the speech recognition device 10 can also perform learning using the correct answer data at that time.

[0060] After a program ends, when the recognition engine corresponding to that program is not in use, the speech recognition engine in the speech recognition device 10 is retrained using the speech recognition correct answer text and the speaker's correct answer data. However, retraining may be performed after a certain amount of speech recognition correct answer text has been generated. Note that "after a certain amount of speech recognition correct answer text has been generated" refers to after a certain amount of the program has been broadcast.

[0061] When re-learning is performed after a certain amount of correct answer text for speech recognition has been generated, the speech recognition engine starts re-learning when a certain amount of correct answer text for speech recognition has been generated.

[0062] Note that although FIG. 6 shows the part that stores data and the E2E learning model as existing outside the speech recognition device 10, in reality, at least the E2E learning model is built into the speech recognition device 10.

[0063] Furthermore, the audio data and video data shown in FIG. 6 are time-synchronized.

[0064] As described above, in the broadcasting system 1 of the first embodiment and the broadcasting system 2 of the second embodiment, the voice recognition device 10 creates subtitle data using machine learning, specifically, voice recognition processing using a voice recognition model. A voice recognition engine (voice recognition model) is prepared for each program type. This improves the accuracy of creating subtitle data. Higher creation accuracy means that subtitle data with fewer errors can be created.

[0065] Furthermore, the voice recognition engine re-learns using the corrected correct text as training data. The text data that has been corrected is more likely to be correct data. Therefore, the accuracy of voice recognition by the voice recognition device 10 is further improved.

[0066] Furthermore, whether the system is configured to re-learn the correct answer text for speech recognition and the correct speaker data as training data after the program ends, or whether the system is configured to re-learn the correct answer text for speech recognition and the correct speaker data as training data after a certain amount of correct answer text for speech recognition has been generated, there is no difference in terms of continuous learning while the broadcasting system is in operation.

[0067] Furthermore, in the broadcasting system 2 of the second embodiment, the speech recognition device 10 can use a speech recognition engine suited to the features of a program in cooperation with the DS / APS 20, so a more appropriate speech recognition engine can be used compared to when a speech recognition engine is manually specified. As a result, the accuracy of the speech recognition engine can be improved more appropriately.

[0068] Next, an example of using video data for speech recognition and speaker recognition will be described. Figure 7 is an explanatory diagram showing an example of using video data.

[0069] The left side of Fig. 7 shows an example of the configuration and learning when the broadcasting system 1 of the first embodiment or the broadcasting system 2 of the second embodiment includes the above-mentioned voice recognition device 10. The right side of Fig. 7 shows an example of the configuration and learning when the broadcasting system 1 of the first embodiment or the broadcasting system 2 of the second embodiment includes the voice image recognition device 30 instead of the above-mentioned voice recognition device 10.

[0070] Also, Figure 7 shows an example in which the speech recognition device 10 outputs text data of "Goodbye" as a speech made by person C, and the text data is proofread to "Goodbye" as a speech made by person B, an operator at the proofreading terminal 11.

[0071] In addition to voice data, video data is also input to the voice image recognition device 30. In addition to the voice recognition function of the voice recognition device 10, the voice image recognition device 30 also has, for example, an image recognition function for recognizing people (e.g., human faces) from video. Hereinafter, it is assumed that the voice image recognition device 30 has a face recognition function for recognizing faces. For example, the voice image recognition device 30 has a general face recognition function that uses deep learning, and a function for comparing a face detected by face recognition with faces registered in a database to identify the person whose face is detected.

[0072] The voice / image recognition device 30 also has a function of associating a person with the content of an utterance based on the timing of the utterance and the timing of the person's appearance in the video.

[0073] FIG. 7 illustrates that the voice image recognition device 30 recognizes "Mr. B" in response to the utterance of "Goodbye."

[0074] Facial recognition by the voice and image recognition device 30 can be applied in a variety of ways. For example, Fig. 7 shows that the operator of the proofreading terminal 11 or a voice recognition engine including a speaker recognition function identifies "Mr. B" corresponding to the timing of the utterance of "Goodbye," and corrects "Mr. C" to "Mr. B." In this case, the voice and image recognition device 30 is expected to improve the accuracy of speaker recognition by also using the results of facial recognition.

[0075] In other words, it is expected that the accuracy of the learning model will be improved by also using the results of face recognition. For example, the voice / image recognition device 30 can also extract features by using video data in speaker identification.

[0076] Although FIG. 7 shows the E2E learning model as existing outside the speech recognition device 10 and the speech image recognition device 30, at least the E2E learning model is built into the speech recognition device 10 and the speech image recognition device 30.

[0077] Next, the operations of the speech recognition device 10 and the proofreading terminal 11 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing the operations of the speech recognition device 10 and the proofreading terminal 11 when creating subtitles.

[0078] The voice recognition device 10 performs voice recognition processing on the material voice data using a voice recognition engine corresponding to the program (step S101). As described above, in the first embodiment, the voice recognition engine is specified from outside the broadcasting system 1. In the second embodiment, the voice recognition device 10 specifies the voice recognition engine to be used based on program information and the like input from the broadcasting system 1.

[0079] The speech recognition device 10 outputs text data, which is the result of the speech recognition processing, to the proofreading terminal 11 (step S102). If the speech recognition device 10 also has a speaker identification function, the speech recognition device 10 also supplies information indicating the speaker (speaker data) to the proofreading terminal 11.

[0080] The operator of the proofreading terminal 11 proofreads the text data (step S103). If speaker data is also supplied from the speech recognition device 10, the operator also proofreads the speaker data. The proofreading terminal 11 outputs the proofread text data to the subtitle sending device 12 (step S104). If the speaker data has also been proofread, the proofreading terminal 11 also outputs the proofread speaker data to the subtitle sending device 12.

[0081] The speech recognition device 10 stores the proofread text data as correct answer data (step S105). If the speech recognition device 10 also has a speaker identification function, the speech recognition device 10 stores the proofread speaker data as correct answer data for the speaker.

[0082] The above process is repeated until the program ends or the subtitle insertion period in the program ends.

[0083] FIG. 9 is a flowchart showing the re-learning operation of the voice recognition device 10.

[0084] When the time for re-learning arrives, the speech recognition device 10 re-trains the speech recognition engine using the correct answer data of the text data (steps S201 and S202). If the speech recognition device 10 also has a speaker identification function, the speech recognition device 10 performs re-learning using the correct answer data of the speaker (step S203).

[0085] Each of the above embodiments can be configured by hardware, but can also be realized by a computer program.

[0086] Fig. 10 is a block diagram showing an example configuration of an information processing system (computer) capable of realizing the voice recognition device 10 and the voice and image recognition device 30 in the broadcasting systems 1 and 2. The information processing system shown in Fig. 10 includes a processor 701 such as a CPU (Central Processing Unit), a program memory 702, and a storage medium 703. The storage medium may be a semiconductor memory such as a flash ROM (Read Only Memory) or a magnetic storage medium such as a hard disk.

[0087] In the information processing system, a program memory 702 stores a program (subtitle creation program) for realizing the functions of the voice recognition device 10 and the voice and image recognition device 30 shown in the above embodiments.

[0088] The processor 701 executes processing in accordance with the programs stored in the program memory 702, thereby realizing the functions of the speech recognition device 10 and the speech and image recognition device 30 described in the embodiments.

[0089] At least the program memory 702 is a non-transitory computer-readable medium. However, the program may be stored in various types of transitory computer-readable medium. The program is supplied to the transitory computer-readable medium, for example, via a wired or wireless communication path, i.e., via an electrical signal, an optical signal, or an electromagnetic wave.

[0090] Fig. 11 is a block diagram showing the main parts of a broadcasting system. The broadcasting system 100 shown in Fig. 11 includes a speech recognition means 101 (implemented by a speech recognition device 10 in the embodiment) that has a plurality of end-to-end speech recognition models according to the features of a program and recognizes the speech of a speaker in the program using any of the end-to-end (E2E) speech recognition models, and a feedback means 102 (implemented by a proofreading terminal 11 in the embodiment) that outputs text data that is proofread based on the recognition result of the speech recognition means 101 as subtitles and also outputs the text data to the speech recognition means 101 as correct data for the subtitles, and the end-to-end speech recognition model is trained using the text data as correct data. [Explanation of symbols]

[0091] 1,2 Broadcasting systems 10 Voice recognition device 11 Proofreading Terminal 12 Subtitle transmission device 13. Audio-visual equipment (audio-visual processing equipment) 14 Subtitle Insertion Device 15 Compression / Multiplexing Device 20 DS / APS 30 Voice and image recognition device 100 Broadcast Systems 101 Voice Recognition Means 102 Feedback Tools 201DS 202 APS 701 processor 702 program memory 703 Storage medium

Claims

1. A broadcasting system having a function of generating subtitles corresponding to the voice of a speaker in a program, a speech recognition means having a plurality of end-to-end speech recognition models according to the characteristics of the program, and recognizing the speech of a speaker in the program using any of the end-to-end speech recognition models; a feedback means for outputting text data obtained by correcting text data based on the recognition result of said voice recognition means as subtitles and for outputting the correct data for the subtitles to said voice recognition means, The end-to-end speech recognition model is trained using the text data as correct answer data. Broadcasting system.

2. At least, a storage means for storing the audio data of the live broadcast material and the correct answer data is provided.

2. The broadcasting system of claim 1.

3. an automatic program control device for transmitting data including program schedules and recorded material; The speech recognition means identifies a speech recognition model to be used based on the data.

2. The broadcasting system of claim 1.

4. At least, a storage means for storing the voice data input from the automatic program control device and the correct answer data is provided.

4. The broadcasting system according to claim 3.

5. a speaker recognition means having a speaker recognition model for identifying a speaker in a program and outputting speaker data indicative of the speaker; the feedback means outputs the calibrated speaker data to the speaker recognition means; The speaker recognition model is trained using the calibrated speaker data as corrective data. A broadcasting system according to any one of claims 1 to 4.

6. Equipped with image recognition means to identify speakers from video data of live broadcast materials or recorded materials A broadcasting system according to any one of claims 1 to 4.

7. A subtitling method for generating subtitles corresponding to the voice of a speaker in a program, comprising: Recognizing the voice of a speaker in a program using one of a plurality of end-to-end voice recognition models according to the characteristics of the program; The text data corrected based on the recognition result is output as subtitles, and is fed back to the speech recognition model as correct data for the subtitles. The end-to-end speech recognition model is trained using the text data as correct answer data. How to create subtitles.

8. receiving data from an automatic program control device that transmits data including program schedules and recording material; Identifying a speech recognition model to be used based on the data The subtitling method according to claim 7.

9. On the computer, Recognizing the voice of a speaker in the program using one of a plurality of end-to-end voice recognition models according to the characteristics of the program; The text data corrected based on the recognition result is output as subtitles, and is fed back to the speech recognition model as correct data for the subtitles. The end-to-end speech recognition model is trained using the text data as correct answer data. Subtitling program for.

Citation Information

Patent Citations

  • Voice recognition device and voice recognition program

    JP2010054685A

  • Voice recognition device, voice recognition system, and voice recognition program

    JP2011197410A

  • Speech language corpus generation device and its program

    JP2017045027A