Subtitle generation apparatus, method and storage medium

By constructing a subtitle generation device and using historical data to infer the segmentation and combination positions, easy-to-read subtitles are generated, solving the problem of heavy manual correction burden in real-time subtitle generation and realizing automated subtitle production.

CN116072120BActive Publication Date: 2026-01-06KK TOSHIBA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211050256.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-01
Filing Date
2022-08-31
Publication Date
2026-01-06
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

In real-time caption generation, existing technologies require a large amount of manual correction of sound recognition results to ensure the readability of captions, resulting in a significant correction burden.

Method used

By constructing a subtitle generation device, the text of the sound recognition results stored in the history department is used as historical data. The generation department infers the segmentation and combination positions based on the historical data, generates subtitle text, and prompts the subtitles through the user interface department, reducing the burden of correction.

Benefits of technology

It enables the automatic generation of easy-to-read subtitles in real-time subtitle generation, reducing manual correction costs and improving subtitle production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116072120B_ABST
    Figure CN116072120B_ABST
Patent Text Reader

Abstract

This invention provides a subtitle generation apparatus, method, and storage medium that reduces the burden of corrections to make subtitles easier to read when generating subtitles in real time based on sound recognition results. The subtitle generation apparatus according to the embodiment includes an acquisition unit, a history unit, a generation unit, a history update unit, and a prompting unit. The acquisition unit acquires text from the sound recognition results sequentially. The history unit saves the text as historical data. The generation unit infers the segmentation and combination positions of the text based on one or more saved historical data points, and generates subtitle text based on the segmentation and combination positions and the historical data points. The history update unit updates the historical data based on the segmentation and combination positions. The prompting unit prompts the user with the subtitle text.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based on Japanese Patent Application 2021-178498 (filed on November 1, 2021), from which it enjoys priority. This application incorporates the entire contents of that application by reference. Technical Field

[0002] The implementation methods involve subtitle generation apparatus, methods, and storage media. Background Technology

[0003] In the production of subtitles for television programs, there are techniques that utilize sound recognition technology to automatically generate subtitles. In recent years, with the development of deep learning technology, the accuracy of sound recognition has improved rapidly. However, due to factors such as background noise, speech errors, and mispronunciation of terminology, sound recognition can still produce errors. Therefore, in actual subtitle production, not all subtitles are created using sound recognition. Instead, in most cases, the subtitles based on the sound recognition results are manually reviewed to correct errors before the final subtitles are produced.

[0004] Furthermore, in applications such as news programs where subtitles are generated in real-time based on sound recognition results, it is desirable to allow sufficient time for manual corrections, thus requiring the sound recognition results to be obtained as early as possible. An effective method for this is early determination. Early determination refers to a technique in pattern recognition, including sound recognition, where the search ends when the number of candidates in the search space decreases and the recognition result is output at that point. This technique is suitable for applications utilizing sound recognition in real-time. However, when using early determination, the search ends before the sound ends, resulting in a shortened sound recognition result, or sometimes conversely, the search space is not reduced while the sound recognition result becomes longer. Therefore, when using early determination techniques for subtitle generation, the burden of correcting the sound recognition results by combining or segmenting them to make the subtitles easier for humans to read is significant.

[0005] As mentioned above, in conventional subtitle generation devices, when generating subtitles in real time based on sound recognition results, the burden of making corrections to make the subtitles easier to read is heavy. Summary of the Invention

[0006] The problem to be solved by the present invention is to provide a subtitle generation apparatus, method and storage medium that can reduce the burden of corrections to make the subtitles easier to read when generating subtitles in real time based on sound recognition results.

[0007] The subtitle generation apparatus according to the embodiment includes an acquisition unit, a history unit, a generation unit, a history update unit, and a prompting unit. The acquisition unit sequentially acquires text based on sound recognition results. The history unit saves the text as historical data. The generation unit infers the segmentation and combination positions of the text based on one or more saved historical data points, and generates subtitle text based on the segmentation and combination positions and the historical data points. The history update unit updates the historical data based on the segmentation and combination positions. The prompting unit prompts the user with the subtitle text. Attached Figure Description

[0008] Figure 1 This is a block diagram illustrating an example of a subtitle generation apparatus according to the first embodiment.

[0009] Figure 2 This is a schematic diagram used to illustrate the historical data involved in the first embodiment.

[0010] Figure 3 This is a flowchart illustrating an example of the operation in the first embodiment.

[0011] Figure 4 This is a schematic diagram illustrating an example of the operation in the first embodiment.

[0012] Figure 5 This is a block diagram illustrating an example of a subtitle generation apparatus according to the second embodiment.

[0013] Figure 6 This is a flowchart illustrating an example of the operation in the second embodiment.

[0014] Figure 7 This is a schematic diagram illustrating an example of the operation in the second embodiment.

[0015] Figure 8 This is a block diagram illustrating an example of a subtitle generation apparatus according to the third embodiment.

[0016] Figure 9 This is a flowchart illustrating an example of the operation in the third embodiment.

[0017] Figure 10 This is a schematic diagram illustrating an example of the operation in the third embodiment.

[0018] Figure 11 This is a diagram illustrating the hardware structure of the subtitle generation apparatus according to the fourth embodiment.

[0019] (Explanation of symbols)

[0020] 10: Sound recognition unit; 20: Subtitle generation device; 21: History unit; 21a: Historical data; 22: Acquisition unit; 23: History update unit; 24: Subtitle processing unit; 24a: Generation unit; 24b: Regeneration unit; 24c: Correction candidate generation unit; 25: User interface unit; 26: Correction expectation receiving unit; 201: CPU; 202: RAM; 203: Program memory; 204: Auxiliary storage device; 205: Input / output interface; UA1: Sliding motion; UA2: Sliding motion; UA3: Indication motion. Detailed Implementation

[0021] Hereinafter, embodiments will be described with reference to the accompanying drawings. The subtitle generation apparatuses involved in these embodiments include, for example, systems used in the production of subtitles for television news programs, where sound is identified in real time and then manually checked and corrected. Furthermore, "real time" in these embodiments does not mean instantaneous or simultaneous, but includes delays caused by data processing and communication. Additionally, the term "subtitle generation apparatus" may be appropriately changed to any name such as "automatic subtitle generation apparatus."

[0022] <First Embodiment>

[0023] Figure 1 This is a block diagram illustrating the subtitle generation apparatus according to the first embodiment. Figure 1 In this device, the sound recognition unit 10 is connected to the subtitle generation device 20. The subtitle generation device 20 includes a history unit 21, an acquisition unit 22, a history update unit 23, a subtitle processing unit 24, and a user interface unit 25. The subtitle processing unit 24 includes a generation unit 24a.

[0024] Here, the voice recognition unit 10 continuously acquires voice data, performs voice recognition on the acquired voice data, and generates text of the voice recognition results.

[0025] The history section 21 is a memory that can be read from / written to the history update section 23 and the generation section 24a. The history section 21 stores, for example, historical data related to the text of the acquired sound recognition results, and data related to the generation of subtitle text based on the historical data. Historical data 21a, such as... Figure 2 The image shows the text of the sound recognition results successively acquired by the acquisition unit 22, and the information corresponding to the number of output characters and the position of the segmented characters, which are related to the subtitle text generated by the generation unit 24a. In addition, as historical data 21a, for each text of the sound recognition results, in addition to the number of output characters and the position of the segmented characters, it may also include the recognition ID, acquisition time (acquisition date and time), string length (number of characters), etc.

[0026] Identification ID is information used to identify each piece of text obtained sequentially.

[0027] The acquisition time is the time when each text is acquired, including the acquisition date not shown. That is, the acquisition time is an example of the acquisition date and time.

[0028] A text is a string representing the voice recognition results obtained sequentially. For example, a Japanese sentence such as "新型コロナウイルスワクチンの大規模接種センターの予約が今日からスタートしました" is pronounced as "shingata-korona-uirusu-wakuchin-no-daikibo-sessyu-sentaa-no-yoyaku-ga-kyoo-kara-sutaato-shimashita" in Japanese, and its Chinese meaning is "The reservation for the large-scale vaccination center of the new coronavirus vaccine has started today". Regarding this sentence, "新型コロナウイルスワクチン" (utterance; shingata-korona-uirusu-wakuchin: meaning; new coronavirus vaccine), "の大規模接種" (no-daikibo-sessyu: large-scale vaccination), "センターの" (sentaa-no: of the center), "予約が今日" (yoyaku-ga-kyoo: the reservation today), and "からスタートしました" (kara-sutaato-shimashita: has started) are obtained sequentially. In addition, a part of a Japanese sentence such as "東京会場となる合同庁舎" is pronounced as "tookyoo-kaijyoo-tonaru-goodoochoosya" in Japanese, and its Chinese meaning is "The comprehensive government building that will be the Tokyo venue". A part of this sentence is obtained sequentially as "東京会場" (tookyoo-kaijyoo: Tokyo venue) and "となる合同庁舎" (tonaru-goodoochoosya: the comprehensive government building that will be).

[0029] The string length is the length of the string of the text, in other words, the number of characters in the text.

[0030] The number of characters already output is the number of characters of the text that has been output from the generation unit 24a to the user interface unit 25. In Figure 2 the example, the characters already output are "新型コロナウイルスワクチンの大規模接種センターの予約が" (shingata-korona-uirusu-wakuchin-no-daikibo-sessyu-sentaa-no-yoyaku-ga: The reservation for the large-scale vaccination center of the new coronavirus vaccine).

[0031] The split character position is the position of the character that splits (newlines) the text when it is combined with other text.

[0032] As data related to the generation of caption text, the following can be used: current position CSI of display candidates, display candidate arrangement, string length of linked text, maximum number of characters, word recognition WID, feature vector, update information, segmentation prediction score, word segmentation position SWI, etc.

[0033] The acquisition unit 22 sequentially acquires the text of the sound recognition results from the sound recognition unit 10. Alternatively, the subtitle generation device 20 may replace the sound recognition unit 10 and the acquisition unit 22, and be capable of continuously acquiring sound data and performing sound recognition on the acquired sound data, and sequentially sending the text of the sound recognition results to the sound recognition unit of the history update unit 23.

[0034] The history update unit 23 saves the text of the sound recognition results obtained by the acquisition unit 22 as history data 21a to the history unit 21. In addition, the history update unit 23 updates the history data 21a of the history unit 21 based on information related to the subtitle text generated by the generation unit 24a.

[0035] The generation unit 24a infers the segmentation and combination positions of the text based on one or more historical data 21a stored in the history unit 21, and generates subtitle text based on the inferred segmentation and combination positions and the historical data 21a. Alternatively, the generation unit 24a can infer the combination position by combining historical data with text up to a threshold number of characters (maximum number of characters) in the linked text obtained from the combined text. Furthermore, the generation unit 24a can perform one or both of lexical parsing and dependency parsing on the linked text obtained from the combined text, and generate feature vectors for each word based on the execution result. In this case, the generation unit 24a uses the feature vectors to calculate a segmentation prediction score for each word, and infers the segmentation position based on the segmentation prediction score.

[0036] User interface unit 25 prompts the user with the subtitle text generated by generation unit 24a. User interface unit 25 can be, for example, a display showing the subtitle text, or a touch panel that combines the display and input functions. Furthermore, user interface unit 25 can also include a keyboard and mouse for input, in addition to the display. This input function can be appropriately used to correct subtitle text segmentation, combination, and sound recognition errors. User interface unit 25 is an example of a prompting unit.

[0037] Next, use Figure 3 Flowchart and Figure 4 Schematic diagram showing the operation of the subtitle generation device configured as described above.

[0038] (Step ST1)

[0039] The acquisition unit 22 sequentially acquires the text of the voice recognition result. Specifically, the voice recognition unit 10 performs voice recognition on continuously received voice data and generates the text of the voice recognition result as shown Figure 4 . The acquisition unit 22 sequentially acquires the text of the voice recognition result from the voice recognition unit 10. In addition, regarding the acquired text, from the viewpoint of having less time lag, it is preferably the text of the voice recognition result using the above-mentioned early determination technology. However, it is not limited to this, and the acquired text may also be the text of the voice recognition result when using a general voice interval estimation technology.

[0040] In addition, in Figure 4 , the Japanese sentence of "The reservation for the large-scale vaccination center for the new coronavirus vaccine has started today" is as described above. Also, in Figure 4 , the pronunciation of the Japanese sentence of "At the Ote-machi Joint Government Building in Chiyoda Ward, which will be the Tokyo venue, medical officers and nurses of the Self-Defense Forces who will receive the vaccination held a formation completion ceremony" is "tookyoo-kaijyoo-tonaru-chiyodaku-no-ootemachi-goodoochoosya-dewa-sessyu-ni-ataru-jieitai-no-ikan-ya-kangokan-ra-ga-hensei-kanketsushiki-o-okonaimashita", and the Chinese meaning is "At the Ote-machi Joint Government Building in Chiyoda Ward, which will be the Tokyo venue, medical officers and nurses of the Self-Defense Forces who will receive the vaccination held a formation completion ceremony".

[0041] (Step ST2)

[0042] The history update unit 23 saves the acquired text of the voice recognition result as historical data 21a in the history unit 21. In Figure 2In the example shown, for the utterance "The reservation for the large-scale vaccination center of the new coronavirus vaccine has started today. The joint government building in Tokyo" (shingata-korona-uirusu-wakuchin-no-daikibo-sessyu-sentaa-no-yoyaku-ga-kyoo-kara-sutaato-shimashita.tookyoo-kaijyoo-tonaru-goodoochoosya), the text of the voice recognition results obtained sequentially for "new coronavirus vaccine" (shingata-korona-uirusu-wakuchin), "large-scale vaccination" (no-daikibo-sessyu), "center of" (sentaa-no), etc. is saved as historical data 21a. In the historical data 21a, together with the text of the voice recognition results, at least the recognition ID for identifying the historical data 21a, the number of characters already output, and the split character position are retained. When the historical update unit 23 has obtained the voice recognition text, it newly creates data with a new recognition ID, the number of characters already output being 0, and no split character position, and saves it as historical data 21a in the history unit 21. The above steps ST1 to ST2 are repeatedly executed in parallel with the processing after the next step ST3.

[0043] (Step ST3)

[0044] The generation unit 24a extracts one or more pieces of historical data 21a from the historical data 21a in the history unit 21 that have not been presented as subtitle text.

[0045] For example, the arrangement of the historical data 21a in the history unit 21 is taken as the historical data arrangement. Specifically, for example, the historical data arrangement includes seven historical data 21a each identified by an identification ID = 0 to 6. Each historical data 21a is, for example, {identification ID = 0 and "新型コロナウイルスワクチン" (shingata-korona-uirusu-wakuchin)}, {identification ID = 1 and "の大規模接種" (no-daikibo-sessyu)}, {identification ID = 2 and "センターの" (sentaa-no)}, {identification ID = 3 and "予約が今日" (yoyaku-ga-kyoo)}, {identification ID = 4 and "からスタートしました" (kara-sutaato-shimashita)}, {identification ID = 5 and "東京会場" (tookyoo-kaijyoo)}, {identification ID = 6 and "となる合同庁舎" (tonaru-goodoochoosya)}. In addition, it is assumed that the maximum number of characters of the generated subtitle text is predetermined to be 30 characters.

[0046] First, the generation unit 24a initializes the current position CSI of the display candidate in the historical data arrangement, the display candidate arrangement that holds one or more historical data 21a in the historical data arrangement, and the string length of the concatenated text obtained by concatenating the texts of the historical data 21a in the display candidate arrangement.

[0047] In addition, when the size of the historical data arrangement is greater than the current position CSI of the display candidate, that is, when there is historical data 21a in the historical data arrangement that has not been displayed as subtitle text, the generation unit 24a repeats the following.

[0048] First, when the generation unit 24a obtains the CSI-th historical data 21a from the historical data arrangement, it increments the current position CSI of the display candidate by 1 and adds the historical data 21a to the display candidate arrangement.

[0049] Next, the generation unit 24a adds the value obtained by subtracting the number of characters already output of the historical data 21a from the string length of the text of the voice recognition result of the historical data 21a to the string length of the concatenated text. Here, when the string length of the concatenated text is greater than the maximum number of characters of the generated subtitle text, in order to make the concatenated text easier to read, it proceeds to step ST4.

[0050] For example, when the current position CSI of the display candidate is 0, the historical data 21a of the 0th (recognition ID = 0) is retrieved from the historical data arrangement, and the retrieved historical data 21a is appended to the display candidate arrangement. At this time, the string length of the historical data 21a with the recognition ID = 0 is 13, and the number of characters already output is 0, so the string length of the concatenated text becomes 13. Next, the current position CSI of the display candidate is incremented by 1, and similarly, the historical data 21a of the CSI is retrieved and appended to the display candidate arrangement. Thus, when the string length of the concatenated text exceeds the maximum number of characters (30 characters), the historical data 21a with the recognition ID = 0 to 4 is appended to the display candidate arrangement. In addition, the text of each historical data 21a with the recognition ID = 0 to 4 is "新型コロナウイルスワクチン" (shingata-korona-uirusu-wakuchin), "の大規模接種" (no-daikibo-sessyu), "センターの" (sentaa-no), "予約が今日" (yoyaku-ga-kyoo), "からスタートしました" (kara-sutaato-shimashita). Therefore, in this example, the string length of the concatenated text is 39 (= 13 + 6 + 5 + 5 + 10).

[0051] (Step ST4)

[0052] The generation unit 24a estimates the combination position and the division position based on the text of the extracted historical data 21a. For example, the generation unit 24a estimates the combination position in such a way that the number of characters of the concatenated text obtained until the concatenated text becomes equal to or more than the maximum number of characters (threshold number of characters) by combining the texts of the historical data 21a. In addition, for example, the generation unit 24a estimates the division position based on the display candidate arrangement and proceeds to step ST5. Specifically, for example, the division position is estimated based on the context before and after.

[0053] First, the generation unit 24a initializes the concatenated text obtained by concatenating the texts in the display candidate arrangement and the update information of the historical data 21a.

[0054] Next, for each historical data 21a in the display candidate arrangement, the generation unit 24a stores the recognition ID of each historical data and the text start position in the update information, and concatenates the texts until the number of characters becomes equal to or more than the maximum number of characters to create a concatenated text.

[0055] For example, assume that the text for identifying historical data 21a with identification IDs = 0 to 4 is "新型コロナウイルスワクチン" (shingata-korona-uirusu-wakuchin), "の大規模接種" (no-daikibo-sessyu), "センターの" (sentaa-no), "予約が今日" (yoyaku-ga-kyoo), "からスタートしました" (kara-sutaato-shimashita). In this case, {identification ID = 0 and text start position "0"}, {identification ID = 1 and text start position "13"}, {identification ID = 2 and text start position "19"}, {identification ID = 3 and text start position "24"}, {identification ID = 4 and text start position "29"} are saved in the update information. In addition, the text start position corresponds to the string length before each text within the concatenated text. Also, the concatenated text "新型コロナウイルスワクチンの大規模接種センターの予約が今日からスタートしました" (shingata-korona-uirusu-wakuchin-no-daikibo-sessyu-sentaa-no-yoyaku-ga-kyoo-kara-sutaato-shimashita) is created. As described above, the string length of this concatenated text is 39, and the text of the voice recognition result is concatenated until it becomes 30 characters or more, which is the maximum number of characters.

[0056] Next, the generation unit 24a performs either or both of morphological analysis and dependency analysis on the concatenated text obtained by combining the texts, and generates feature vectors for each word based on the execution result. In addition, the generation unit 24a uses the feature vectors to calculate the segmentation speculation score for each word, and speculates the segmentation position based on the segmentation speculation score.

[0057] Specifically, for example, the generation unit 24a performs morpheme analysis on the linked text, and generates a word arrangement that associates each word with a word recognition WID based on the result of the morpheme analysis. For example, as the word arrangement, {word recognition WID = 0 and "新型コロナウイルス" (shingata-korona-uirusu: novel coronavirus)}, {word recognition WID = 1 and "ワクチン" (wakuchin: vaccine)}, {word recognition WID = 2 and "の" (no: of)}, {word recognition WID = 3 and "大規模" (daikibo: large-scale)}, {word recognition WID = 4 and "接種" (sessyu: vaccination)}, {word recognition WID = 5 and "センター" (sentaa: center)}, {word recognition WID = 6 and "の" (no: of)}, {word recognition WID = 7 and "予約" (yoyaku: reservation)}, {word recognition WID = 8 and "が" (ga: is)}, {word recognition WID = 9 and "今日" (kyoo: today)}, {word recognition WID = 10 and "から" (kara: from)}, {word recognition WID = 11 and "スタート" (sutaato: start)}, {word recognition WID = 12 and "し" (shi: has)}, {word recognition WID = 13 and "まし" (mashi: become)}, {word recognition WID = 14 and "た" (ta: ended)} are generated.

[0058] Next, for the word arrangement, the generation unit 24a uses the text start position held in the update information to associate the recognition ID of each history data 21a with the word recognition WID and saves it to the update information.

[0059] For example, when recognizing historical data 21a with IDs = 0 to 4 ("shingata-korona-uirusu" (new coronavirus vaccine), "no-daikibo-sessyu" (mass vaccination), "sentaa-no" (of the center), "yoyaku-ga-kyoo" (reservations started today), "kara-sutaato-shimashita" (started from today)), the following is saved in the update information. {Recognize ID = 0 and "shingata-korona-uirusu" (WID = 0) (new coronavirus), "wakuchin" (WID = 1) (vaccine)}, {Recognize ID = 1 and "no" (WID = 2), "daikibo" (WID = 3) (large scale), "sessyu" (WID = 4) (vaccination)}, {Recognize ID = 2 and "yoyaku" (WID = 7) (reservation), "ga" (WID = 8), "kyoo" (WID = 9) (today)}, {Recognize ID = 3 and "sentaa" (WID = 5) (center), "no" (WID = 6)}, {Recognize ID = 4 and "kara" (WID = 10) (from), "sutaato" (WID = 11) (start), "shi" (WID = 12), "mashi" (WID = 13), "ta" (WID = 14)}.

[0060] Next, the generation unit 24a generates feature vectors for each word for the word arrangement. Here, as features used in the feature vectors, for example, the word index from the beginning of the article, character length, word class, dependency parsing result, etc. of each word can be appropriately used.

[0061] Next, the generation unit 24a calculates a segmentation speculation score for whether to perform segmentation before each word through the learned model based on the generated feature vectors. Here, as the learned model, for example, a deep learning model can be appropriately used. As the deep learning model, for example, an N-gram-based deep neural network (DNN) model can also be used. In this case, the generation unit 24a can also use the feature vectors of the N words before and after each word as one vector, input this vector into the DNN model, and calculate the segmentation speculation score for each word through this DNN model.

[0062] Next, the generation unit 24a determines the position of the word to be segmented, that is, the word segmentation position SWI, based on the segmentation speculation score. For example, the position before the word with the maximum segmentation speculation score is set as the word segmentation position SWI. Or, the position before the top several words above the threshold is set as the word segmentation position SWI.

[0063] Specifically, for example, let the word segmentation position SWI = [3, 9]. In this case, it means that the word segmentation position SWI is the position where segmentation is performed before the word "daikibo" (large scale) with word recognition WID = 3 and before the word "kyoo" (today) with word recognition WID = 9. Supplementary, SWI = [3] in the word segmentation position SWI = [3, 9] indicates between the word "no" (of) with word recognition WID = 2 and the word "daikibo sesshu sentaa" (large scale vaccination center) with word recognition WID = 3. Similarly, SWI = [9] in SWI = [3, 9] indicates between the word "ga" (is) with word recognition WID = 8 and the word "kyoo" (today) with word recognition WID = 9.

[0064] Next, when the word arrangement and the word segmentation position SWI are obtained, it proceeds to the step ST5 of generating subtitle text.

[0065] (Step ST5)

[0066] The generation unit 24a generates subtitle text based on the extracted historical data 21a and the speculated combination position and segmentation position. In addition, the generation unit 24a creates update information required for updating each historical data 21a. For example, the generation unit 24a extracts the update information in which each word appears by starting from the word start position of the word index of each word and the historical data 21a of the update information, and adds the output character count to the corresponding update information, thereby creating the update information.

[0067] For example, let the above-mentioned word segmentation position SWI = [3, 9]. The generation unit 24a creates update information that holds the subtitle text "shingata-korona-uirusu-wakuchin-no(¥n)daikibo-sessyu-sentaa-no-yoyaku-ga" (reservation for the new coronavirus vaccine large scale vaccination center of) and the recognition ID and output character count of each historical data 21a based on the display candidate arrangement of the historical data 21a including recognition IDs = 0 to 4. In addition, "¥n" indicates a line break at the segmentation position. In addition, (¥n) within the phonetic notation is not pronounced, but is described corresponding to the line break of the subtitle text. The update information includes, for example, {recognition ID = 0 and output character count 13}, {recognition ID = 1 and output character count 6}, {recognition ID = 2 and output character count 4}, {recognition ID = 3 and output character count 3}, {recognition ID = 4 and output character count 0}.

[0068] (Step ST6)

[0069] User interface unit 25 displays the generated subtitle text to the user. For example, user interface unit 25... Figure 4 As shown, up to two lines of subtitle text are displayed sequentially. Additionally, the generation unit 24a initializes the candidate arrangement and the string length of the linked text.

[0070] (Step ST7)

[0071] The history update unit 23 updates the historical data 21a stored in the history unit 21 based on the generated subtitle text. For example, the history update unit 23 updates the historical data arrangement based on the update information of the historical data 21a. In addition, the history update unit 23 updates the current position CSI of the displayed candidates to the position where there is no displayed historical data 21a based on the update information.

[0072] For example, regarding the current position CSI of the candidate, in the case of the update information described in step ST5, the historical data 21a with ID=3 is updated because the string length of the text "5" is inconsistent with the number of characters output "3".

[0073] (Step ST9)

[0074] The generation unit 24a determines whether there is any unprompted historical data 21a. If there is unprompted historical data 21a, steps ST3 to ST9 are repeatedly executed. On the other hand, if the determination result is that there is no unprompted historical data 21a, the processing ends.

[0075] As described above in the first embodiment, the acquisition unit 22 sequentially acquires the text obtained from the sound recognition results. The history unit 21 saves this text as history data 21a. The generation unit 24a infers the segmentation and combination positions of the text based on one or more saved history data 21a. Furthermore, the generation unit 24a generates subtitle text based on the segmentation and combination positions and the one or more history data 21a. The history update unit 23 updates the history data 21a based on the segmentation and combination positions. The user interface unit 25, which serves as a prompting unit, displays the subtitle text. Therefore, when generating subtitles in real time based on the sound recognition results, the burden of making corrections to make the subtitles easier to read can be reduced.

[0076] To elaborate further, real-time subtitle creation utilizing sound recognition technology can automatically generate readable subtitle text, thus reducing revision costs for subtitle creators. Furthermore, using sound recognition results that utilize early-determined technology allows for more time to be devoted to revisions.

[0077] Furthermore, comparative examples (1) to (3) for such a first embodiment will be described.

[0078] (1) The first comparative example is a technique for predicting line break positions from already written text to improve the readability of subtitles. In the technique of the first comparative example, to make the subtitles of a speech easier to read, a recurrent neural network (RNN), a deep learning technique, is used to predict whether a line break should be inserted before each phrase in the already written speech text. However, in the first comparative example, the line break positions are predicted entirely from the already written speech text, so it is not suitable for generating subtitles in real time. Specifically, the technique of the first comparative example is not suitable for obtaining recognition results one by one, displaying them, and manually correcting them, so it is difficult to apply to the real-time subtitle production of news programs, etc.

[0079] (2) The second comparative example is a technique for adjusting the output reproduction of audio data based on the number of characters in the text obtained from audio recognition in order to create subtitles. This second comparative example is unrelated to the readability of the subtitles produced. Therefore, in the second comparative example, manual correction is required to make the subtitles take readability into account, so the cost of manual correction is high.

[0080] (3) The third comparative example is a technique that uses a touch panel or similar device to specify the correction points and reproduce the speech after correcting the sound recognition results. However, the third comparative example is not a structure that performs corrections sequentially and continuously, so it cannot be applied to real-time subtitle production.

[0081] Thus, none of the comparative examples (1) to (3) are suitable for real-time generation of easy-to-read subtitles, and therefore the effects of the first embodiment described above cannot be obtained.

[0082] Furthermore, according to the first embodiment, the generation unit 24a estimates the joining position by combining historical data text up to a threshold number of characters (maximum number of characters) in the linked text obtained by combining the combined text. Therefore, in addition to the effects described above, since the segmentation position is estimated based on linked text with a character count slightly exceeding the threshold number of characters, it is also expected that subtitle text with a character count close to the threshold number of characters can be generated. That is, a large amount of subtitle text can be generated within the displayable range, and a reduction in the subtitle display delay time can be expected.

[0083] Furthermore, according to the first embodiment, the generation unit 24a performs one or both of lexical analysis and dependency analysis on the linked text obtained from the combined text, and generates feature vectors for each word based on the execution result. Additionally, the generation unit 24a uses the feature vectors to calculate a segmentation prediction score for each word, and predicts the segmentation position based on the segmentation prediction score. Therefore, in addition to the effects described above, since the segmentation position can be predicted from the linked text without manual intervention, the burden of correcting subtitles to make them easier to read can be further reduced.

[0084] (A variation of the first embodiment)

[0085] The first embodiment can also be implemented as in the following variations. These variations can also be applied to the following embodiments in the same way.

[0086] That is, feature vectors were used in the first embodiment, but it is not limited to this. For example, word vectors utilizing word embedding methods can be used instead of feature vectors. In this case, the generation unit 24a performs one or both of lexical parsing and dependency parsing on the linked text obtained from the combined text, and generates embedding vectors for each word based on the execution result. The embedding vectors are used to calculate the segmentation prediction score for each word, and the segmentation position is predicted based on the segmentation prediction score. This can also achieve the same effect as the first embodiment.

[0087] Next, an N-gram DNN model was used in the first embodiment, but it is not limited to this. For example, a recurrent neural network (RNN) model can be used instead of an N-gram DNN model. In this case, an RNN model that uses the feature vectors of each word as input to calculate the segmentation prediction score of each word can be used. In addition, the RNN model can be a unidirectional structure that considers the context on one side of the word of interest, or a bidirectional structure that considers the context on both sides. Specifically, for example, the generation unit 24a can also use an LSTM model that considers the context before each word, the context after each word, or the context before and after each word when calculating the segmentation prediction score of each word. Furthermore, LSTM is short for "Long Short-Term Memory".

[0088] Furthermore, while the first embodiment uses word-based processing for generating subtitle text, it is not limited to this. For example, instead of word-based processing, phrases obtained by summarizing words can be used as processing units, and feature vectors obtained by summarizing phrase units can be used. In this case, the generation unit 24a performs one or both of lexical analysis and dependency analysis on the linked text obtained by combining text, summarizes the results of the execution by phrase units to generate feature vectors for each phrase, uses the feature vectors to calculate the segmentation prediction score for each phrase, and predicts the segmentation position based on the segmentation prediction score. This achieves the same effect as the first embodiment.

[0089] Furthermore, in the first embodiment, unnecessary words such as filler characters or symbols, or specific strings, are not deleted when generating subtitle text, but this is not a limitation. That is, it is also possible to delete these unnecessary words when generating subtitle text. In this case, the generation unit 24a performs the process of deleting unnecessary words after lexical analysis, so that the position of the deleted words in the text is included in the updated data and reflected in each historical data during historical updates.

[0090] Furthermore, in the first embodiment, a segmentation prediction score is calculated during the process of generating subtitle text, and the segmentation position is predicted based on the segmentation prediction score, but this is not limited to this. For example, it may be configured such that the segmentation position is predetermined during the process of generating subtitle text, and segmentation is performed at the predetermined fixed segmentation position.

[0091] Furthermore, in the first embodiment, the generation unit 24a infers the joining position by combining the text of historical data 21a up to a threshold number of characters (maximum number of characters) in the linked text obtained by combining the combined text, but is not limited to this. For example, the generation unit 24a may also infer the segmentation position for the linked text obtained by combining the combined text based on a fixed number of characters different from the threshold number of characters. The fixed number of characters is, for example, any number of characters smaller than the maximum number of characters.

[0092] Furthermore, in the first embodiment, the segmentation position is inferred regardless of the amount of historical data not displayed in the historical data arrangement, but this is not a limitation. That is, the model and processing for inferring the segmentation position can be changed according to the amount of historical data not displayed in the historical data arrangement. In addition, when the amount of historical data 21a increases, the generation unit 24a preferably prompts the user as early as possible, so it infers the segmentation position based on a fixed number of characters with low computational complexity. That is, when the amount of historical data 21a stored in the historical data 21a of the history unit 21, including text not used as subtitle text prompts, is above a threshold, the generation unit 24a infers the segmentation position based on a fixed number of characters. On the other hand, when the amount of historical data 21a is below a certain value, the generation unit 24a uses a segmentation inference model with high accuracy to calculate the segmentation inference score, although the computational complexity is high. That is, when the amount of historical data 21a stored in the historical data 21a of the history unit 21, including text not used as subtitle text prompts, is less than a threshold, the generation unit 24a infers the segmentation position by calculating the segmentation inference score.

[0093] Alternatively, in the first embodiment, when estimating the segmentation position, it can be estimated within a range longer than the area displayed on a single screen and including both the preceding and following sections. Specifically, for example, in determining whether to add historical data 21a to the candidate display arrangement, the candidate display arrangement is made to include historical data 21a until the number of characters in the linked text exceeds the maximum number of characters displayed on a single screen by adding a few characters (+α characters), and the segmentation position is estimated in the same way as described above. In this case, when generating the subtitle text, the segmentation position can be estimated in such a way that the previously displayed portion of the initial string and the portion exceeding the maximum number of characters in the final string are not included in the subtitle text.

[0094] In addition, in the first embodiment, as a method for determining the addition of historical data 21a to the display candidate arrangement, a method based on the maximum number of characters displayed on one screen is shown, but it is not limited to this. For example, instead of determining based on the maximum number of characters, a method can be used as follows: the date and time of obtaining the text of the voice recognition result is associated with the text and saved to the history unit 21, and historical data that will not be displayed even after a certain period of time has elapsed since the date and time of obtaining the text is added to the display candidate arrangement. In this case, the historical data 21a includes the text of the voice recognition result and the date and time of obtaining the text. The generation unit 24a infers the combination position by combining the text of the historical data 21a that will not be displayed even after a certain period of time has elapsed since the date and time of obtaining the text. The certain period of time is, for example, 10 seconds. However, it is not limited to this, and any time in the range of 1 to 10 seconds can be appropriately used as the certain period of time.

[0095] <Second Implementation>

[0096] The second embodiment, compared to the first embodiment, includes the following structure: it is used to correct the segmentation and combination positions of the subtitle text based on the user's action (operation) on the subtitle text portion displayed on the user interface section 25. Furthermore, the action may also be referred to as an operation.

[0097] Figure 5 This is a block diagram illustrating the subtitle generation apparatus according to the second embodiment. Figure 1 The same components are marked with the same symbol, and their detailed descriptions are omitted. The main difference is described here. The following embodiments also omit repeated descriptions.

[0098] Figure 5 The subtitle generation device 20 shown is compared to Figure 1 As shown in the structure, the subtitle processing unit 24 also includes a regeneration unit 24b, and a correction expectation receiving unit 26 is provided between the regeneration unit 24b and the user interface unit 25.

[0099] In addition to the above-described structure, the user interface unit 25 also detects user actions on a portion of the displayed subtitle text and generates action data. This user interface unit 25 is an example of a detection unit.

[0100] The correction expectation receiving unit 26 determines the scope of the subtitle text to be corrected and the expectation type for any aspect of the segmentation and combination of the subtitle text based on the motion data. Here, the correction expectation receiving unit 26 performs this determination, for example, upon receiving motion data as a correction expectation. This correction expectation receiving unit 26 is an example of a determination unit.

[0101] The regeneration unit 24b re-infers the segmentation or combination position of the subtitle text based on the scope of the correction object, the expected type, and the historical data 21a, and regenerates the subtitle text based on the re-infermentation result.

[0102] use Figure 6 Flowchart and Figure 7 The diagram illustrates the operation of the subtitle generation device configured as described above. Furthermore, the operation of this embodiment is performed as step ST8, which is between steps ST7 and ST9 described above. Step ST8 is performed as steps ST8-1 to ST8-6 below.

[0103] (Step ST8-1)

[0104] When the user performs an action on a portion of the subtitle text displayed in step ST6, the user interface unit 25 detects the action and generates action data. Specifically, this is assuming the user performs an action such as moving the cursor, clicking, dragging, flicking, or swiping on the displayed subtitle text. In this case, the user interface unit 25 generates information such as the start and end positions of the action on the screen as action data.

[0105] (Step ST8-2)

[0106] User interface unit 25 notifies the correction expectation receiving unit 26 of the action data. Correction expectation receiving unit 26 receives the action data.

[0107] (Step ST8-3)

[0108] Next, the correction expectation receiving unit 26 determines the scope of the subtitle text to be corrected, and which side of the subtitle text segmentation and combination the expectation type is directed towards, based on the motion data.

[0109] (Step ST8-4)

[0110] The regeneration unit 24b re-speculates the segmentation position or combination position of the subtitle text based on the correction target range, desired type, and historical data 21a of the subtitle text, and regenerates the subtitle text according to the re-speculation result.

[0111] (Step ST8-5)

[0112] The user interface unit 25 prompts the user with the regenerated subtitle text.

[0113] (Step ST8-6)

[0114] The history update unit 23 updates the historical data 21a stored in the history unit 21 according to the regenerated subtitle text. Thus, the subtitle generation device 20 ends step ST8 composed of steps ST8-1 to ST8-6 and transfers to the above step ST9.

[0115] Next, regarding the above actions, a specific example of the desired types of segmentation and combination for subtitle text will be described with reference to Figure 7 the schematic diagram. In the following description, the user interface unit 25 is installed as a touch panel, but it is not limited thereto. This is the same in each of the following embodiments. In addition, an example of segmentation into times t1 to t3 and an example of combination into times t4 to t6 are shown.

[0116] (Time t1)

[0117] At time t1, the user interface unit 25 displays on the screen the subtitle text "新型コロナウイルスワクチンの¥n大規模接種センターの予約が今日から" (shingata-korona-uirusu-wakuchin-no(¥n)daikibo-sessyu-sentaa-no-yoyaku-ga-kyoo-kara) in two lines. At this time, for a part of the subtitle text being displayed, "今日から" (kyoo-kara), it is assumed that the user performed a diagonal drag action UA1 of dragging from the upper right to the lower left while touching with a finger.

[0118] The user interface unit 25 detects this diagonal drag action UA1, obtains the start position and end position of the diagonal drag action UA1, and generates action data including the start position and end position (step ST8-1). Then, the user interface unit 25 notifies the correction expectation receiving unit 26 of the action data.

[0119] The correction expectation receiving unit 26 receives the action data (step ST8-2) and determines whether it is a diagonal drag action based on the start position and end position of the action data (step ST8-3). Specifically, it is determined as a diagonal drag action when the start position of the action data is in the upper right and the end position is in the lower left.

[0120] In the case where it is determined as a diagonal drag action, the correction expectation receiving unit 26 determines the correction target range of the subtitle text based on the display position of the subtitle text and the peripheries of the start position and the end position of the action data. For example, the correction expectation receiving unit 26 calculates the center position based on the start position and the end position of the action data, obtains the character position of the subtitle text closest to this center position, and sets several characters before and after this character position as the correction target range. In addition, in the case where it is determined as a diagonal drag action, the correction expectation receiving unit 26 sets the expectation type as the split type, and sends the expectation type and the correction target range to the regeneration unit 24b.

[0121] When the expectation type is the split type, the regeneration unit 24b speculates a reasonable split position within the correction target range based on the obtained correction target range and the history data 21a, and regenerates the subtitle text (step ST8-4).

[0122] Specifically, for the obtained correction target range "の予約が今日から" (no-yoyaku-ga-kyoo-kara), the regeneration unit 24b calculates the split speculation score of each word in the same way as above, and within the correction target range, speculates the position with the highest split speculation score after removing the current split position as the split position. In this example, the position before the subtitle text "今日から" (kyoo-kara) is speculated as the split position. In addition, when the number of lines after splitting (3 lines) exceeds the number of lines that can be displayed in one screen (2 lines), the regeneration unit 24b deletes the excess part (the "今日から" (kyoo-kara) in the third line) from the subtitle text being displayed, and creates update information for the history data 21a. Also, although not in this example, when the number of characters after splitting exceeds the number of characters that can be displayed in one line, the regeneration unit 24b deletes the excess part from that line, and creates update information for the history data 21a. For example, among the subtitle texts in two lines being displayed, when a part of the subtitle text in the first line is split and transferred to the second line, sometimes the subtitle text in the second line exceeds the number of characters that can be displayed in one line. In such a case, the excess part can be deleted from the subtitle text in the second line.

[0123] In addition, based on the speculated split position (before "kyoo-kara"), the regeneration unit 24b regenerates the subtitle text according to the historical data 21a corresponding to the subtitle text "shingata-korona-uirusu-wakuchin-no (¥n) daikibo-sessyu-sentaa-no-yoyaku-ga-kyoo-kara" being displayed. In this case, the regeneration unit 24b regenerates the subtitle text "shingata-korona-uirusu-wakuchin-no (¥n) daikibo-sessyu-sentaa-no-yoyaku-ga".

[0124] (Time t2)

[0125] Then, the user interface unit 25 presents the regenerated subtitle text "shingata-korona-uirusu-wakuchin-no (¥n) daikibo-sessyu-sentaa-no-yoyaku-ga" to the user (step ST8-5).

[0126] The history update unit 23 updates the historical data 21a stored in the history unit 21 according to the regenerated subtitle text (step ST8-6). For example, the history update unit 23 updates the arrangement of the historical data according to the update information of the historical data 21a.

[0127] (Time t3)

[0128] After the display of the subtitle text displayed at time t2 ends, the user interface unit 25 presents the newly generated subtitle text "kyoo-kara-sutaato-shimashita" to the user. In the subtitle text shown at time t3, at the beginning, it includes the subtitle text "kyoo-kara" that was deleted along with the diagonal drag action UA1 at time t1. In addition, when assuming that the case without the diagonal drag action UA1 is set as time t3p, at time t3p, the subtitle text "sutaato-shimashita" that does not include the subtitle text "kyoo-kara" at the beginning will be presented.

[0129] The above is an example of splitting the subtitle text at times t1 to t3. Next, an example of combining the subtitle text at times t4 to t6 will be described.

[0130] (At time t4)

[0131] At time t4, the user interface unit 25 displays on the screen the subtitle text “In the Daiichi Building in Otemachi, Chiyoda Ward, which will be the Tokyo venue, vaccination is underway” in two lines. At this time, it is assumed that the user has performed a sliding operation UA2 of sliding from right to left while touching with a finger in the blank area to the right of a part “··· of Chiyoda Ward” of the subtitle text being displayed.

[0132] The user interface unit 25 detects this sliding operation UA2, obtains the start position and end position of the sliding operation UA2, and generates motion data including the start position and end position (step ST8-1). Then, the user interface unit 25 notifies the correction expectation receiving unit 26 of the motion data.

[0133] The correction expectation receiving unit 26 receives the motion data (step ST8-2), and determines whether it is a sliding operation based on the start position and end position of the motion data (step ST8-3). Specifically, when the height of the start position of the motion data and the height of the end position are substantially the same, and the start position is to the right of the end position, it is determined as a sliding operation.

[0134] When it is determined as a sliding operation, the correction expectation receiving unit 26 determines the correction object range of the subtitle text based on the display position of the subtitle text and the start position of the motion data. For example, the correction expectation receiving unit 26 sets several characters before and after the character position of the subtitle text closest to the start position of the motion data as the correction object range. In addition, when it is determined as a sliding operation, the correction expectation receiving unit 26 sets the expectation type as the combination type, and notifies it to the regeneration unit 24b together with the correction object range.

[0135] When the expectation type is the combination type, the regeneration unit 24b speculates the combination position and the split position after combination based on the obtained correction object range and the history data 21a, and regenerates the subtitle text (step ST8-4).

[0136] Specifically, the regeneration unit 24b re-speculates the current segmentation position included in the historical data 21a corresponding to the obtained correction target range "at the Ote-machi Joint Government Office in Chiyoda Ward, which is the venue" (jyoo-tonaru-chiyodaku-no-ootemachi-goodoochoosya-dewa) as the combination position. In addition, the regeneration unit 24b calculates the segmentation speculation score of each word for the correction target range "at the Ote-machi Joint Government Office in Chiyoda Ward, which is the venue" (jyoo-tonaru-chiyodaku-no-ootemachi-goodoochoosya-dewa) in the same manner as described above, and re-speculates the position with the highest segmentation speculation score after removing the current segmentation position in the correction target range as the segmentation position after combination. In this example, the position before the subtitle text "dewa" is re-speculated as the segmentation position. In addition, although not in this example, when the number of characters after segmentation exceeds the number of characters that can be displayed in one line, the excess part is deleted from that line, and update information of the historical data 21a is created. For example, assume that the number of characters in the subtitle text of the first line among the two lines of subtitle text being displayed is close to the number of characters that can be displayed in one line. In this case, when calculating the segmentation speculation score of each word in the correction target range according to the expectation of the combination type for the subtitle text of the first line, the number of characters in the subtitle text of the first line may exceed the number of characters that can be displayed in one line at the speculated segmentation position. In such a case, the excess part can be deleted from the subtitle text of the first line and moved to the subtitle text of the second line.

[0137] Then, based on the re-speculated combination position, the regeneration unit 24b combines the subtitle text according to the historical data 21a corresponding to the subtitle text "at the Ote-machi Joint Government Office in Chiyoda Ward, which is the Tokyo venue, for vaccination" (tookyoo-kaijyoo-tonaru-chiyodaku-no(¥n)ootemachi-goodoochoosya-dewa-sessyu-ni-ataru) currently being displayed.

[0138] Next, the regeneration unit 24b regenerates the subtitle text based on the re-speculated segmentation position (before "では" (dewa)) according to the historical data 21a corresponding to the combined subtitle text "東京会場となる千代田区の大手町合同庁舎では接種に当たる" (tookyoo-kaijyoo-tonaru-chiyodaku-no-ootemachi-goodoochoosya-dewa-sessyu-ni-ataru). In this case, the regeneration unit 24b regenerates the subtitle text "東京会場となる千代田区の大手町合同庁舎¥nでは接種に当たる" (tookyoo-kaijyoo-tonaru-chiyodaku-no-ootemachi-goodoochoosya(¥n)dewa-sessyu-ni-ataru).

[0139] (Time t5)

[0140] Then, the user interface unit 25 presents the regenerated subtitle text "東京会場となる千代田区の大手町合同庁舎¥nでは接種に当たる" (tookyoo-kaijyoo-tonaru-chiyodaku-no-ootemachi-goodoochoosya(¥n)dewa-sessyu-ni-ataru) to the user (step ST8-5).

[0141] The history update unit 23 updates the historical data 21a stored in the history unit 21 according to the regenerated subtitle text (step ST8-6). For example, the history update unit 23 updates the arrangement of the historical data according to the update information of the historical data 21a.

[0142] (Time t6)

[0143] After the display of the subtitle text displayed at time t5 ends, the user interface unit 25 presents the newly generated subtitle text "自衛隊のいかんや看護官らが¥n編成完結式を行ないました" (jieitai-no-ikan-ya-kangokan-ra-ga(¥n)hensei-kanketsushiki-o-okonaimashita: The medical officers and nurses of the Self-Defense Forces held a completion ceremony for the formation) to the user.

[0144] The above is an example of combining subtitle texts at times t4 to t6. In this way, it is possible to segment or combine subtitle texts according to the actions of the user.

[0145] As described above, according to the second embodiment, the user interface unit 25, which is the detection unit, detects the user's actions on a portion of the prompted subtitle text and generates action data. The correction expectation receiving unit 26, which is the determination unit, determines the scope of correction for the subtitle text and the expected type for either the segmentation or combination of the subtitle text based on the action data. The regeneration unit 24b re-estimates the segmentation or combination position of the subtitle text based on the correction scope, the expected type, and historical data, and regenerates the subtitle text based on the re-estimated result. Therefore, in addition to the effects of the first embodiment, the segmentation and combination positions of the prompted subtitle text can be easily corrected based on the user's actions.

[0146] (A variation of the second embodiment)

[0147] Furthermore, in the second embodiment, the operation of dragging from the upper right to the lower left is detected as a diagonal dragging action UA1, but it is not limited to this. For example, the operation of dragging from the upper left to the lower right can also be detected as a diagonal dragging action UA1. In addition, for example, the operation of dragging vertically from top to bottom can also be detected as a diagonal dragging action UA1. That is, if it is an action indicating a division position, any action can be set as a diagonal dragging action UA1. In addition, if it is an action indicating a division position, other terms can be used instead of diagonal dragging action.

[0148] Furthermore, in the second embodiment, the operation of sliding from right to left is detected as a sliding action UA2, but it is not limited to this. For example, the operation of flicking from right to left (lightly brushing) can also be detected as a sliding action UA2. Additionally, for example, the operation of narrowing two fingers that are spread out to the left and right while in a touching state (pinching) can also be detected as a sliding action UA2. That is, if it is an action indicating a joining position, any action can be used as a sliding action UA2. Furthermore, if it is an action indicating a joining position, other terms can be used instead of the term "sliding action".

[0149] (Third Implementation)

[0150] The third embodiment, compared to the second embodiment, includes the following structure: it provides suggestions for correcting subtitle text based on the user's actions regarding the subtitle text displayed on the user interface section 25. However, the third embodiment is not limited to this, and this structure may also be added to the first embodiment.

[0151] Figure 8 This is a block diagram illustrating the subtitle generation apparatus according to the third embodiment.

[0152] Figure 8 The subtitle generation device 20 shown is compared to Figure 5As shown in the structure, the subtitle processing unit 24 also includes a correction candidate generation unit 24c.

[0153] In conjunction with this, the user interface unit 25 detects the user's actions on a portion of the displayed subtitle text, as described above, and generates action data. This user interface unit 25 is an example of a detection unit.

[0154] The correction expectation receiving unit 26 determines the scope of the subtitle text to be corrected based on the motion data. Here, the correction expectation receiving unit 26 performs this decision, for example, upon receiving motion data as a correction expectation. This correction expectation receiving unit 26 is an example of a decision-making unit.

[0155] The correction candidate generation unit 24c generates correction candidates for the correction target range based on the correction target range and historical data 21a, and infers the segmentation position of the subtitle text including the correction candidate.

[0156] use Figure 9 Flowchart and Figure 10 The diagram illustrates the operation of the subtitle generation device configured as described above. Furthermore, the operation of this embodiment is performed as step ST8A between steps ST7 and ST9. Step ST8A is performed as steps ST8A-1 to ST8A-7, and ST8-3 to ST8-6 below.

[0157] (Steps ST8A-1 to ST8-2)

[0158] Steps ST8A-1 to ST8A-2 are performed in the same manner as steps ST8-1 to ST8-2 described above. That is, the user interface unit 25 detects the user's action on a portion of the displayed subtitle text and generates action data. The correction expectation receiving unit 26 receives this action data from the user interface unit 25.

[0159] (Step ST8A-2-1)

[0160] The correction expectation receiving unit 26 determines whether the motion data is an instruction to prompt a correction candidate based on the motion data. If not, it executes steps ST8-3 to ST8-6 as described in the second embodiment and ends step ST8A, then proceeds to step ST9. On the other hand, if an instruction to prompt a correction candidate is received, it proceeds to step ST8A-3. Furthermore, step ST8A-2-1 is omitted if the processing of the second embodiment cannot be performed.

[0161] (Step ST8A-3)

[0162] Next, the correction expectation receiving unit 26 determines the scope of correction targets for the subtitle text based on the motion data and notifies the correction candidate generation unit 24c of the scope of correction targets.

[0163] (Step ST8A-4)

[0164] The correction candidate generation unit 24c generates correction candidates for the correction target range based on the correction target range of the subtitle text and the history data 21a.

[0165] (Step ST8A-5)

[0166] For each generated correction candidate, the correction candidate generation unit 24c estimates the segmentation position and combination position of the subtitle text when the correction candidate is selected, and generates candidate subtitle texts based on the segmentation position and combination position. Then, the correction candidate generation unit 24c sends the correction candidates and the candidate subtitle texts to the user interface unit 25. At this time, one or more correction candidates and the candidate subtitle texts corresponding to each correction candidate are sent.

[0167] (Step ST8A-6)

[0168] The user interface unit 25 presents one or more generated correction candidates to the user.

[0169] (Step ST8A-7)

[0170] When the user interface unit 25 selects any candidate among the correction candidates according to the user's action, the user interface unit 25 presents the candidate subtitle text corresponding to the selected correction candidate to the user. Then, the history update unit 23 updates the history data 21a stored in the history unit 21 based on the presented candidate subtitle text. Thus, the subtitle generation device 20 ends Step ST8A and transfers to the above-mentioned Step ST9.

[0171] Next, regarding the above actions, a specific example of presenting and correcting correction candidates for subtitle text will be described with reference to Figure 10 the schematic diagram shown below. In the following description, an example of correcting subtitle text at times t6 to t8 is shown.

[0172] (Time t6)

[0173] At time t6, the user interface unit 25 displays on the screen the two-line subtitle text "jieitai-no-ikan-ya-kangokan-ra-ga(¥n)hensei-kanketsushiki-o-okonaimashita" (Self-Defense Force's situation and nurses etc. completed the (¥n) formation completion ceremony). At this time, it is assumed that the user performed an instruction action UA3 of pressing and then releasing a finger on a part "no" of the displayed subtitle text.

[0174] User interface unit 25 detects the instruction action UA3, obtains the instruction position of instruction action UA3 within the screen, and generates motion data including the instruction position (step ST8A-1). Then, user interface unit 25 notifies the correction expectation receiving unit 26 of the motion data.

[0175] The correction expectation receiving unit 26 receives motion data (step ST8A-2) and determines whether it is an indication (indication action) for prompting a correction candidate based on the indication position of the motion data (step ST8A-2-1). Specifically, if the indication position of the motion data is a point, it is determined to be an indication action. In other words, if the start position and end position of the motion data are the same, it is determined to be an indication action.

[0176] If the action is determined to be an instruction action, the correction expectation receiving unit 26 determines the scope of correction targets for the subtitle text based on the action data and notifies the correction candidate generation unit 24c of the scope of correction targets.

[0177] For example, the correction expectation receiving unit 26 estimates the interval with the highest error rate around the indication position where "の" (no) is indicated, and sets the string "のいかんや" (no-ikan-ya) in that interval as the correction target range (step ST8A-3). However, it is not limited to this; the correction expectation receiving unit 26 may also set the characters before and after the character position of the subtitle text displayed closest to the indication position as the correction target range. In addition, when the correction expectation receiving unit 26 determines that an indication action has been performed, it sends the correction target range "のいかんや" (no-ikan-ya) to the correction candidate generation unit 24c.

[0178] Next, the correction candidate generation unit 24c obtains the history data 21a corresponding to the subtitle text being displayed from the history unit 21. In addition, the correction candidate generation unit 24c masks the part of the subtitle text being displayed that becomes the correction target range, and then uses the pre-trained model BERT to infer (generate) candidate strings (correction candidates) for the correction target range (step ST8A-4). Here, BERT is an abbreviation for "Bidirectional Encoder Representations from Transformer". In this example, in the correction candidate generation unit 24c, as correction candidates for the correction target range "のいかんや" (no-ikan-ya), it infers "の医官や" (no-ikan-ya) (replaced with the Japanese word "医官" (ikan: medical officer)), "の" (no) (deleting "いかんや"), and "の、いかんや" (no,ikan-ya) (insertion of "、"). In addition, the Japanese "、" is not pronounced, just like the English "," (comma), and represents a separator within a single sentence. In addition, the correction candidate generation unit 24c infers the segmentation positions of the subtitle texts "自衛隊の医官や看護官らが編成完結式を行ないました" (jieitai-no-ikan-ya-kangokan-ra-ga-hensei-kanketsushiki-o-okonaim ashita: The medical officers and nurses of the Self-Defense Forces held a formation completion ceremony), "自衛隊の看護官らが編成完結式を行ないました" (jieitai-no-kangokan-ra-ga-hensei-kanketsushiki-o-okonaimashita: The nurses of the Self-Defense Forces held a formation completion ceremony), and "自衛隊の、いかんや看護官らが編成完結式を行ないました" (jieitai-no,ikan-ya-kangokan-ra-ga-hensei-kanketsushiki-o-okonaimashita: The Self-Defense Forces', medical officers and nurses held a formation completion ceremony), each including 3 correction candidates. In this case, the segmentation position is before "編成完結式" (hensei-kanketsushiki: formation completion ceremony), the same as before the correction.

[0179] Next, the correction candidate generation unit 24c generates the first to third candidate subtitle texts based on the three correction candidates and the segmentation positions (step ST8A-5). The first candidate subtitle text is "Self-Defense Force doctors and nurses completed the compilation completion formula (¥n)" (jieitai-no-ikan-ya-kangokan-ra-ga(¥n)hensei-kanketsushiki-o-okonaimashita). The second candidate subtitle text is "Self-Defense Force nurses completed the compilation completion formula (¥n)" (jieitai-no-kangokan-ra-ga(¥n)hensei-kanketsushiki-o-okonaimashita). The third candidate subtitle text is "Self-Defense Force, doctors and nurses completed the compilation completion formula (¥n)" (jieitai-no,ikan-ya-kangokan-ra-ga(¥n)hensei-kanketsushiki-o-okonaimashita). In addition, the correction candidate generation unit 24c sends the correction candidates and the candidate subtitle texts to the user interface unit 25.

[0180] (Time t7)

[0181] The user interface unit 25 presents the three generated correction candidates to the user (step ST8A-6).

[0182] (Time t8)

[0183] When the user interface unit 25 selects the correction candidate "no-ikan-ya" (replaced with the Japanese "ikan: medical officer") based on the user's action, it presents the candidate subtitle text corresponding to the correction candidate to the user. In this example, the first candidate subtitle text "Self-Defense Force doctors and nurses completed the compilation completion formula (¥n)" (jieitai-no-ikan-ya-kangokan-ra-ga(¥n)hensei-kanketsushiki-o-okonaimashita) is displayed.

[0184] The above is an example of correcting the subtitle text at times t6 to t8. In this way, the subtitle text can be corrected according to the user's actions.

[0185] As described above in the third embodiment, the user interface unit 25, acting as the detection unit, detects the user's actions on a portion of the prompted subtitle text and generates action data. The correction expectation receiving unit 26, acting as the decision unit, determines the scope of subtitle text to be corrected based on this action data. The correction candidate generation unit generates correction candidates for the correction scope based on the correction target range and historical data 21a, and infers the segmentation position of the subtitle text containing the correction candidate. Therefore, in addition to the effects of the first embodiment, correction candidates for the prompted subtitle text can be easily generated based on the user's actions. Furthermore, by selecting correction candidates based on the user's actions, the subtitle text can be easily corrected.

[0186] <Fourth Implementation>

[0187] Figure 11 This is a block diagram illustrating the hardware structure of the subtitle generation apparatus according to the fourth embodiment. The fourth embodiment is a specific example of the first to third embodiments, and is a form in which the subtitle generation apparatus 20 is implemented by a computer.

[0188] The subtitle generation apparatus 20 includes, as hardware, a CPU (Central Processing Unit) 201, RAM (Random Access Memory) 202, program memory 203, auxiliary storage device 204, and input / output interface 205. The CPU 201 communicates with the RAM 202, program memory 203, auxiliary storage device 204, and input / output interface 205 via a bus. In other words, the subtitle generation apparatus 20 of this embodiment is implemented using a computer with this hardware structure.

[0189] CPU 201 is an example of a general-purpose processor. CPU 201 uses RAM 202 as its working memory. RAM 202 includes volatile memory such as SDRAM (Synchronous Dynamic Random Access Memory). RAM 202 can also be used as a history unit 21 to store historical data. Program memory 203 stores a subtitle generation program for implementing the various parts corresponding to each embodiment. This subtitle generation program can be, for example, a program for enabling the computer to function as each of the aforementioned units such as the history unit 21, acquisition unit 22, history update unit 23, subtitle processing unit 24, and user interface unit 25. Furthermore, program memory 203 may be, for example, a ROM (Read-Only Memory), part of an auxiliary storage device 204, or a combination thereof. Auxiliary storage device 204 stores data in a non-transitory manner. Auxiliary storage device 204 includes non-volatile memory such as HDD (hard disk drive) or SSD (Solid State Drive). Furthermore, RAM 202, program memory 203, and auxiliary storage device 204 are not limited to being built into the computer, but can also be external to the computer.

[0190] Input / output interface 205 is an interface used for connecting to other devices. Input / output interface 205 is used, for example, for connecting to a keyboard, mouse, and monitor.

[0191] The subtitle generation program stored in program memory 203 includes computer-executable commands. When the subtitle generation program (computer-executable commands) is executed by CPU 201, which is a processing circuit, it causes CPU 201 to perform predetermined processing. For example, when the subtitle generation program is executed by CPU 201, it causes CPU 201 to perform processing related to... Figure 1 , Figure 5 as well as Figure 8 The process involves a series of steps, each part of which is described. For example, when a computer-executable command included in the subtitle generation program is executed by the CPU 201, the CPU 201 executes the subtitle generation method. The subtitle generation method may also include steps corresponding to the functions of the aforementioned history unit 21, acquisition unit 22, history update unit 23, subtitle processing unit 24, generation unit 24a, and user interface unit 25. Here, the subtitle processing unit 24 may also include, in addition to the generation unit 24a, a regeneration unit 24b and a candidate generation correction unit 24c. Furthermore, the subtitle generation method may also include steps corresponding to the functions of the desired reception unit 26. Additionally, the subtitle generation method may appropriately include... Figure 3 , Figure 6 , Figure 9The steps shown.

[0192] The subtitle generation program can also be provided to the subtitle generation apparatus 20, which is a computer, in the form of storage on a computer-readable storage medium. In this case, for example, the subtitle generation apparatus 20 also includes a drive (not shown) for reading data from the storage medium and retrieving the subtitle generation program from the storage medium. As the storage medium, for example, a magnetic disk, optical disk (CD-ROM, CD-R, CD-RW, DVD-RAM, DVD-ROM, DVD-R, etc.), optical disc (MO, etc.), semiconductor memory, etc., can be appropriately used. The storage medium can also be referred to as a non-transitory computer-readable storage medium. Alternatively, the subtitle generation program can be stored on a server on a communication network, and the subtitle generation apparatus 20 can download the subtitle generation program from the server using the input / output interface 205.

[0193] The processing circuitry executing the subtitle generation program is not limited to general-purpose hardware processors such as the CPU201, but can also use dedicated hardware processors such as ASICs (Application Specific Integrated Circuits). The term "processing circuitry" (processing unit) includes at least one general-purpose hardware processor, at least one dedicated hardware processor, or a combination of at least one general-purpose hardware processor and at least one dedicated hardware processor. Figure 11 In the example shown, CPU201, RAM202, and program memory203 are equivalent to the processing circuit.

[0194] According to at least one of the embodiments described above, when generating subtitles in real time based on sound recognition results, the burden of making corrections to make the subtitles easier to read can be reduced.

[0195] Furthermore, while several embodiments of the invention have been described, these embodiments are merely illustrative and not intended to limit the scope of the invention. These embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included within the scope and spirit of the invention, and likewise within the scope of the claims and their equivalents.

[0196] Furthermore, the above-described embodiments can be summarized into the following technical solutions.

[0197] Technical Solution 1

[0198] A subtitle generation device, comprising:

[0199] The acquisition section sequentially obtains the text from the sound recognition results;

[0200] The history department will save the text as historical data.

[0201] The generation unit infers the segmentation and combination positions of the text based on one or more of the saved historical data, and generates subtitle text based on the segmentation and combination positions and the historical data.

[0202] The history update unit updates the historical data based on the segmentation position and the combination position; and

[0203] The prompt section displays the subtitle text.

[0204] Technical Solution 2

[0205] According to the above technical solution 1, the generation unit infers the combination position by combining the text of the historical data until the number of characters of the linked text obtained by combining the text becomes a threshold number of characters or more.

[0206] Technical Solution 3

[0207] According to the above technical solution 1, wherein,

[0208] The historical data includes the text and the date and time the text was obtained.

[0209] The generation unit infers the combination position by combining text of historical data that will not be displayed even after a certain period of time has elapsed since the date and time of acquisition.

[0210] Technical Solution 4

[0211] According to the above technical solution 2 or 3, the generation unit performs one or both of lexical parsing and dependency parsing on the linked text obtained by combining the text, generates feature vectors for each word based on the execution result, uses the feature vectors to calculate the segmentation prediction score for each word, and predicts the segmentation position based on the segmentation prediction score.

[0212] Technical Solution 5

[0213] According to the above technical solution 2 or 3, the generation unit performs one or both of lexical parsing and dependency parsing on the linked text obtained by combining the text, summarizes the results of the execution by phrase unit to generate feature vectors for each phrase, uses the feature vectors to calculate the segmentation prediction score for each phrase, and predicts the segmentation position based on the segmentation prediction score.

[0214] Technical Solution 6

[0215] According to the above technical solution 2 or 3, the generation unit performs one or both of lexical parsing and dependency parsing on the linked text obtained by combining the text, generates the embedding vector of each word based on the result of the execution, uses the embedding vector to calculate the segmentation prediction score of each word, and infers the segmentation position based on the segmentation prediction score.

[0216] Technical Solution 7

[0217] According to the above technical solution 4 or 6, the generation unit calculates the segmentation prediction score for each word using an LSTM model that considers the context before the word, the context after the word, or the context before and after the word.

[0218] Technical Solution 8

[0219] According to any one of the above technical solutions 4 to 7, when the number of historical data, including texts not used as subtitle text prompts, stored in the historical data of the historical department is less than a threshold, the generation unit infers the segmentation position by calculating the segmentation inference score.

[0220] Technical Solution 9

[0221] According to the above technical solution 2, the generation unit infers the segmentation position based on a fixed number of characters different from the threshold number of characters for the linked text obtained by combining the text.

[0222] Technical Solution 10

[0223] According to the above technical solution 9, when the number of historical data, including texts not used as subtitle text prompts, stored in the historical data of ...

[0224] Technical Solution 11

[0225] According to any one of the above technical solutions 1 to 10, wherein,

[0226] The subtitle generation device also features:

[0227] The detection department detects the user's actions in response to a portion of the subtitle text and generates action data.

[0228] The determination unit, based on the motion data, determines the scope of the subtitle text to be corrected, and the desired type of any party in the segmentation and combination of the subtitle text; and

[0229] The regeneration unit re-infers the segmentation or combination position of the subtitle text based on the range of the correction object, the expected type, and the historical data, and regenerates the subtitle text based on the re-infermentation result.

[0230] Technical Solution 12

[0231] According to any one of the above technical solutions 1 to 10, wherein,

[0232] The subtitle generation device also features:

[0233] The user interface department generates action data from the user's actions in response to a portion of the subtitle text.

[0234] The decision-making unit determines the scope of correction for the subtitle text based on the motion data; and

[0235] The correction candidate generation unit generates correction candidates for the correction target range based on the correction target range and the historical data, and infers the segmentation position of the subtitle text containing the correction candidate.

[0236] Technical Solution 13

[0237] A subtitle generation method, comprising:

[0238] Text from which sound recognition results are obtained sequentially;

[0239] Save the text as historical data;

[0240] The segmentation and combination positions of the text are inferred based on one or more of the saved historical data, and subtitle text is generated based on the segmentation and combination positions and one or more of the historical data.

[0241] The historical data is updated based on the segmentation position and the combination position; and

[0242] The text in the subtitle is displayed.

[0243] Technical Solution 14

[0244] A storage medium storing a subtitle generation program, the subtitle generation program being used to enable a computer to function as a unit:

[0245] A unit of text that obtains sound recognition results sequentially;

[0246] The unit that saves the text as historical data;

[0247] The segmentation and combination positions of the text are inferred based on one or more of the saved historical data, and subtitle text units are generated based on the segmentation and combination positions and one or more of the historical data.

[0248] The unit that updates the historical data based on the segmentation position and the combination position;

[0249] The unit that indicates the subtitle text.

Claims

1. A subtitle generation apparatus comprising: an acquisition unit that sequentially acquires a text of a sound recognition result; a history unit that stores the text as history data; a generation unit that estimates a division position and a combination position of the text from one or more pieces of the stored history data, and generates a subtitle text from one or more pieces of the history data based on the division position and the combination position; a history update unit that updates the history data based on the division position and the combination position; and a presentation unit that presents the subtitle text, in the generation unit, the combination position is estimated in a manner in which a character number of a concatenated text obtained by combining the text becomes a threshold character number or more, for the concatenated text obtained by combining the text, morpheme analysis and dependency analysis are performed, a feature vector of each word that uses a word index from the beginning of an article, a character length, a part of speech, and a dependency analysis result of each word as features is generated based on a result of the performance, a division estimation score of each word is calculated using the feature vector, and the division position is estimated based on the division estimation score, in a case where the division estimation score of each word is calculated, the division estimation score is calculated using an LSTM model that takes into account a context before the each word, a context after the each word, or a context before and after the each word. 2.The subtitle generation apparatus according to claim 1, wherein the history data includes the text and a date and time of acquisition of the text, the generation unit estimates the combination position in a manner in which the text of the history data that is not displayed even if a certain time elapses from the date and time of acquisition is combined. 3.The subtitle generation apparatus according to claim 1 or 2, wherein the subtitle generation apparatus further comprises: a detection unit that detects an action of a user with respect to a part of the presented subtitle text, and generates action data; a determination unit that determines a correction target range of the subtitle text and a desired type of any of division and combination of the subtitle text based on the action data; and a regeneration unit that re-estimates a division position or a combination position of the subtitle text based on the correction target range, the desired type, and the history data, and regenerates a subtitle text based on a re-estimation result. 4.The subtitle generation apparatus according to claim 1 or 2, wherein the subtitle generation apparatus further comprises: a user interface unit that generates an action of a user with respect to a part of the presented subtitle text as action data; a decision unit that decides a correction target range of the subtitle text based on the action data; and a correction candidate generation unit that generates a correction candidate of the correction target range based on the correction target range and the history data, and estimates a division position of a subtitle text including the correction candidate. 5.A subtitle generation method comprising: sequentially acquiring a text of a sound recognition result; storing the text as history data; ​ ​ speculating a split position and a combination position of the text from one or more of the saved history data, and generating a caption text from one or more of the history data based on the split position and the combination position; updating the history data based on the split position and the combination position; and presenting the caption text, in the processing of generating the caption text, speculating the combination position in a manner that combines the text of the history data until the number of characters of a concatenated text obtained by combining the text becomes a threshold number of characters or more, performing morphological analysis and dependency analysis on a concatenated text obtained by combining the text, generating a feature vector of each word using a word index from the beginning of the article, a character length, a part of speech, and a dependency analysis result of each word as features from a result of the performing, calculating a split speculation score of each word using the feature vector, and speculating the split position from the split speculation score, in a case where the split speculation score of each word is calculated, calculating the split speculation score of the each word using an LSTM model that takes into account a context before the each word, a context after the each word, or a context before and after the each word.

6. A storage medium storing a caption generation program for causing a computer to function as: a unit that sequentially acquires a text of a sound recognition result; a unit that saves the text as history data; a unit that speculates a split position and a combination position of the text from one or more of the saved history data, and generates a caption text from one or more of the history data based on the split position and the combination position; a unit that updates the history data based on the split position and the combination position; and a unit that presents the caption text, in the unit that generates the caption text, speculating the combination position in a manner that combines the text of the history data until the number of characters of a concatenated text obtained by combining the text becomes a threshold number of characters or more, performing morphological analysis and dependency analysis on a concatenated text obtained by combining the text, generating a feature vector of each word using a word index from the beginning of the article, a character length, a part of speech, and a dependency analysis result of each word as features from a result of the performing, calculating a split speculation score of each word using the feature vector, and speculating the split position from the split speculation score, in a case where the split speculation score of each word is calculated, calculating the split speculation score of the each word using an LSTM model that takes into account a context before the each word, a context after the each word, or a context before and after the each word.

Citation Information

Patent Citations

  • Molded body

    JP2021178498A

  • Subtitle generation method and device, computer storage medium and electronic equipment

    CN112002328A