Synchronization control device
The synchronization control device synchronizes translated subtitles with video content by delaying the video based on the total processing time of the longest audio segment, ensuring the integrity and synchronized playback of translated subtitles.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NTT DOCOMO INC
- Filing Date
- 2023-01-06
- Publication Date
- 2026-07-23
AI Technical Summary
Existing technologies fail to synchronize translated subtitles with video content by translating audio segments.
A synchronization control device that synchronizes translated subtitles with video content by detecting and translating audio segments, ensuring the integrity of the video content.
Ensures the integrity of the video content while adding translated subtitles in sync, using a synchronization control device that delays the video content by the total processing time of the longest audio segment.
Smart Images

Figure 0007894325000001 
Figure 0007894325000002 
Figure 0007894325000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a synchronization control device that performs control for synchronously reproducing video content and translated subtitle data.
Background Art
[0002] In recent years, services for delivering video content in real time have been spreading. In such services, after speech recognition of the audio included in the video content, translated subtitles translated into other languages are provided to the video content on demand, and it has been long awaited to synchronously reproduce the translated subtitles and the video content.
[0003] In relation to this, technologies for using, as subtitles, text data obtained by speech recognition of the audio included in video content or subtitle data prepared in advance and included in video content, and synchronously displaying the video and the subtitles on a terminal have been proposed in Patent Documents 1 and 2 below.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, Patent Documents 1 and 2 did not assume a process of further translating the text data obtained by speech recognition from video content into other languages to obtain translated subtitles.
[0006] Based on the above, this disclosure aims to ensure the integrity of the video content, add translated subtitles to the video content by speech recognition of the audio contained in the video content and translating it into another language, and then play the translated subtitles and the video content in sync. [Means for solving the problem]
[0007] The synchronization control device according to this disclosure includes: a detection unit that detects a series of multiple audio segments in video content and detects the longest audio segment among the multiple audio segments; an audio recognition unit that acquires text data for each audio segment as a result of audio recognition processing targeting the audio of the detected multiple audio segments; a translation unit that acquires translated subtitle data for each audio segment as a result of translation processing targeting the acquired text data into another language; and a control unit that calculates the total processing time for the longest audio segment, including the processing time by the detection unit, the processing time by the audio recognition unit, and the processing time by the translation unit, delays the video content by the obtained total processing time, and adds the translated subtitle data for each audio segment to the delayed video content and plays it back.
[0008] In the synchronization control device described above, the detection unit detects a series of multiple audio segments in the video content, and detects the longest audio segment among these multiple audio segments. The speech recognition unit acquires text data for each audio segment as a result of speech recognition processing on the audio of the detected multiple audio segments. The translation unit acquires translated subtitle data for each audio segment as a result of translating the acquired text data into another language. Here, the control unit calculates the total processing time, which includes (1) the processing time by the detection unit targeting the longest audio segment, (2) the processing time by the speech recognition unit, and (3) the processing time by the translation unit. The control unit then delays the video content by the obtained total processing time, adds translated subtitle data for each audio segment to the delayed video content, and plays it back. In this way, the video content is delayed by the above total processing time, and translated subtitle data for each audio segment is added to the delayed video content and played back, so that the translated subtitles and video content can be played back in sync. Furthermore, since no changes are made to the video content, the integrity of the video content is ensured. [Effects of the Invention]
[0009] According to this disclosure, while ensuring the integrity of the video content, it is possible to add translated subtitles to the video content, obtained by speech recognition of the audio contained in the video content and translating it into another language, and to play the translated subtitles and the video content in sync. [Brief explanation of the drawing]
[0010] [Figure 1] This is a functional block diagram showing the configuration of the synchronous control device and peripheral devices in the first embodiment. [Figure 2] This is a diagram illustrating the synchronous playback method in the first embodiment. [Figure 3] This is a flowchart showing the processing performed by the synchronous control device and peripheral devices in the first embodiment. [Figure 4] This diagram illustrates an example of effective use of non-voice sections. [Figure 5]This diagram illustrates another example of effective use of non-voice sections. [Figure 6] This is a functional block diagram showing the configuration of the synchronous control device and peripheral devices in the second embodiment. [Figure 7] This is a diagram illustrating the synchronous playback method in the second embodiment. [Figure 8] This is a flowchart showing the processing performed by the synchronous control device and peripheral devices in the second embodiment. [Figure 9] This figure shows another example of the configuration of the synchronous control device in the first embodiment. [Figure 10] This figure shows another example of the configuration of the synchronous control device in the second embodiment. [Figure 11] This figure shows an example of the hardware configuration of a synchronous control device. [Modes for carrying out the invention]
[0011] Hereinafter, the first and second embodiments of the synchronization control device according to this disclosure will be described with reference to the drawings. In the first embodiment, the video content is delayed by the total processing time for the longest audio section among the multiple audio sections included in the video content, and the translated subtitle data in "one language" obtained from the audio of each audio section is played back in sync with the video content. In the second embodiment, the video content is delayed by the total processing time including the longest translation processing time among the processing time including translation into "multiple languages" for the longest audio section, and the translated subtitle data in the selected language for each audio section is played back in sync with the video content.
[0012] (First Embodiment) As shown in Figure 1, the synchronization control device 10 in the first embodiment includes, as functional blocks, a detection unit 11, a speech recognition unit 12, a translation unit 13, a control unit 14, and a data storage unit 15 in order to perform the functions according to this disclosure. Such a synchronization control device 10 is composed of various information processing devices equipped with communication functions (for example, smartphones, mobile phones, tablet terminals, notebook personal computers, desktop computers, servers, etc.) as hardware. The functional blocks in Figure 1 will be described below.
[0013] The data storage unit 15 is a functional unit that stores various types of data, such as video content and translated subtitle data, which will be described later.
[0014] The detection unit 11 is a functional unit that receives video content including video data and audio data from the outside, detects a series of multiple audio segments in the video content, and detects the longest audio segment with the longest time among the multiple audio segments. Here, the "audio segment" means the time period during which audio is output when the video content is played back in chronological order. On the other hand, the time period during which no audio is output is referred to as a "silent segment". Further, the detection unit 11 assigns a sequence number to each of the detected multiple audio segments, outputs the audio data of each audio segment with the assigned sequence number to the audio recognition unit 12 described later, and saves the video content in the data storage unit 15. The above "sequence number" is effective for identifying the audio data for each audio segment. Note that it is not essential for the detection unit 11 to assign a sequence number. Instead of the detection unit 11, the audio recognition unit 12 described later may assign a sequence number when receiving the audio data of each audio segment. Also, when the data volume of the target video content is huge, the processing time for detecting multiple audio segments in the video content may become long. In such a case, for example, a reference value regarding the data volume of the target video content is preset in advance. When the data volume of the video content exceeds the reference value, the detection unit 11 divides the video content with a large data volume into multiple parts along the chronological order so that the data volume of each divided video content is below the reference value, and the divided individual video contents may be used as detection processing targets. In this way, even if the data volume of the original target video content is huge, by using the divided individual video contents as detection processing targets, it is possible to prevent inconveniences such as the detection processing time of the audio segment becoming long in advance.
[0015] The speech recognition unit 12 is a functional unit that acquires text data for each speech segment as a result of speech recognition processing targeting the speech of multiple speech segments detected by the detection unit 11. Specifically, the speech recognition unit 12 sends the speech data for each speech segment with a sequence number and the speech recognition processing request received from the detection unit 11 to the external speech recognition API (Application Programming Interface) server 20. After that, it receives the text data as a result of speech recognition processing by the speech recognition API server 20, and further verifies that the received text data is valid as a sentence in light of the grammar of the speech language, and corrects any deficiencies as appropriate (hereinafter referred to as "verification processing"). The speech recognition unit 12 also assigns the sequence number of the corresponding speech segment to the speech recognition result (text data) acquired in the manner described above, and outputs the speech recognition result (text data) after the assignment of the sequence number to the translation unit 13 described later. By assigning a sequence number, the speech recognition result (text data) for each speech segment can be identified. The speech recognition processing applied to the audio data of each audio segment as described above may be performed sequentially for each audio segment, or, if processing capacity allows, it may be performed concurrently for the audio data of multiple audio segments. In this embodiment, an example of sequential execution for each audio segment will be described later, following the flowchart in Figure 3.
[0016] The translation unit 13 is a functional unit that obtains translation subtitle data as a result of translation into another language for the text data of each voice section acquired by the voice recognition unit 12. Specifically, the translation unit 13 transmits the text data and translation processing request of each voice section received from the voice recognition unit 12 to an external translation API (Application Programming Interface) server 30, and then receives the translation subtitle data as a result of the translation by the translation API server 30. Furthermore, the received translation subtitle data is given the sequence number of the corresponding voice section, and the translation subtitle data after the sequence number is given is output to a control unit 14 described later and stored in a data storage unit 15. By assigning the sequence number, the translation subtitle data for each voice section can be identified. The translation processing for the text data of each voice section as described above may be sequentially executed for each voice section, or may be executed in parallel for the text data of a plurality of voice sections if there is processing capacity. In the present embodiment, an example of sequentially executing for each voice section will be described later along the flowchart of FIG. 3. Also, the languages of the subtitle original text and the translation subtitle to be targeted can assume any combination. For example, various forms such as a form of obtaining Chinese translation subtitle data from a Japanese subtitle original text, a form of obtaining Japanese translation subtitle data from an English subtitle original text, etc. can be adopted.
[0017] The control unit 14 is a functional unit that calculates the total processing time of the detection unit 11, the elapsed processing time of the longest audio section by the speech recognition unit 12, and the processing time of the translation unit 13, delays the video content by the obtained total processing time, and then plays back the delayed video content with the translation subtitle data for each audio section. Specifically, the control unit 14 obtains the processing time of the detection unit 11, the elapsed processing time of the longest audio section by the speech recognition unit 12, and the processing time of the translation unit 13 by querying the detection unit 11, the speech recognition unit 12, and the translation unit 13, and calculates the total processing time. Then, the control unit 14 obtains the video content and the translation subtitle data after the sequence number has been assigned to each audio section from the data storage unit 15, delays the video content by the above total processing time, and plays back the delayed video content with the translation subtitle data for each audio section.
[0018] Next, the delay control of video content by the control unit 14 will be explained using Figure 2. As shown in Figure 2, the synchronization control device 10 performs the following processes on the audio data of each detected audio section in the video content: "A. Audio section detection processing" by the detection unit 11, "B. Audio recognition processing" by the speech recognition unit 12, and "C. Translation processing" by the translation unit 13. Of the above, "A. Audio section detection processing" and "C. Translation processing," which targets text data, have short processing times, while "B. Audio recognition processing" has a long processing time because the amount of audio data to be processed is enormous. Specifically, "B. Processing time required for speech recognition processing" includes (1) the time for sending and receiving data with the external speech recognition API server 20 (including data transmission time in the communication environment), (2) the time for speech recognition processing for each audio section by the external speech recognition API server 20, and (3) the time for confirmation processing by the speech recognition unit 12 on the speech recognition results. In this case, steps (1) to (3) above may be executed sequentially, or some may be executed in parallel.
[0019] Here, if (1) to (3) above are executed in series, the control unit 14 calculates the total processing time by the speech recognition unit 12 as the sum of the processing time for the transmission / reception process, the processing time for the speech recognition process for the longest audio section, and the processing time for the confirmation process. On the other hand, if (1) to (3) above are executed partially in parallel, the control unit 14 calculates the total processing time from the start to the end of the transmission / reception process, the speech recognition process for the longest audio section, and the confirmation process as the processing time by the speech recognition unit 12. Thus, the control unit 14 has the advantage of being able to appropriately calculate "B. Processing time required for speech recognition processing" depending on the execution mode, whether (1) to (3) above are executed in series or partially in parallel. Furthermore, the control unit 14 acquires the video content and translated subtitle data for each audio section from the data storage unit 15, delays the video content by the total processing time, and adds the translated subtitle data for each audio section to the delayed video content before playback.
[0020] Next, a series of processes performed by the synchronization control device 10 and related devices will be explained in accordance with the flowchart in Figure 3. First, in the synchronization control device 10, the detection unit 11 receives the target video content, including video data and audio data, from an external source, detects a series of audio segments in the video content using existing voice activity detection (VAD) technology, and detects the longest audio segment among the multiple audio segments (step S1). At this time, the detection unit 11 assigns a sequence number to each of the detected audio segments, outputs the audio data of each audio segment with the assigned sequence number to the voice recognition unit 12 (described later), and saves the video content to the data storage unit 15.
[0021] The speech recognition unit 12 sends the audio data of one unprocessed audio segment from among multiple audio segments and a speech recognition processing request to the speech recognition API server 20 (step S2). The speech recognition API server 20 receives the audio data of the audio segment and the speech recognition processing request (step S3), and, in response to the request, performs speech recognition processing on the audio data of the audio segment (step S4), and then sends the text data of the audio segment obtained from the speech recognition processing to the speech recognition unit 12 (step S5). The speech recognition unit 12 receives the text data of the audio segment (speech recognition processing result) from the speech recognition API server 20 (step S6), and performs a verification process to confirm that the received text data is valid as a sentence in light of the grammar of the spoken language, and corrects any deficiencies as appropriate (step S7). The text data after the above verification process (correct text data in light of the language grammar, etc.) is assigned the sequence number of the corresponding audio segment and then output to the translation unit 13. Furthermore, the speech recognition unit 12 determines whether speech recognition processing has been completed for all speech segments (step S8). If there are any unprocessed speech segments (NO in step S8), it returns to step S2 and executes the processing in steps S2 to S7 on the speech data of the one unprocessed speech segment. If it is then determined that speech recognition processing has been completed for all speech segments (YES in step S8), it proceeds to step S9.
[0022] The translation unit 13 sends text data of one unprocessed audio segment from among multiple audio segments (text data after confirmation processing output from the speech recognition unit 12) and a translation processing request to the translation API server 30 (step S9). The translation API server 30 receives the text data of the audio segment and the translation processing request (step S10), and in response to the request, performs translation processing of the text data of the audio segment into a predetermined language (step S11), and then sends the translated subtitle data of the audio segment obtained from the translation processing to the translation unit 13 (step S12). The translation unit 13 receives the translated subtitle data of the audio segment (translation processing result) from the translation API server 30, assigns a sequence number to the corresponding audio segment, and saves the translated subtitle data after the sequence number has been assigned to the data storage unit 15 (step S13). Furthermore, the translation unit 13 determines whether all audio segments have been translated (step S14). If there are any unprocessed audio segments (NO in step S14), it returns to step S9 and executes the processes in steps S9 to S13 on the audio data of the one unprocessed audio segment. If it is then determined that all audio segments have been translated (YES in step S14), it proceeds to step S15.
[0023] The control unit 14 queries the detection unit 11, the speech recognition unit 12, and the translation unit 13 to obtain the processing time for the longest audio section (i.e., the processing time by the detection unit 11, the elapsed processing time by the speech recognition unit 12, and the processing time by the translation unit 13), calculates the total processing time (step S15), obtains the video content and subtitle translation data for each audio section from the data storage unit 15, delays the video content by the total processing time as explained with reference to Figure 2, adds the translated subtitle data to the delayed video content and plays it back (step S16).
[0024] As described above, since no changes are made to the video content, the integrity of the video content is ensured. Translated subtitles, created by speech recognition of the audio contained in the video content and translating it into another language, are added to the video content with a delay equal to the total processing time of the longest audio section. This allows for synchronized playback of the translated subtitles and the video content.
[0025] (Regarding the effective use of silent sections) Generally, the subtitle display area in video content is fixed, and there is an upper limit to the number of translated subtitle characters that can be displayed in the subtitle display area. Therefore, if the number of characters in the translated subtitle obtained during the translation process exceeds the upper limit, the translated subtitle is generally displayed by scrolling in the subtitle display area. When the control unit 14 detects that the above situation is occurring (i.e., the number of characters in the translated subtitle exceeds the upper limit), that is, when the translated subtitle is displayed by scrolling in the video content, the control unit 14 may perform control to add the translated subtitle data from the preceding audio section to the "silent section" of the video content and play it back. Specifically, examples include (A) repeatedly displaying the translated subtitle data from the preceding audio section by scrolling, and (B) without scrolling, dividing the translated subtitle data from the preceding audio section, displaying the divided translated subtitle data (part of the translated subtitle data) in the preceding audio section, and displaying the remainder of the divided translated subtitle data in the silent section.
[0026] Figure 4 shows an example of (A) above. In a video content where audio section A, a silent section, and audio section B follow in chronological order, the control unit 14 repeatedly scrolls and displays lines 1 to 4 of the translated subtitle data in audio section A. It also continues to repeatedly scroll and display lines 1 to 4 of the translated subtitle data in the silent section immediately following audio section A. This prevents inconveniences such as not being able to display all of the translated subtitle data in audio section A alone, and difficulty for viewers to read all of the translated subtitle data when only audio section A is displayed.
[0027] Figure 5 shows an example of (B) above. In a video content where audio section A, a silent section, and audio section B follow in chronological order, the control unit 14 splits the first to fourth lines of the translated subtitle data in audio section A, displays the first to second lines of the split translated subtitle data in audio section A, and displays the remaining lines (third to fourth lines) of the split translated subtitle data in the silent section immediately following audio section A. This prevents the inconvenience of not being able to display all of the translated subtitle data in audio section A alone, as in example (A) above, and also prevents the inconvenience of viewers having difficulty reading all of the translated subtitle data when only audio section A is displayed.
[0028] (Second Embodiment) Next, as a second embodiment, we will describe a configuration in which the video content is delayed by the total processing time, including the longest translation processing time among the processing times for translating the longest audio section into "multiple languages," and the translated subtitle data for the selected language for each audio section is played back in sync with the video content. Since the second embodiment has many parts in common with the first embodiment described above, we will focus on explaining the points specific to the second embodiment in order to avoid redundant explanations.
[0029] Figure 6 shows the configuration of the synchronization control device 10 in the second embodiment. The synchronization control device 10 in the second embodiment is characterized in that the translation unit 13 makes translation requests to each of the translation API servers 30A to 30C for each translation language and obtains the translation processing result (translated subtitle data for the target language), which is different from the synchronization control device in the first embodiment. The other components are the same as in the first embodiment.
[0030] As an example, let's assume that translation API server 30A has the function of translating Japanese text into English and outputting translated subtitle data in English, translation API server 30B has the function of translating Japanese text into Chinese and outputting translated subtitle data in Chinese, and translation API server 30C has the function of translating Japanese text into Korean and outputting translated subtitle data in Korean. In this case, the synchronization control device 10 is configured to acquire Japanese text data from video content including Japanese audio data, acquire translated subtitle data in multiple desired languages from English, Chinese, and Korean from the Japanese text data, and synchronize the translated subtitles with the video content for playback. The above configuration of translation API server 30 is just an example, and in addition to the above, • A function to translate English text into Japanese and output Japanese translated subtitle data. • A function to translate English text into Chinese and output translated subtitle data in Chinese. • A function to translate Chinese text into Japanese and output Japanese translated subtitle data. Features include translating Chinese text into English and outputting translated subtitle data in English. A configuration may be adopted in which the translation unit 13 requests translation from each individual translation API server that has each of the respective functions.
[0031] Next, using Figure 7, the delay control of video content by the control unit 14 of the second embodiment will be explained. As shown in Figure 7, in the synchronization control device 10, the following are performed on the audio data of each detected audio section in the video content: "A. Audio section detection processing" by the detection unit 11, "B. Audio recognition processing" by the speech recognition unit 12, and "C. Translation processing" by the translation unit 13. Of these, "C. Translation processing" differs from the first embodiment in that in the second embodiment, it is translation processing into "multiple languages". Since the translation processing into multiple languages is performed in parallel by multiple translation API servers 30A to 30C, the translation processing time for each language is different from each other, but the "longest translation processing time" among the translation processing times for each language is used.
[0032] In other words, the second embodiment is characterized by calculating the delay time for delaying the video content using the "longest translation processing time" among the translation processing times for multiple languages targeting the longest audio segment. Specifically, the control unit 14 delays the video content by the total processing time of "A. Processing time required for audio segment detection processing," "B. Processing time required for speech recognition processing," and "C. Longest translation processing time among the translation processing times for multiple languages" targeting the longest audio segment, and then adds translated subtitle data for each audio segment to the delayed video content and plays it back.
[0033] Next, the processing of the synchronization control device 10 in the second embodiment will be explained using Figure 8. In the second embodiment, the "translation processing" in steps S9 to S14 and the "control processing" in steps S15A and S16A differ from those in the first embodiment, so these processes will be explained below. Here, as an example, the process of requesting translation processing from translation API servers 30A that translate from Japanese to English, translation API server 30B that translates from Japanese to Chinese, and translation API server 30C that translates from Japanese to Korean will be explained, assuming that the source text for the subtitles is in Japanese.
[0034] In step S9, the translation unit 13 sends the text data of one unprocessed audio segment from among multiple audio segments (the text data after confirmation processing output from the speech recognition unit 12) and a translation processing request to the aforementioned translation API servers 30A to 30C (step S9). Each of the translation API servers 30A to 30C receives the text data of the audio segment and the translation processing request (step S10), and, in response to the request, performs translation processing of the text data of the audio segment into a predetermined language (the language of the translation function that each of the translation API servers 30A to 30C has) (step S11), and then sends the translated subtitle data in English, Chinese, and Korean obtained from the translation processing to the translation unit 13 (step S12). The translation unit 13 receives translated subtitle data (translation processing results) in English, Chinese, and Korean for the above audio sections from each of the translation API servers 30A to 30C, assigns a sequence number to the above audio section, and saves the translated subtitle data after the sequence number has been assigned to the data storage unit 15 (step S13). Furthermore, the translation unit 13 determines whether all audio sections have been translated or not (step S14), and if there are any unprocessed audio sections (NO in step S14), it returns to step S9 and executes the processing in steps S9 to S13 on the audio data of the one unprocessed audio section. After that, if it is determined that all audio sections have been translated (YES in step S14), it proceeds to step S15A.
[0035] The control unit 14 queries the detection unit 11, the speech recognition unit 12, and the translation unit 13 to obtain the processing time for the longest audio section (i.e., the processing time by the detection unit 11, the processing time elapsed by the speech recognition unit 12, and the translation processing time by the translation unit 13, including the translation processing time at each of the translation API servers 30A to 30C), identifies the longest translation processing time among the translation processing times at each of the translation API servers 30A to 30C, calculates the total processing time including the longest translation processing time for the longest audio section (step S15A), obtains the video content and English, Chinese, and Korean subtitle translation data for each audio section from the data storage unit 15, delays the video content by the above total processing time as explained with reference to Figure 7, and plays the delayed video content with the English, Chinese, and Korean translation subtitle data added (step S16A).
[0036] The embodiments described above ensure the integrity of the video content while allowing the audio contained within the video content to be translated into multiple languages (for example, English, Chinese, and Korean) using speech recognition. The translated subtitles are then added to the video content, delayed by a total processing time that includes the longest translation processing time for the longest audio segment. The translated subtitles and the video content can then be played back in sync.
[0037] The above example shows how to obtain translated subtitles in English, Chinese, and Korean from the original Japanese subtitle text and synchronize the translated subtitles with the video content for playback. However, this is just one example, and any combination of languages can be used for the original subtitle text and translated subtitles.
[0038] (Regarding the effective use of silent sections) The method for effectively utilizing the silent sections is substantially the same as that described in the first embodiment using Figures 4 and 5, and produces similar functions and effects. Therefore, redundant explanations are omitted here.
[0039] (modified version) The following describes modified configurations of the synchronization control device 10 shown in Figures 1 and 6. Figure 1 shows an example configuration in which speech recognition processing is requested from an external speech recognition API server 20 to obtain speech recognition results, and translation processing is requested from an external translation API server 30 to obtain translated subtitle data. However, as shown in Figure 9, the synchronization control device 10 may adopt a configuration in which the speech recognition unit 12 contains a speech recognition processing unit 12A that performs speech recognition processing, and the translation unit 13 contains translation processing units 13A to 13C that perform translation processing. Similarly, the configuration shown in Figure 10 is an example of a modified configuration of the synchronization control device 10 shown in Figure 6. That is, as shown in Figure 10, the synchronization control device 10 may adopt a configuration in which the speech recognition unit 12 contains a speech recognition processing unit 12A that performs speech recognition processing, and the translation unit 13 contains translation processing units 13A to 13C that perform translation processing.
[0040] Furthermore, as a variation other than those shown in Figures 9 and 10, the speech recognition unit 12 may include a speech recognition processing unit 12A, but the translation unit 13 may not include a translation processing unit 13A. Alternatively, the speech recognition unit 12 may not include a speech recognition processing unit 12A, but the translation unit 13 may include a translation processing unit 13A. In either of the above configurations, the synchronization control device 10 can perform the same processing as in the above embodiments and achieve the same effects.
[0041] Furthermore, in the first and second embodiments described above, the speech recognition unit 12 is shown to verify that the text data received as a result of speech recognition is valid as a sentence in light of the grammar of the spoken language, and to correct any defects as appropriate (verification process). However, a configuration may also be adopted in which the speech recognition unit 12 requests the verification process from an external server that performs such verification processing upon request, and obtains the text data after the verification process (text data without defects) from the external server.
[0042] The gist of this disclosure is found in the following [1] to [7]. [1] A detection unit that detects a series of multiple audio segments in video content and detects the longest audio segment among the multiple audio segments, A speech recognition unit that obtains text data for each speech segment as a result of speech recognition processing targeting the audio of multiple detected speech segments, A translation unit that obtains translated subtitle data for each audio segment as a result of processing the acquired text data into another language, A control unit calculates the total processing time for the longest audio segment, including the processing time by the detection unit, the processing time by the speech recognition unit, and the processing time by the translation unit; delays the video content by the obtained total processing time; and adds translated subtitle data for each audio segment to the delayed video content before playback. A synchronous control device equipped with the following features. [2] Each of the plurality of audio segments is assigned a sequence number, The speech recognition unit acquires text data for each speech segment accompanied by the sequence number, The translation unit acquires translated subtitle data for each audio segment accompanied by the sequence number, The control unit plays back the audio section of the delayed video content by adding translated subtitle data for each audio section, which is associated with the sequence number. [1] [3] The translation unit obtains translated subtitle data for each audio segment and each language as a result of processing the acquired text data into multiple other languages, The control unit calculates the total processing time for the longest audio section, which is the processing time by the detection unit, the processing time by the speech recognition unit, and the processing time for each language by the translation unit, and delays the video content by the obtained total processing time, and plays back the delayed video content with the translated subtitle data for each audio section, as described in [2]. [4] The speech recognition is performed in series, Data transmission and reception processing with an external speech recognition server, Speech recognition processing by an external speech recognition server, and, Verification process by the speech recognition unit targeting the speech recognition results If it includes, The control unit calculates the sum of the processing time of the transmission / reception process, the processing time of the voice recognition process, and the processing time of the confirmation process as the processing time elapsed by the voice recognition unit, according to any one of [1] to [3]. [5] The speech recognition is performed in parallel, at least partially. Data transmission and reception processing with an external speech recognition server, Speech recognition processing by an external speech recognition server, and, Verification process by the speech recognition unit targeting the speech recognition results If it includes, The control unit calculates the total processing time for the transmission / reception process, the voice recognition process and the confirmation process as the processing time by the voice recognition unit, according to any one of [1] to [3]. [6] If the translated subtitles are displayed by scrolling in the video content, The control unit is a synchronization control device according to any one of [1] to [5], which adds the translated subtitle data of the immediately preceding audio section to the video content and plays it back during silent periods in the video content that are not audio sections. [7] The synchronization control device according to any one of [1] to [6], wherein the detection unit processes individual video content divided into multiple parts along a time series.
[0043] (Explanation of terms, explanation of hardware configuration (Figure 11), etc.) The block diagrams used in the description of the above embodiments show functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may also be realized by combining the above one device or the above multiple devices with software.
[0044] Functions include, but are not limited to, judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration part) that enables transmission is called a transmitting unit or transmitter. As mentioned above, the method of implementation is not particularly limited.
[0045] For example, the synchronous control device 10 in this embodiment may function as a computer that performs the processing of the disclosure. Figure 11 is a diagram showing an example of the hardware configuration of the synchronous control device 10. The synchronous control device 10 described above may be physically configured as a computer device including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, etc.
[0046] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the synchronous control device 10 may include one or more of the devices shown in the figure, or it may be configured to omit some of the devices.
[0047] Each function of the synchronization control device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, which allows the processor 1001 to perform calculations, control communication by the communication device 1004, and control at least one of data reading and writing in the memory 1002 and storage 1003.
[0048] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, and so on.
[0049] Furthermore, the processor 1001 reads programs (program code), software modules, data, etc., from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. Although the above processes have been described as being executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The program may also be transmitted from a network via a telecommunications line.
[0050] Memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 1002 may also be called a register, cache, main memory, etc. Memory 1002 can store executable programs (program code), software modules, etc., for carrying out a wireless communication method according to one embodiment of the present disclosure.
[0051] Storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. Storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of memory 1002 and storage 1003.
[0052] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc. The communication device 1004 may be configured to include, for example, a high-frequency switch, duplexer, filter, frequency synthesizer, etc., in order to implement at least one of frequency division duplex (FDD) and time division duplex (TDD).
[0053] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).
[0054] Furthermore, each device, such as the processor 1001 and memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or different buses may be configured for each device.
[0055] Furthermore, the synchronous control device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.
[0056] The notification of information is not limited to the embodiments described herein and may be carried out by other means. For example, the notification of information may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.
[0057] Each aspect / embodiment described in this disclosure includes LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), 6th generation mobile communication system (6G), xth generation mobile communication system (xG) (xG (where x is, for example, an integer or decimal)), FRA (Future Radio Access), NR (new Radio), New radio access (NX), Future generation radio access (FX), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), and IEEE This may apply to at least one system utilizing 802.20, UWB (Ultra-WideBand), Bluetooth®, or other appropriate systems, and to next-generation systems extended, modified, created, or defined based thereon. It may also apply to a combination of multiple systems (for example, a combination of at least one of LTE and LTE-A with 5G).
[0058] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described herein may be reordered, provided they are consistent with each other. For example, the methods described herein present various step elements in an exemplary order and are not limited to that specific order.
[0059] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be transmitted to other devices.
[0060] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).
[0061] Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).
[0062] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.
[0063] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.
[0064] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0065] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0066] In addition, terms used in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of the communication channel and the symbol may be a signal (signaling). Also, the signal may be a message. Furthermore, the component carrier (CC) may be called a carrier frequency, cell, frequency carrier, etc.
[0067] The terms “system” and “network” as used in this disclosure are interchangeable.
[0068] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values from a given value, or other corresponding information. For example, wireless resources may be indicated by an index.
[0069] The names used for the parameters described above are not restrictive in any way. Furthermore, the formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Various communication channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable name, and therefore, the various names assigned to these various communication channels and information elements are not restrictive in any way.
[0070] As used in this disclosure, the terms “determining” and “determining” may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiry (e.g., searching in a table, database, or other data structure), and ascertaining. “Determining” may also include, for example, receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."
[0071] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."
[0072] Any reference to elements using the designations “first,” “second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements do not imply that only two elements may be employed, or that the first element must precede the second element in any way.
[0073] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.
[0074] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.
[0075] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different." [Explanation of symbols]
[0076] 10...Synchronization control unit, 11...Detection unit, 12...Speech recognition unit, 12A...Speech recognition processing unit, 13...Translation unit, 13A...Translation processing unit, 14...Control unit, 15...Data storage unit, 20...Speech recognition API server, 30, 30A, 30B, 30C...Translation API server, 1001...Processor, 1002...Memory, 1003...Storage, 1004...Communication device, 1005...Input device, 1006...Output device, 1007...Bus.
Claims
1. A detection unit that detects a series of multiple audio segments in video content and detects the longest audio segment among the multiple audio segments, A speech recognition unit that obtains text data for each speech segment as a result of speech recognition processing targeting the audio of multiple detected speech segments, A translation unit that obtains translated subtitle data for each audio segment as a result of processing the acquired text data into another language, A control unit calculates the total processing time for the longest audio segment, including the processing time by the detection unit, the processing time by the speech recognition unit, and the processing time by the translation unit; delays the video content by the obtained total processing time; and adds translated subtitle data for each audio segment to the delayed video content before playback. A synchronous control device equipped with the following features.
2. Each of the aforementioned plurality of audio segments is assigned a sequence number, The speech recognition unit acquires text data for each speech segment accompanied by the sequence number, The translation unit acquires translated subtitle data for each audio segment accompanied by the sequence number, The control unit plays back the audio sections of the delayed video content, assigning translated subtitle data to each audio section, which is associated with the sequence number. The synchronous control device according to claim 1.
3. The translation unit acquires translated subtitle data for each audio segment and each language as a result of processing the acquired text data into multiple other languages. The control unit calculates the total processing time for the longest audio segment, which is the longest of the processing time by the detection unit, the processing time by the speech recognition unit, and the processing time for each language by the translation unit. The control unit then delays the video content by the obtained total processing time, and plays back the delayed video content with the translated subtitle data for each audio segment. The synchronous control device according to claim 2.
4. The aforementioned speech recognition is performed in series. Data transmission and reception processing with an external speech recognition server, Speech recognition processing by an external speech recognition server, and, Verification process by the speech recognition unit targeting the speech recognition results If it includes, The control unit calculates the total processing time by the speech recognition unit as the sum of the processing time for the transmission / reception process, the processing time for the speech recognition process, and the processing time for the confirmation process. The synchronous control device according to claim 1.
5. The aforementioned speech recognition is performed in parallel, at least partially. Data transmission and reception processing with an external speech recognition server, Speech recognition processing by an external speech recognition server, and, Verification process by the speech recognition unit targeting the speech recognition results If it includes, The control unit calculates the total processing time for the transmission / reception process, the voice recognition process, and the confirmation process as the processing time elapsed by the voice recognition unit. The synchronous control device according to claim 1.
6. When the translated subtitle data is displayed by scrolling in the video content, The control unit, during silent periods in the video content that are not audio sections, adds the translated subtitle data from the preceding audio section to the video content and plays it back. The synchronous control device according to claim 1.
7. The detection unit processes individual video content that has been divided into multiple parts along a time series. The synchronous control device according to claim 1.