Moving image automatic translation device and method
The real-time video translation device synchronizes translated audio with the original speech duration using multiple translation models and speech synthesis, addressing the discrepancy in speech length to provide a natural and accurate viewing experience.
Patent Information
- Application Number
- JP2024039827
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-29
AI Technical Summary
Existing real-time video translation technologies fail to synchronize audio and image accurately due to discrepancies in speech length caused by language translation, leading to an awkward viewing experience.
A real-time automatic video translation device that includes a speech recognition unit, automatic translation unit, and a dubbing buffer to store and synchronize translated texts with the original speech duration, using multiple translation models and speech synthesis to adjust playback speed for precise audio-image alignment.
Ensures synchronized and natural playback of translated audio with the video, maintaining the original content's integrity and reducing viewer discomfort.
Smart Images

Figure 2025140425000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to automatic speech translation technology, and more particularly to a real-time video automatic translation device and method that automatically translates speech accompanying a video in real time and plays it back in synchronization with the image. [Background technology]
[0002] Live streaming is an increasingly important communication technology. Live streaming is a type of video that includes audio and images and is used for news reports, speeches, virtual meetings, and more. Furthermore, with the development of global communication technologies, information sent by people in different countries can now be viewed and heard simultaneously around the world.
[0003] However, a major obstacle to using live streaming across borders is the issue of language. Even if someone can see the images of a live stream, if they cannot understand the audio or can only partially understand it, it will be difficult for them to properly understand the content being streamed. As a result, the purpose of live streaming will be halved.
[0004] To solve this problem, technologies have been developed that translate the audio accompanying a video into another language, synthesize the audio, and replace the delayed audio with the synthesized audio. However, such technologies have the problem of not matching the image and the translated audio well due to the change in the length of the audio expression that occurs when translating languages.
[0005] A technique for solving such problems is disclosed in Patent Document 1, which is listed below. The technique disclosed in Patent Document 1 mainly relates to a technique for dubbing in television. In Patent Document 1, the length of text obtained by speech recognition is predicted, and if the predicted length is significantly longer than the length of the original utterance, a process is performed to change the text to be spoken to a shorter expression. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-322077 [Non-patent literature]
[0007] [Non-Patent Document 1] Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021.Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021 Summary of the Invention [Problem to be solved by the invention]
[0008] Literature 1 states that if the text after speech translation is significantly longer than the original speech, the text expression is shortened. However, the details of the method for shortening the text expression are not disclosed. Furthermore, there is no guarantee that the shortened expression will accurately correspond to the audio of the original video, or that the synthesized audio will be of an appropriate length. Even for utterances with the same content, the difference in length between languages can be significant. Therefore, if the technology disclosed in Literature 1 is applied to real-time streaming, there is a high possibility that the final audio and video will not correspond correctly. As a result, viewers may feel uncomfortable with the dubbed video and may not be able to properly evaluate the video.
[0009] Therefore, an object of the present invention is to provide a real-time automatic video translation device and method that converts the audio of a video into audio in another language without creating an awkward feeling. [Means for solving the problem]
[0010] A first aspect of the present invention provides an automatic video translation device that includes a speech recognition device that outputs a plurality of texts in a first language divided into predetermined processing units by speech recognition of speech waveforms from a video that includes images and speech waveforms in a first language; an automatic translation unit that outputs a plurality of translated texts in a second language by translating each of the plurality of texts in the first language into a second language; and a dubbing buffer for storing unit videos corresponding to each of the processing units, wherein the unit videos include a plurality of translated texts obtained by the automatic translation unit from the processing units included in the unit videos, a predicted duration at the time of speech synthesis for each of the plurality of translated texts, an image and speech waveform in the first language corresponding to the processing unit in the video, and the duration of speech based on the speech waveform in the first language of the processing unit, and the automatic video translation device further includes a playback unit that selects, from the plurality of translated texts included in the first unit video stored at the beginning of the dubbing buffer, a translated text that has a predicted duration closest to the duration of speech based on the speech waveform in the first language of the processing unit, synthesizes speech, and plays the synthesized speech in synchronization with the image of the processing unit.
[0011] Preferably, the playback unit includes: a target playback speed ratio determination means for determining a target playback speed ratio to be a target for playing back the audio included in the first unit video based on the total duration of the audio of the unit videos stored in the dubbing buffer; a standard duration conversion means for using the target playback speed ratio to convert the duration of the audio of the unit videos to a standard duration, which is the duration when played back at a playback speed determined by the target playback speed ratio; a selection means for selecting from a plurality of translated texts one whose predicted duration is closest to the standard duration; a playback speed adjustment means for adjusting the playback speed of the translated text selected by the selection means within a predetermined allowable range of audio playback speed based on the ratio between the predicted duration and the standard duration of the translated text selected by the selection means; and a speech synthesis means for synthesizing speech based on the playback speed adjusted by the playback speed adjustment means, based on the translated text selected by the selection means.
[0012] More preferably, the speech synthesis means synthesizes the speech of the translation text using speech embeddings and speaker embeddings extracted from speech waveforms in the first language included in the unit moving image corresponding to the processing unit.
[0013] More preferably, the speech synthesis means outputs a predicted duration of the phonemes constituting each of the plurality of translated texts, and the selection means selects one of the translated texts using the predicted duration output by the speech synthesis means.
[0014] Preferably, the speech synthesis means includes phoneme output means for identifying and outputting a plurality of phonemes that constitute the selected translation text; duration prediction means for outputting a predicted duration of each of the plurality of phonemes output by the phoneme output means; and optimization means for optimizing the value of each of the plurality of phonemes' duration so that the sum of the ceiling values of the predicted durations of each of the plurality of phonemes approaches a value obtained by dividing the sum of the predicted durations by a reference playback speed ratio, the reference playback speed ratio being the ratio between the predicted duration of the selected translation candidate and the standard duration, and being a value limited within a range between a lower limit and an upper limit of the speech playback speed adjustment range.
[0015] A video automatic translation method according to a second aspect of the present invention includes the steps of: a computer outputting a plurality of texts in a first language divided into predetermined processing units by performing speech recognition on speech waveforms of a video including images and speech waveforms in a first language; a computer outputting a plurality of translated texts in a second language by translating each of the plurality of texts in the first language into a second language; and a computer storing unit videos corresponding to each of the predetermined processing units in a dubbing buffer, wherein the unit videos are generated by extracting a plurality of translation texts from the processing units included in the processing units. The output includes a plurality of translated texts obtained in the outputting step, a predicted duration after speech synthesis for each of the plurality of translated texts, an image and a speech waveform in the first language corresponding to the processing unit in the video, and the duration of the speech based on the speech waveform in the first language of the processing unit.The automatic video translation method further includes a step in which the computer selects, from the plurality of translated texts included in the first unit video stored at the beginning of the dubbing buffer, a translated text having a predicted duration closest to the duration of the speech based on the speech waveform in the first language of the processing unit, synthesizes the speech, and plays the synthesized speech in synchronization with the image of the processing unit.
[0016] The above and other objects, features, aspects and advantages of the present invention will become apparent from the following detailed description of the invention taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a block diagram showing the input / output relationship of a real-time automatic video translation system 50 according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the functional configuration of the real-time video machine translation system. [Figure 3] FIG. 3 is a schematic diagram showing an example of data storage in the dubbing buffer. [Figure 4] FIG. 4 is a diagram showing the data structure of data corresponding to each sentence stored in the dubbing buffer. [Figure 5] FIG. 5 is a flowchart showing a control structure of a program for implementing the functions of real-time video machine translation system 50 by a computer. [Figure 6] FIG. 6 is a flowchart showing a control structure of a routine for determining a preferred playback speed ratio in the program shown in FIG. [Figure 7] FIG. 7 is a flowchart showing a control structure of a routine that determines which translation to adopt in the program shown in FIG. [Figure 8] FIG. 8 is a block diagram showing a functional configuration of the speech synthesizer shown in FIG. [Figure 9] FIG. 9 is a block diagram showing the functional configuration of a conventional voice synthesizer. [Figure 10] FIG. 10 is a block diagram showing a functional configuration of the speed adjustment processing unit shown in FIG. [Figure 11] FIG. 11 is an external view of a computer system for realizing the real-time video machine translation system shown in FIG. [Figure 12] FIG. 12 is a block diagram showing the configuration of the computer system shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0018] In the following description and drawings, the same components are designated by the same reference numerals. Therefore, detailed description thereof will not be repeated. In the following embodiments, "dubbing" refers to translating the audio of a video into another language, so-called "voiceover." Furthermore, a video includes image data and audio data. However, in the following, for the sake of simplicity, image data will be simply referred to as "image" and audio data will be simply referred to as "audio."
[0019] 1. First embodiment 1.1 Configuration FIG. 1 shows the relationship between input and output of a real-time video machine translation system 50 according to the first embodiment. Referring to FIG. 1, the input to the real-time video machine translation system 50 is a video 52, which includes an image 60 and audio 62 in a source language (hereinafter referred to as the "source language"). The output from the real-time video machine translation system 50 is also a video 54, which includes an image 64 that is substantially the same as the video 52 and audio 66 in a language to which it is translated (hereinafter referred to as the "target language"). For example, the audio 62 is in English and the audio 66 is in Japanese. There are no particular restrictions on the languages, and the languages may be similar to each other, such as different dialects of the same language.
[0020] The real-time video machine translation system 50 converts each utterance of the speech 62 into speech 66 in the target language and outputs it together with the image 64. At this time, the real-time video machine translation system 50 adjusts the speaking time of the speech 66 to match the speaking time of the speech 62 as closely as possible. In addition, the real-time video machine translation system 50 uses, as the original text for the speech 66, a text selected from a plurality of translation texts for the speech 62 in accordance with predetermined criteria.
[0021] However, the real-time video machine translation system 50 differs from that described in Patent Document 1 in the following respects: The real-time video machine translation system 50 optimizes the playback of the audio 66 as follows so that the dubbed video 54 sounds natural and expresses the content of the video 52 with high reproducibility.
[0022] (1) The real-time video machine translation system 50 selects from among multiple translation texts obtained by translating an utterance in a source language using multiple machine translation models. This selection is based on the criterion of matching the speech time after speech synthesis with the speech time in the source language as closely as possible.
[0023] (2) The real-time video automatic translation system 50 further controls the playback time of the speech waveform synthesized from the selected translation text so that the difference between the playback time and the speaking time of the source language becomes smaller.
[0024] (3) The real-time video automatic translation system 50 further adjusts the playback speed of the image 64 and the audio 66 so that the dubbed video does not cause discomfort to the viewer.
[0025] In addition to these, this embodiment has various other features, which will be described in the following description of the embodiment.
[0026] 2, the real-time video machine translation system 50 includes a speech recognition device 102 that receives and recognizes speech 62 and outputs text in a source language, and a sentence segmentation unit 104 that segments the text output by the speech recognition device 102 into sentences. The real-time video machine translation system 50 further includes an automatic translation unit 106 that generates multiple translation texts for each sentence output by the sentence segmentation unit 104, and a first-in, first-out dubbing buffer 100 that stores the image 60, the speech 62, the original text output by the sentence segmentation unit 104, and the multiple translations generated by the automatic translation unit 106, as described below. In the following description, the text is segmented by sentence, but the text may be segmented by sentence or by units called translatable chunks that are shorter than sentences.
[0027] The real-time video machine translation system 50 further includes a speech synthesizer 108 that, for each sentence stored in the dubbing buffer 100, selects a translation text from several translation texts that has a duration close to that of the speech 62 and performs speech synthesis while adjusting the duration; a dubbing playback controller 110 that plays back the image of each sentence stored in the dubbing buffer 100 and the speech synthesized by the speech synthesizer 108 and outputs them as image 64 and speech 66; and a video synthesis unit 112 that synthesizes the image 64 and speech 66 in synchronization with each other and outputs them as video 54.
[0028] In this embodiment, the automatic translation unit 106 includes a first translation model 130, a second translation model 132, and a third translation model 134. In this embodiment, these models have different parameters and often output different translation results from the same source language sentence. These models can be realized, for example, by having different network structures or network hyperparameters, or even if the network structure and hyperparameters are the same, by using different initial values for the model during training, or by using different training data and training methods.
[0029] The automatic translation unit 106 further has the function of causing each of the first translation model 130, the second translation model 132, and the third translation model 134 to create multiple translation texts for each sentence segmented by the sentence segmentation unit 104, and providing these to the dubbing buffer 100. In this embodiment, the first translation model 130, the second translation model 132, and the third translation model 134 each output the top three translation texts (selected based on the probability that they are suitable as translations). These translation texts are hereinafter referred to as "translation candidates." As a result, the translation generation unit 136 outputs nine translation candidates for each sentence output by the sentence segmentation unit 104.
[0030] The speech synthesizer 108 performs speech synthesis using a mechanism similar to that of conventional VITS (Variational Inference with Adversarial Learning for End-to-End Text-to-Speech) TTS (Text-To-Speech), but differs from conventional systems in the part that determines the duration of each phoneme in the utterance to be generated. This difference will be described later with reference to Figures 8 and 9.
[0031] 3 shows how video data corresponding to each sentence (which corresponds to a processing unit and will hereinafter be referred to as a "video unit") is stored in the dubbing buffer 100 at the start of processing of the audio 62. Referring to FIG. 3, video units 170, 172, and 174 are stored in the dubbing buffer 100 in this order. Of course, as processing progresses and the sentence division unit 104 outputs a new sentence, the data for that sentence is added to the end of the dubbing buffer 100, while the first sentence is read and processed, and the data is deleted from the dubbing buffer 100. Furthermore, the playback speed can be adjusted by changing the ratio between the speed at which data is written to the dubbing buffer 100 and the speed at which data is read from the dubbing buffer 100.
[0032] FIG. 4 shows the structure of a unit video 200 (data structure of unit video) generally relating to the i-th sentence, which is stored in the dubbing buffer 100 (see FIG. 3). i ) is the image 210 (IMAGE i ) and the speech waveform 212 (WAVE) of the source language speech 62 corresponding to the sentence. i ) and the duration of the sound 214(O i ) and translation candidate information 216 relating to a plurality of (nine in this embodiment) translation candidates generated by the automatic translation unit 106 for the sentence.
[0033] The translation candidate information for each translation candidate is i and its TTS predicted duration u i Since there are nine translation candidates, the translation candidate information is the following pair (p i1 , u i1 ),…,(p i9 , u i9) In this embodiment, each predicted TTS duration is generated using the function of a speech synthesizer (described later) when generating the translated text output from the first translation model 130, the second translation model 132, and the third translation model 134, and is stored in association with the translated text. This predicted TTS duration is assigned by the speech synthesizer when each translated text is output from each model. Other methods for generating the predicted TTS duration can also be used; for example, a configuration in which the translation unit predicts the standard duration of the generated translation candidates can be considered.
[0034] 5 is a flowchart showing the control structure of a program for implementing, on a computer, the dubbing playback controller 110 of the real-time video machine translation system 50. By executing the process of this program once, processing is executed on the unit moving image stored at the top of the dubbing buffer 100. The following description will be given assuming that this unit moving image is the ith unit moving image.
[0035] The parameters required to run this program are as follows:
[0036] A lower threshold T1 (e.g., 2.0 seconds) and an upper threshold T2 (e.g., 8.0 seconds) for the duration of the dubbing buffer 100. Note that the duration of the dubbing buffer 100 refers to the total duration of the images of each unit video stored in the dubbing buffer 100.
[0037] - Audio playback speed adjustment range lower limit a1 (e.g. 0.98) and upper limit a2 (e.g. 1.04) - Lower limit v1 (e.g. 0.96) and upper limit v2 (e.g. 1.08) of the image playback speed adjustment range For users, fluctuations in audio playback speed are more noticeable than fluctuations in image playback speed, so v1 <a1<1.0<a2<v2となるようにしてある。
[0038] 5, this program 250 includes step 260, which performs initial processing when the program starts to run, step 262, which opens dubbing buffer 100 shown in FIG. 2, and step 264, which reads out unit moving images from dubbing buffer 100. However, at this stage, the data in dubbing buffer 100 is not deleted. For the sake of simplicity in the following explanation, dubbing buffer 100 will be referred to simply as the "buffer."
[0039] The program 250 further calculates the total duration of the images of the unit video present in the buffer. S The calculation in step 266 is performed as follows: Referring to Figures 3 and 4, the durations of the audio of the video units 170, 172, and 174 are O1, O2, and O3, respectively. In step 266, the duration of the buffer O S Calculate O1 + O2 + O3. Note that the duration of the image and the duration of the audio in the unit video in the buffer are essentially the same, so they are not distinguished from each other.
[0040] Returning to FIG. 5, the program 250 further calculates the buffer duration O(s) calculated in step 266 based on the data read in step 264. S and determining 268 a target playback speed ratio r, which is a preferred playback speed ratio of the speech in the target language, based on
[0041] 6 shows the control structure of the routine executed in step 268. Referring to FIG. 6, the routine in step 268 determines the duration of the buffer O S is smaller than the lower limit threshold T1 of the buffer duration, and if the determination in step 320 is affirmative, step 322 substitutes the lower limit v1 of the image playback speed adjustment range for the target playback speed ratio r and returns to the original routine of FIG. 5.
[0042] The routine further includes, in response to a negative determination in step 320, determining the duration of the buffer O. S is greater than the upper limit threshold value T2 of the duration; step 326, in response to the affirmative determination in step 324, assigning the upper limit v2 of the image playback speed adjustment range to the target playback speed ratio r and returning to the original routine; and step 328, in response to the negative determination in step 324, assigning 1.00 to the target playback speed ratio r and returning to the original routine.
[0043] This process limits the target playback speed ratio r to between the lower limit v1 and the upper limit v2 of the image playback speed adjustment range. This target playback speed ratio r becomes a desirable playback speed ratio.
[0044] 5, program 250 further includes, following step 268, a step 270 for determining a standard duration d, which is a preferred playback time, by converting the duration of the translated text using the target playback speed ratio r determined in step 268. In step 270, the standard duration d is calculated as follows: d=0 i / r. As a result, if the target playback speed ratio r is large, the standard duration d will be short. If the target playback speed ratio r is small, the standard duration d will be long.
[0045] The program 250 further calculates the standard duration d obtained by the above process and the duration O of the audio of the unit moving image to be processed. i The method includes step 272 of determining one translation candidate from among the sub-number of translation candidates based on the translation candidate. The details of step 272 are shown in FIG.
[0046] Referring to FIG. 7, the routine for implementing step 272 includes step 350, which executes process 352 for calculating a score for each of the nine translation candidates, and step 354, which adopts the translation candidate with the highest score calculated in step 350 and returns to the routine of FIG. 5.
[0047] The routine of process 352 assumes that the translation candidate to be processed is the i-th one, and calculates the predicted duration u of the j-th translation candidate for the i-th unit video. ij Step 370 branches the flow of control according to whether or not the j-th translation candidate is shorter than the standard duration d. If the determination in step 370 is affirmative, j Niu ij Step 372 substitutes / d and terminates the execution of the process 352; and when the determination in step 370 is negative, the score j 1.0-α(u ij -d) / d and terminates the execution of the process 352. i is the predicted TTS duration of the translation candidate being processed.
[0048] The above process is as follows: First, when the predicted TTS duration of a translation candidate is shorter than the standard duration d, the score of the translation candidate is a monotonically increasing function of the predicted TTS duration. When the predicted TTS duration of a translation candidate is longer than the standard duration d, the score of the translation candidate is a monotonically decreasing function of the predicted TTS duration. As a result, the translation candidate whose predicted TTS duration is closest to the standard duration d is adopted. Here, it is assumed that the kth translation candidate is adopted.
[0049] Referring again to FIG. 5, the program 250 continues at step 272 and calculates the standard duration d, the TTS predicted duration u of the accepted translation candidate, and i , and the lower limit a1 and upper limit a2 of the audio playback speed adjustment range are used, and the reference playback speed ratio a r The reference playback speed ratio a is determined in step 274. r is a coefficient used to convert the audio playback speed so that it falls within an acceptable range. r is determined by the following formula:
[0050] a r =min(max(d / u k ,a1),a2) That is, the reference playback speed ratio a r is the TTS predicted duration u of the selected translation candidate (kth) k The standard playback speed ratio a is a ratio of the standard duration d to the standard playback speed d, and is limited to a value within the range between the lower limit a1 and the upper limit a2 of the audio playback speed adjustment range. r is used to change the speech duration of the selected translation candidate.
[0051] The program 250 further calculates the reference playback speed ratio a determined in step 274. r the step 276 of synthesizing speech by VITS TTS based on the text of the selected translation candidate using the above, the step 278 of synchronously playing back the video in the i-th unit video and the synthesized speech, and the step 280 of updating the dubbing buffer 100 by deleting the i-th unit video from the dubbing buffer 100, and then returning control to step 264 to start processing the (i+1)-th unit video. Note that unit videos of each sentence of the speech recognition result from the input video 52 are added to the dubbing buffer 100 as needed.
[0052] The step 278 of synchronizing the image and the audio includes adjusting the playback speed of the image based on the duration of the audio after audio synthesis. r The actual duration of the adjusted synthesized voice is expressed as a u Then, the duration of the image v u is defined by the following formula:
[0053] v u =max(min(O1 / v1,a u ),O1 / v2) That is, the duration of the image is, in principle, equal to the actual duration of the adjusted synthesized speech a u is used, but the value is limited within the range of the lower limit time obtained by dividing the duration O1 of the original image by the lower limit v1 of the image playback speed, and the upper limit time obtained by dividing the duration O1 of the original image by the upper limit v2 of the image playback speed. r is the duration of the original image O1 and the v calculated above. uand are calculated according to the following formula:
[0054] v r =O1 / v u In this way, by setting the duration of the image (playback speed) based on the duration of the synthesized audio, the duration of both can be made approximately the same, and by matching the start of audio and image readout, synchronized output can be achieved.
[0055] VITS TTS is used for the speech synthesizer 108 (corresponding to step 278 in FIG. 5) according to this embodiment. FIG. 8 shows the functional configuration of a program that realizes the speech synthesizer 108. Note that this speech synthesizer 108 is an improvement over the conventional VITS TTS, with the following improvements: First, the duration of the synthesized speech is made closer to the duration of the original utterance. Second, the synthesized speech has phonetic characteristics similar to those of the speech in the input source language. Third, for each translation candidate, information about the duration of each phoneme in the synthesized speech is provided to the dubbing playback controller 110. Therefore, the conventional VITS TTS will first be described.
[0056] Figure 9 shows a functional block diagram of a conventional VITS TTS. VITS TTS is disclosed in Non-Patent Document 1, which also describes a training method for each component of VITS TTS. The disclosure of Non-Patent Document 1 is incorporated herein by reference. VITS TTS consists of a neural network trained by adversarial learning.
[0057] Referring to FIG. 9, a conventional VITS TTS speech synthesizer 500 includes an embedding 510 that has been trained from a corpus (e.g., speech data by a specific speaker and its text) and indicates the acoustic characteristics of the speech of each utterance in the corpus, a text encoder 414 that receives text 412 to be synthesized and outputs an intermediate feature representation of the speech, and a trained projection layer 416 that receives the intermediate feature representation output by the text encoder 414 and the embedding 510, identifies a plurality of phonemes that make up the text 412, and outputs a phoneme sequence 418.
[0058] The speech synthesizer 500 further includes a probabilistic duration predictor for outputting a predicted duration 422 for each phoneme in the phoneme sequence 418 to be synthesized based on the intermediate feature representation output by the text encoder 414 and a pre-trained probability distribution. The predicted duration 422 output by the probabilistic duration predictor 420 is obtained as a real number with one decimal point. The speech synthesizer 500 applies a ceiling function to each value constituting the predicted duration 422. That is, as shown in FIG. 9, 1.8 and 1.2 are both converted to 2, and 0.9 is converted to 1. The units of each value are a predetermined constant time for speech synthesis. By applying the duration 512 after applying the ceiling function to each phoneme in the phoneme sequence 418, each phoneme is converted into an output phoneme sequence 514 consisting of as many copies as indicated by the duration 512.
[0059] The speech synthesizer 500 further includes a flow layer 432 that receives the output phoneme sequence 514, inversely generates and outputs acoustic features corresponding to each phoneme constituting the output phoneme sequence 514, and a decoder 434 that decodes the acoustic features output by the flow layer 432 and outputs a waveform signal 516. The flow layer 432 is trained in advance to generate acoustic features corresponding to the output phoneme sequence 514.
[0060] The configuration of the speech synthesizer 108 according to this embodiment shown in Fig. 8 will be described below in comparison with Fig. 9. The speech synthesizer 108 differs from the speech synthesizer 500 shown in Fig. 9 in the following three points.
[0061] The first embodiment includes a linear mapping layer 402 that performs linear mapping on speech embeddings 400, which represent the acoustic characteristics of speech extracted from the speech, instead of the embedding 510 in FIG. 9 ; a linear mapping layer 406 that performs linear mapping on speaker embeddings 404, which represent the acoustic characteristics of the speaker of the speech; and an adder 408 that adds the outputs of the linear mapping layer 402 and the linear mapping layer 406 and outputs the combined embedding 410. The speech embeddings 400 are sequentially generated by an embedding unit (not shown) using the speech 62 shown in FIG. 2 and the results of sentence segmentation by the sentence segmentation unit 104. The speaker embeddings 404 are obtained by adding values obtained by multiplying the speech embeddings 400 of each utterance by the same speaker by the duration of the utterance. In this embodiment, a mel spectrogram is used as the embedding. However, other features, such as a row spectrogram, can also be used as the embedding.
[0062] The speech synthesizer 108 further includes the same text encoder 414, projection layer 416, and probabilistic duration predictor 420 as shown in Figure 9, which receive the combined embedding 410 as an embedding. However, this embodiment differs from Figure 9 in that the predicted duration 422 output by the probabilistic duration predictor 420 for each translation candidate is provided to the dubbing buffer 100 shown in Figure 2 and used as the estimated duration of the sentence to be processed.
[0063] The speech synthesizer 108 also differs from the speech synthesizer 500 of Fig. 9 in that it includes a speed adjustment processor 426 for optimizing and adjusting the speech duration, i.e., playback speed, of a translation candidate selected by the dubbing playback controller 110 shown in Fig. 2 within an allowable range of speech playback speed based on the calculation result of the ratio between the predicted duration and standard duration of the selected translation text. In the example shown in Fig. 8, the predicted duration 422 value of (1.8 1.2 0.9) is adjusted by the speed adjustment processor 426 to an estimated duration 428 of (2 1 1), which is different from the duration 512 in Fig. 9.
[0064] The estimated duration 428 is applied to the sequence of phonemes 418 output by the projection layer 416 to generate a sequence of phonemes 430 according to the adjusted duration. The speech synthesizer 108 further includes a flow layer 432 and a decoder 434 for generating a speech waveform 436 based on the sequence of phonemes 430.
[0065] In this embodiment, the speaker embedding is calculated as follows. For example, assume that unit moving images 170, 172, and 174 are stored in the dubbing buffer 100 as shown in FIG. 3. As can be seen from FIG. 4, the audio waveforms 212 contained in the unit moving images 170, 172, and 174 are the first utterance, the second utterance, and the third utterance, respectively. These utterance embeddings are e1, e2, and e3, respectively. Furthermore, their durations are d1, d2, and d3, respectively. Then, in this embodiment, the speaker embedding e spk is calculated using the following formula:
[0066] e spk =(e1*d1+e2*d2+e3*d3) / (d1+d2+d3) 10 is a flowchart showing a control structure of a program for implementing processing for adjusting the duration of audio, executed by dubbing playback controller 110. This program selects a translation candidate for the i-th utterance, and adjusts the standard playback speed ratio a ris determined, it is executed. As hyperparameters of this program, it is necessary to specify the maximum number of iterations N for the iteration process to search for the upper limit value of the phoneme in the program, and the value of the search step during the iteration (corresponding to the learning rate in machine learning) as β (0<β<1). For example, N=100. β is set according to the application.
[0067] The inputs to this program are the output of the probabilistic duration predictor 420 for the selected translation candidate, a vector of predicted durations of synthesized speech (a vector with predicted durations of each phoneme as elements) (V1, V2, V3, ...), and a reference playback speed ratio a r The output of this program is an audio waveform with the desired duration.
[0068] In FIG. 10, u' is the speed at which u is reduced to the reference playback speed ratio a r The value obtained by dividing V1 by V2 and rounding it off is shown as u. The value obtained by adding the ceiling function values of V1, V2, and V3 is shown as u = ceil(V1) + ceil(V2) + ceil(V3) + .... The program includes step 550, which performs initial processing, and step 552, which begins each iteration.
[0069] In step 550, the variable x representing the magnification of the duration of the output of the phoneme to be processed is set to the reference playback speed ratio a r In step 550, 0 is further substituted into a variable count for controlling the number of repetitions.
[0070] In step 552, the duration value u corresponding to the value of the variable x is calculated. x is calculated by the following formula:
[0071] u x =ceil(V1*x)+ceil(V2*x)+ceil(V3*x)+… In step 552, the error between ux and u' is further calculated. x) is calculated. In addition, 1 is added to the value of the variable count.
[0072] The program then continues to step 552 and calculates the error error(u x ) is 0 or not, step 560 generates a speech waveform of the target phoneme so that the duration is determined by the variable x at that time when the determination in step 554 is affirmative, and step 556 branches the control flow depending on whether the value of the variable count is greater than N or not when the determination in step 554 is negative.
[0073] The program further stores the error (u) so far in the variable x if the determination in step 556 is positive. x ) is given the smallest value, and control passes to step 560.
[0074] This program further includes step 558, which records the history of the repetitions when the determination in step 556 is negative, and further modifies the processing conditions after the repetitions, and returns control to step 552. In step 558, the following processing is performed: x ) is greater than or equal to u', the error error(u x The value of x that gives the minimum value of error(u) is stored in the variable x0. x ) is smaller than u', the error error(u x ) is stored in variable x1. Furthermore, when values are assigned to both the variables x0 and x1, (x0+x1) / 2 is assigned to variable x. When a value is assigned only to variable x0 and not to variable x1, x0*(1.0+β) is assigned to variable x. When a value is assigned only to variable x1 and not to variable x0, x1*(1.0-β) is assigned to variable x.
[0075] By performing this process, the duration of the utterance is determined according to the same concept as that used when selecting translation candidates.
[0076] 1.2 Operation The real-time video machine translation system 50 (see FIGS. 1 and 2), the configuration of which has been described above, operates as follows. Referring to FIG. 2, when a video 52 is provided, speech 62 is provided to the speech recognition device 102. The speech recognition device 102 performs speech recognition on the input speech 62 and provides text in the source language to the sentence segmentation unit 104.
[0077] The sentence dividing unit 104 divides the input text in the source language into sentences and provides them to the translation generating unit 136 and the dubbing buffer 100. At this time, the duration of the speech 62 is assigned to each sentence.
[0078] The translation generation unit 136 provides the received single sentence of source language text to the first translation model 130, the second translation model 132, and the third translation model 134. The first translation model 130, the second translation model 132, and the third translation model 134 each translate the input source language text into three target languages and provide the translation candidates to the translation generation unit 136. In other words, the translation generation unit 136 receives nine translation candidates. Each of these translation candidates is assigned a predicted duration (TTS predicted duration) when speech in the target language is generated from the translation candidate by TTS using the function of the speech synthesizer 108.
[0079] The translation generation unit 136 pairs the source language text provided by the sentence division unit 104 with each of the target language translation candidates received from the first translation model 130, the second translation model 132, and the third translation model 134, and provides the duration of the source language utterance and each predicted duration together to the dubbing buffer 100. The translation generation unit 136 further provides each translation candidate to the speech synthesizer 108.
[0080] 4, the dubbing buffer 100 stores, as a unit moving image 200, information about each translation candidate provided by the translation generation unit 136, the duration of the audio, an image of the period of image 60 corresponding to the duration of the audio, an audio waveform, and pairs of each translation candidate and its predicted TTS duration. The speech recognition device 102, the sentence division unit 104, and the automatic translation unit 106 all perform processing in response to the input of the moving image 52. As a result, unit moving images 170, 172, 174, ..., etc. are added to the dubbing buffer 100 in order.
[0081] 8, speech synthesizer 108 uses text encoder 414 and probabilistic duration predictor 420 to output predicted duration 422 of each phoneme constituting the text of the translation candidate for each translation candidate output by translation generation unit 136, and provides this to dubbing playback controller 110 as predicted duration 424 of synthesized speech. By providing predicted duration 424 of each translation candidate to dubbing playback controller 110 in this way, dubbing playback controller 110 can predict the duration of the synthesized speech without actually performing speech synthesis for each translation candidate. Speech synthesis is a burdensome process, so being able to predict the duration of synthesized speech without speech synthesis reduces the amount of calculation and speeds up the overall processing.
[0082] The speech recognition device 102, sentence division unit 104, automatic translation unit 106, speech synthesizer 108, and dubbing buffer 100 repeat the above-mentioned processes on the input speech. As a result, unit videos 200 for each utterance of the video 52 are stored in the dubbing buffer 100 in order on a first-in, first-out basis.
[0083] 5, when the dubbing playback controller 110 is started, it performs necessary initial processing in step 260 and then opens the dubbing buffer 100 in step 262. The dubbing playback controller 110 then reads all of the moving image units 200 stored in the dubbing buffer 100 in step 264. The dubbing playback controller 110 calculates the total duration of all images of the moving image units 200 stored in the dubbing buffer 100 in step 266.
[0084] The dubbing playback controller 110 further determines the target playback speed ratio r in step 268 by executing a routine whose control structure is shown in FIG.
[0085] The dubbing playback controller 110 further determines the standard duration d in step 270. More specifically, the dubbing playback controller 110 determines the standard duration d as follows: i / r(O i determines the standard duration d according to the duration of the speech of the speech to be processed).
[0086] 7 in the next step 272, the dubbing playback controller 110 determines the translation to be adopted. More specifically, the dubbing playback controller 110 determines the TTS predicted duration u calculated by the probabilistic duration predictor 420 of the speech synthesizer 108 (see FIG. 2) for each of the nine translation candidates j (j=1 to 9) of the i-th unit video 200. ij (Predicted duration 424 shown in FIG. 8) is used to perform process 352. As a result, for each translation candidate j (j=1 to 9), a score j The dubbing playback controller 110 calculates the score j The translation candidate that maximizes is adopted.
[0087] Referring again to FIG. 5, in step 274, the dubbing playback controller 110 calculates the reference playback speed ratio a according to the following formula: rDetermine.
[0088] a r =min(max(d / u k ,a1),a2) Standard playback speed ratio a r is the TTS predicted duration u of the selected translation candidate (kth) k and the standard duration d, where this value is limited to a range between a lower limit a1 and an upper limit a2 of the audio playback speed adjustment range.
[0089] In step 276, the dubbing playback controller 110 performs speech synthesis using the speech synthesizer 108 for the translation candidate adopted in step 272, using the reference playback speed ratio ar determined in step 274. As shown in FIG. 8, the speech synthesizer 108 generates an intermediate feature representation for the target text 412 using joint embedding 410. From this intermediate feature representation, a probabilistic duration predictor 420 outputs a predicted duration 422, and a speed adjustment processor 426 further adjusts the duration of the phonemes to generate an estimated duration 428. The speech synthesizer 108 outputs a time-adjusted phoneme sequence 430 in accordance with the estimated duration 428 for the phoneme sequence 418 output by the projection layer 416 based on the output of the text encoder 414. A flow layer 432 and a decoder 434 output a speech waveform 436 corresponding to the text 412 based on this phoneme sequence 430.
[0090] 5, in step 278, the dubbing playback controller 110 synchronizes and plays back the image and the sound of the sound waveform generated by the sound synthesizer 108 for the unit moving image 200 to be processed. u is calculated according to the following formula:
[0091] v u =max(min(O1 / v1,a u ),O1 / v2) where a u As mentioned above, the standard playback speed ratio ar This process results in the image duration v u is a value limited within the range of a lower limit time obtained by dividing the duration O1 of the original image by the lower limit v1 of the image reproduction speed, and an upper limit time obtained by dividing the duration O1 of the original image by the upper limit v2 of the image reproduction speed.
[0092] The dubbing / playback controller 110 further calculates the duration of the image v u The final image duration (playback speed) v is calculated using the following formula: r Calculate.
[0093] v r =O1 / v u As a result, the duration of the synthesized audio and the duration of the image (playback speed) can be made approximately the same. By matching the start of audio and image readout, the synthesized audio and image can be output in sync.
[0094] Thereafter, the dubbing playback controller 110 updates the dubbing buffer 100 by deleting the unit moving image (the unit moving image that was the subject of processing) at the top of the dubbing buffer 100 from the dubbing buffer 100. Thereafter, the dubbing playback controller 110 starts the processing from step 264 onwards for the new unit moving image at the top of the updated dubbing buffer 100.
[0095] 1.3 Effects of the embodiment As described above, the real-time video machine translation system 50 according to this embodiment generates multiple translation candidates for each utterance in the source language. Among these multiple translation candidates, the translation candidate that is considered to be closest in duration to the source language utterance to be processed is selected, and speech synthesis is performed based on this translation candidate. Furthermore, during the speech synthesis process, the duration of the phoneme waveform is adjusted so that the speech duration of the synthesized utterance is close to the duration of the source language utterance. Furthermore, to select a translation candidate, the output of the probabilistic duration predictor 420 is used to calculate the predicted TTS duration without actually synthesizing the speech waveform. Because the process of synthesizing a speech waveform takes time, this embodiment shortens the prediction of the duration of the translation candidate. As a result, the dubbed audio and video can be played back in sync with each other without creating a strange feeling.
[0096] In addition, speech synthesis is performed using both the source language speech embedding and the speaker embedding, which results in the dubbed speech reproducing the vocal characteristics of the source language speaker and the emotions expressed by the speech.
[0097] Therefore, in the case of live streaming, etc., the audio of a video can be converted into audio in another language without any sense of incongruity.
[0098] Second Variation In the above embodiment, the automatic translation unit 106 includes three translation models, each with different parameters. However, the present invention is not limited to such an embodiment. The translation models included in the automatic translation unit 106 may be capable of generating multiple different translation candidates. For example, the top n (n≧2) translation results from one translation model may be used as translation candidates. Alternatively, the top translation result from each of multiple translation models may be used as a translation candidate. Furthermore, when multiple translation models are used, the number of translation candidates used as the output of each translation model may not be equal.
[0099] In the above embodiment, the speaker embedding is provided to the text encoder 414, the projection layer 416, and the probabilistic duration predictor 420 in the form of a joint embedding 410. However, the present invention is not limited to such a form. Speaker embeddings can also be added to the flow layer 432 to add speaker characteristics to the speech.
[0100] 3. Computer implementation Fig. 11 is an external view of a computer system 600 that realizes the real-time video machine translation system 50 according to the first embodiment of the present invention shown in Figs. 1 and 2. Fig. 12 is a hardware block diagram of the computer system 600. The hardware configuration of the computer system 600 will be described below.
[0101] 11, this computer system 600 includes a computer 650 having a DVD (Digital Versatile Disc) drive 662, and a keyboard 654, a mouse 656, a monitor 652, a microphone 660, and a pair of speakers 658 for interacting with a user, all of which are connected to the computer 650. Of course, these are just examples, and any common hardware and software (e.g., a touch panel, voice input, or a general pointing device) that can be used for interacting with an operator can be used.
[0102] 11 and 12 , the computer 650 includes, in addition to a DVD drive 662, a CPU (Central Processing Unit) 710, a GPU (Graphics Processing Unit) 712, and a bus 720 connected to the CPU 710, the GPU 712, and the DVD drive 662. The computer 650 further includes a ROM (Read-Only Memory) 714 connected to the bus 720 and storing a boot-up program and the like of the computer 650, a RAM (Random Access Memory) 716 connected to the bus 720 and storing instructions constituting a program, a system program, working data, and the like, and an SSD (Solid State Drive) 718, which is nonvolatile memory, connected to the bus 720. The SSD 718 is used to store programs executed by the CPU 710 and the GPU 712, as well as data used by the programs executed by the CPU 710 and the GPU 712. The computer 650 further includes a network I / F (Interface) 726 that provides connection to a network enabling communication with other terminals, and a USB port 664 to which a USB (Universal Serial Bus) memory 702 can be attached / detached and that provides communication between the USB memory 702 and each part within the computer 650.
[0103] The computer 650 further includes an audio I / F 722 that is connected to the microphone 660, the speaker 658, and the bus 720, and has the function of reading out audio signals, image signals, and text data generated by the CPU 710 and stored in the RAM 716 or the SSD 718 according to instructions from the CPU 710, converting them to analog, amplifying them, and driving the speaker 658, and digitizing the analog audio signal from the microphone 660 and storing it at any address in the RAM 716 or the SSD 718 specified by the CPU 710.
[0104] In the above embodiment, the programs and the like that realize the various functions of the real-time video machine translation system 50 shown in Fig. 1 are all stored in, for example, the ROM 714, SSD 718, DVD 700, or USB memory 702 shown in Fig. 12, or in a storage medium of an external device (not shown) connected via the network I / F 726 and the network 704. Typically, these data and parameters are written to the SSD 718 from the outside, for example, and loaded into the RAM 716 when the computer 650 is executed.
[0105] 1 is stored on a DVD 700 inserted into a DVD drive 662 and transferred from the DVD drive 662 to the SSD 718. Alternatively, these programs may be stored on a USB memory 702, which may be inserted into a USB port 664 and transferred to the SSD 718. Alternatively, these programs may be transmitted to the computer 650 via the network 704 and stored in the SSD 718.
[0106] The program is loaded into RAM 716 at the time of execution. Machine learning models such as deep neural networks are used in the speech recognition device 102, sentence segmentation unit 104, first translation model 130, second translation model 132, third translation model 134, speech synthesizer 108, dubbing playback controller 110, etc. shown in Figure 2. In computer system 600, a machine learning model that has already been trained in another device may be used, or the computer system 600 may be used as a training device to train a machine learning model.
[0107] The CPU 710 reads a program from the RAM 716 according to an address indicated by an internal register called a program counter (not shown) and interprets the instructions. The CPU 710 reads data required to execute the instructions from the RAM 716, the SSD 718, or another device according to the address specified by the instruction, and executes the processing specified by the instruction. The CPU 710 stores the execution result data at an address specified by the program, such as in the RAM 716, the SSD 718, or a register within the CPU 710. Depending on the address, the execution result data is output from the computer to an external device via, for example, the network I / F 726. At this time, the program counter value is also updated by the program. The computer program may be loaded directly into the RAM 716 from the DVD 700, the USB memory 702, or via the network 704. Note that some tasks (mainly numerical calculations) of the program executed by the CPU 710 are issued to the GPU 712 according to instructions included in the program or according to the analysis results obtained when the CPU 710 executes the instructions. The first translation model 130, second translation model 132, third translation model 134, etc., associated with the automatic translation unit 106 can be processed in parallel with one another. Therefore, they are suitable for processing by the GPU 712.
[0108] The program that enables the computer 650 to realize the functions of each unit of the real-time video automatic translation system 50 (FIG. 2) according to the embodiment described above includes a plurality of instructions written and arranged to cause the computer 650 to operate to realize those functions. Some of the basic functions required to execute these instructions may be provided by an OS or third-party program running on the computer 650, various toolkit modules installed on the computer 650, or the program execution environment. Therefore, the program does not necessarily include all of the functions required to realize the system and method according to this embodiment. The program may include only instructions that execute the operations of the above-described devices and their components by statically linking appropriate functions or modules at compile time or by dynamically calling them at runtime in a controlled manner to achieve the desired results. The method for operating the computer 650 for this purpose is well known. Therefore, a description of the method for operating the computer 650 will not be repeated here.
[0109] The GPU 712 is capable of parallel processing, and can simultaneously execute a large amount of calculations associated with machine learning and inference in a parallel or pipelined manner. For example, parallel calculation elements discovered in a program when the program is compiled or when the program is executed are dispatched from the CPU 710 to the GPU 712 as needed, and executed. The results are returned to the CPU 710 directly or via a predetermined address in the RAM 716 and assigned to a predetermined variable in the program.
[0110] The embodiments disclosed herein are merely examples, and the present invention is not limited to the above-described embodiments. The scope of the present invention is defined by the claims in the appended claims, taking into consideration the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wordings described therein. [Explanation of symbols]
[0111] 50 Real-time video automatic translation system 52, 54 videos 60, 64, 210 images 62, 66 Audio 100 Dubbing Buffer 102 Voice recognition device 104 Sentence division part 106 Machine Translation Department 108,500 speech synthesizer 110 Dubbing playback controller 112 Video composition unit 214, 512 duration 216 Translation candidate information 412 Text 414 Text Encoder 416 Projection Layer 418, 430 phoneme string 420 Probabilistic Duration Predictor 422 Estimated Duration 426 Speed adjustment processing unit 428 Estimated Duration 514 output phoneme sequence 516 Waveform Signal
Claims
1. a speech recognition device that outputs a plurality of texts in the first language divided into predetermined processing units by speech recognition of the speech waveforms in a video including images and speech waveforms in the first language; an automatic translation unit that translates each of the plurality of texts in the first language into a second language, thereby outputting a plurality of translated texts in the second language; a dubbing buffer for storing unit moving images corresponding to each of the processing units, The unit video is the plurality of translated texts obtained by the automatic translation unit from the processing unit included in the unit video; a predicted duration during speech synthesis for each of the plurality of translated texts; the image corresponding to the processing unit in the video and the audio waveform in the first language; a duration of the speech of the speech waveform in the first language of the processing unit; The video automatic translation device further includes a playback unit that selects, from the plurality of translated texts included in the first unit video stored at the beginning of the dubbing buffer, the translated text having the predicted duration closest to the duration of the speech based on the speech waveform in the first language of the processing unit, synthesizes speech, and plays the synthesized speech in synchronization with the image of the processing unit.
2. The playback unit a target playback speed ratio determination means for determining a target playback speed ratio to be a target for playback of the audio included in the first unit moving image, based on the total duration of the audio of the unit moving images stored in the dubbing buffer; a standard duration conversion means for converting the duration of the sound of the unit moving image into a standard duration when the sound is played back at a playback speed determined by the target playback speed ratio, using the target playback speed ratio; a selection means for selecting, from the plurality of said translated texts, one whose predicted duration is closest to the standard duration; a playback speed adjusting means for adjusting the playback speed of the translation text selected by the selecting means within a predetermined allowable range of audio playback speed based on the ratio between the predicted duration of the translation text selected by the selecting means and the standard duration; 2. The video automatic translation device according to claim 1, further comprising: a speech synthesis means for synthesizing speech based on the translation text selected by the selection means and the playback speed adjusted by the playback speed adjustment means.
3. 3. The video automatic translation device according to claim 2, wherein the speech synthesis means synthesizes the speech of the translation text using speech embeddings and speaker embeddings extracted from the speech waveform of the first language included in the unit video corresponding to the processing unit.
4. the speech synthesis means outputs, for each of the plurality of translated texts, the predicted duration of the phonemes constituting the translated text; 3. The video automatic translation device according to claim 2, wherein said selection means selects one of said translation texts using said predicted duration output by said speech synthesis means.
5. The voice synthesis means phoneme output means for identifying and outputting a plurality of phonemes that make up the selected translation text; duration prediction means for outputting the predicted duration of each of the plurality of phonemes output by the phoneme output means; optimizing means for optimizing the duration value of each of the plurality of phonemes so that a sum of the ceiling values of the predicted durations of each of the plurality of phonemes approaches a value obtained by dividing the sum of the predicted durations by a reference playback speed ratio; 3. The video automatic translation device according to claim 2, wherein the reference playback speed ratio is a ratio between the predicted duration of the selected translation candidate and the standard duration, and is a value limited within a range between a lower limit and an upper limit of an adjustment range of the audio playback speed.
6. a step in which a computer outputs a plurality of texts in the first language divided into predetermined processing units by performing speech recognition on the speech waveforms of a video including images and speech waveforms in the first language; a computer translating each of the plurality of said texts in the first language into a second language to output a plurality of translated texts in the second language; a step of storing unit videos corresponding to each of the predetermined processing units in a dubbing buffer by a computer, The unit video is the plurality of translated texts obtained in the step of outputting the plurality of translated texts from the processing units included in the processing unit; a predicted duration after speech synthesis of each of the plurality of translated texts; the image corresponding to the processing unit in the video and the audio waveform in the first language; a duration of the speech of the speech waveform in the first language of the processing unit; The video automatic translation method further includes a step in which a computer selects, from the plurality of translation texts included in the first unit video stored at the beginning of the dubbing buffer, the translation text having a predicted duration closest to the duration of the speech based on the speech waveform in the first language of the processing unit, synthesizes speech, and plays the synthesized speech in synchronization with the image of the processing unit.
Citation Information
Patent Citations
Television device
JP2000322077A