Video translator support system and dubbed video edition system
The video translator support system and dubbing video editing system address the challenge of producing high-quality dubbed videos by using AI-driven speech recognition and synthesis to generate and refine multilingual content, ensuring natural and skillful dubbing at a lower cost for small businesses.
Patent Information
- Application Number
- JP2024080452
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-28
AI Technical Summary
Small businesses lack the resources and expertise to produce high-quality dubbed videos that maintain the appeal of the original content and provide foreign language audio that matches the visuals, due to the limited availability of skilled translators and high costs.
A video translator support system and dubbing video editing system that utilize AI-driven speech recognition, translation, and voice synthesis to generate and refine multilingual text and audio data, synchronized with lip movements and other visual cues, reducing the workload of translators and enabling high-quality dubbing at a lower cost.
The systems significantly reduce the labor and cost required for producing high-quality dubbed videos, allowing small businesses to create natural and skillful dubbed content efficiently, even without expert translators.
Smart Images

Figure 2025174280000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video translator support system and a dubbing video editing system. The video translator support system of the present invention is specifically a system that supports translation work essential for producing high-quality dubbing videos. The dubbing video editing system of the present invention is specifically a video production system that uses the video translator support system. [Background technology]
[0002] For many years, films and broadcasts have been dubbed, using the original images, videos, and music, while translating the spoken language into another language. For example, when dubbing a film, it is necessary to ensure that the dialogue in each scene is the correct length for the original footage, that the dialogue matches the characteristics of the film, characters, and actors and is understandable to viewers of the dubbed work, and that the subtitles are easy to read and read. In the case of broadcast works such as news programs, highly technical terms must be translated into concise expressions that are easy for viewers to understand, both in the audio and subtitles. This type of dubbing requires advanced skills, and film and broadcast dubbing is supported by a small number of skilled translators.
[0003] In recent years, the spread of internet video distribution sites and smartphones has made it possible for anyone to watch a huge amount of video produced for a variety of purposes, including entertainment, advertising, education, public service, etc. Until now, promotional videos (commercials) were limited to businesses that could afford large expenses, but today even individuals and small businesses can distribute unique videos online for advertising and profit-making purposes.
[0004] Furthermore, with today's globalization of people and goods movement and expectations for inbound tourism in Japan, the need to provide video with audio in multiple languages is rapidly increasing. For example, for businesses such as restaurants and small shops, videos featuring the owner themselves are ideal for promoting unique services and products, and are expected to attract more customers. Many businesses hope to appeal to foreign tourists by dubbing these unique original videos into foreign languages. For example, transportation and accommodation facilities may also require guide videos with multilingual audio. Furthermore, when streaming game videos and animation, high-quality dubbed versions must be released shortly after the original version is released to prevent piracy.
[0005] However, even for relatively short videos used for advertising or informational purposes, maintaining the quality of the original video through dubbing requires the same high level of translation skills as traditional film or broadcast dubbing. Despite recent improvements in machine translation software, machine-translated text can still sound unnatural. Moreover, it is often impossible for the characters in the original video to read the machine-translated text in a natural way. Simply displaying the translated text as subtitles often fails to convey the atmosphere of the original video. For this reason, creating satisfactory dubbed videos is not easy for small businesses that cannot afford the high costs or the hiring of experts, posing a business risk.
[0006] Previously proposed dubbing video production technologies are not easily usable by small businesses such as restaurants and shops. For example, Patent Document 1 describes a cultural school system that translates course content videos stored on a cultural school's server computer into a language selected by the client and distributes the translated videos to the client computer. This system is intended for relatively large cultural schools. This system allows users to select courses for which translated videos are provided, taking into account the course opening date and number of students, and allows students to cover the production costs. However, it is nearly impossible for small businesses that want to distribute promotional videos to an unspecified number of viewers in a short period of time to apply this system.
[0007] For example, Patent Document 2 describes a video editing system that uses text and audio superimposed on an original educational video to create a foreign language version of the video, which is then used to provide technical training to foreign workers. This video editing system allows users to create dubbed educational videos with simple editing operations on their devices. Since the original language of such technical training videos primarily uses specific technical terms and jargon, translating the original text is not difficult. Furthermore, since there is little need for the individuality of the actors featured or for the video to be entertaining, the quality requirements for the dubbed video are not high. However, small businesses distributing the promotional videos described above typically expect the dubbed video to maintain the appeal of the original video and to feature foreign language audio that matches the visuals. The video editing system described in Patent Document 2 cannot meet such high requirements. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-302930 [Patent Document 2] Japanese Patent Publication No. 2023-73184 Summary of the Invention [Problem to be solved by the invention]
[0009] Currently, the work of humans (translators) is essential to producing high-quality dubbed videos that maintain the appeal of the original video and provide foreign language audio that matches the visuals. However, there are only a limited number of skilled translators, and due to delivery times and costs, it is not possible to meet the demand for dubbed videos. [Means for solving the problem]
[0010] The inventors have sought an effective means for providing high-quality dubbed videos in a short period of time to all businesses, including individuals and small businesses with limited financial resources, for original videos of various content, such as entertainment, advertising, education, and public services. As a result, they have discovered a video translator support system (10) that can provide translators with high-precision dubbed video data and reduce their workload, and a dubbed video editing system (100) that can use this system to provide the most natural and skillful dubbed lines. That is, the present invention is as follows.
[0011] (Invention 1) Using a database (1), original video data (2), a text generation unit (3), a translation unit (4), a dubbing unit (5), a communication unit (6), and a user terminal device (7), Store the original video data (2) in the database (1), A text generation unit (3) executes a voice recognition application to generate original language text data (31) from the voice data (21) included in the original video data (2) extracted from the database (1); A translation application is executed in a translation unit (4) to generate text data (41) in another language from the original language text data (31) generated in the text generation unit (3); A dubbing unit (5) executes a reading application to generate dubbed voice data (51) based on the multilingual text data (41) generated by the translation unit (4); A communication unit (6) transmits video editing support data (201) including original video data (2), original language text data (31), and other language text data (41) to a user terminal device (7); Video translator support system (10). (Invention 2) The video translator support system (10) of Invention 1 refers to the lip movements and / or lip shapes of characters appearing in the original video in at least one of the text generation unit (3), translation unit (4), and dubbing unit (5). (Invention 3) The system includes a database (1), original video data (2), a text generation unit (3), a translation unit (4), a dubbing unit (5), a communication unit (6), a user terminal device (7), and a video generation unit (9). The original video data (2) is stored in the database (1), A text generation unit (3) executes a voice recognition application to generate original language text data (31) from the voice data (21) included in the original video data (2) extracted from the database (1); A translation application is executed in a translation unit (4) to generate text data (41) in another language from the original language text data (31) generated in the text generation unit (3); A dubbing unit (5) executes a reading application to generate dubbed voice data (51) based on the multilingual text data (41) generated by the translation unit (4); A communication unit (6) transmits video editing support data (201) including original video data (2), original language text data (31), and other language text data (41) to a user terminal device (7); The user terminal device (7) transmits updated video data (202) including updated original language text data (32) obtained by processing the original language text data (31) and / or updated other language text data (42) obtained by processing the other language text data (41) to the communication unit (6); A video generation unit (8) generates dubbed video data (203) based on the updated video data (202); Dubbing video editing system (100). (Invention 4) The user terminal device (7) transmits the updated video data (202) consisting of the multilingual text data (42) to the communication unit (6); Next, the dubbing unit (5) executes a reading application to generate dubbed voice data (520) based on the multilingual text data (42) acquired by the communication unit (6); Furthermore, a video generation unit (8) generates dubbed video data (203) based on the multilingual text data (42) and the dubbed audio data (520). Invention 3: A dubbing video editing system (100). (Invention 5) The dubbing video editing system (100) of Invention 3 refers to the lip movements and / or lip shapes of characters appearing in the original video in at least one of the text generation unit (3), translation unit (4), and dubbing unit (5). [Effects of the Invention]
[0012] The video translator support system (10) can reduce the effort of translators who participate in dubbing video editing. Using the dubbing video editing system (100), high-quality dubbing videos can be provided at low cost. [Brief explanation of the drawings]
[0013] [Figure 1] A reference diagram for understanding an example of a video translator support system (10). [Figure 2] 1 is a reference diagram for understanding an example of a dubbing video editing system (100). [Figure 3] 1 is a reference diagram for understanding an example of a dubbing video editing system (100). DETAILED DESCRIPTION OF THE INVENTION
[0014] [Video Translator Support System (10)] The video translator support system (10) of the present invention is a method for providing mechanically translated text data and / or mechanically dubbed audio data to translators who participate in the dubbing of original videos containing commentary, speech, and dialogue in a certain language. Here, "translator" means the party in charge of translation, and includes not only the person who actually translates, but also translation agencies, assistants to translators, and people who manage them.
[0015] The system or device on which the video translator support system (10) is executed may be a system or device managed by the translator himself, a system or device managed by a person whose business is supporting translators, or a system or device managed by a person whose business is editing and providing dubbed videos using translators. In the present invention, the system or device on which the video translator support system (10) is executed is collectively referred to as the operator of the video translator support system (10).
[0016] The video translator support system (10) includes a database (1), original video data (2), a text generation unit (3), a translation unit (4), a dubbing unit (5), a communication unit (6), and a user terminal (7). The database (1), original video data (2), text generation unit (3), translation unit (4), dubbing unit (5), and communication unit (6) operate as applications executed by a computer. The above applications are capable of speech recognition, translation, speech synthesis, and image recognition.
[0017] [Database (1)] In the video translator support system (10), original video data (2) is stored in a database (1). The database (1) may be located on a server on the Internet or on a local device managed by the operator of the video translator support system (10). For example, when a Japanese store publishes a store introduction video with Japanese audio on its website and dubs it into a foreign language, the data of the store introduction video is uploaded to a cloud on the Internet as original video data (2) and stored in the database (1) in the cloud. In this case, information about the store that is requesting the service is registered in the database (1) in association with the uploaded original video data (2). The original video data (2) and the data associated with it can be updated at any time.
[0018] [Original video data (2)] The original video data (2) used in the video translator support system (10) is the data of the original video to be dubbed. This original video is any video that includes audio in at least one language. In addition to certain language audio, the original video data (2) can also include data such as images, videos, music (theme songs and background music), sound effects, and subtitles.
[0019] Original videos are, for example, promotional videos for products or services. Examples of original videos include promotional and informational videos provided by businesses such as restaurants, retailers selling daily necessities and food, manufacturers or retailers of traditional crafts, performers and promoters of performing arts such as dance and comedy, massage parlors, hairdressers, saunas, and spas, hospitals, nursing homes, IT service providers, cultural schools, cram schools, and other educational service providers, traditional performing arts practitioners such as hanamichi and sado, sports-related businesses such as sports instructors and sports classes, travel and accommodation providers, and public services such as trains and buses. Original videos are generally shorter in length (with shorter playback and viewing times) than films or broadcast programs. Furthermore, original videos requested by the above businesses often feature the businesses (store owners, producers, performers, practitioners, and those responsible for providing products and services) introducing and explaining their products and services in their own voices.
[0020] [Text Generation Section (3)] In the video translator support system (10), a text generation unit (3) executes a speech recognition application to generate original language text data (31) from the speech data (21) contained in the original video data (2) extracted from the database (1). The speech recognition application is also called a transcription application. The speech recognition application used is generally a transcription application with high recognition accuracy that uses speech recognition AI. The original language text data (31) is original text data for translation.
[0021] For example, the text generation unit (3) generates Japanese text data (31) from Japanese voice data (21) included in the original video data (2). When a conversation scene between three people (persons A, B, and C) appears in the original video, the text generation unit (3) generates speech text data (31) for each of the people A, B, and C.
[0022] The text generation unit (3) can use not only the voices contained in the original video data (2) but also the lip movements of the characters appearing in the original video. For example, in the video translator support system (10), for a time period when a specific lip movement and / or shape of a character is continued or maintained, text that closely matches the emotional expression of the character at that time is registered. The text generation unit (3) can perform voice recognition and image recognition in parallel on the original video data (2).
[0023] When the text generation unit (3) executes a separate image recognition application to detect the specific lip movement and / or lip shape from an image of the lips of a person appearing in a video, the text generation unit (3) can extract the most appropriate text from the text associated with the detected lip movement and / or shape.The text generation unit (3) can then correct the text transcribed by the speech recognition application using the extracted text.The image recognition application used in the video translator support system (10) can use not only the lips of a character, but also faces, hands, fingers, etc. as image targets for detecting movement and shape.
[0024] For example, if a character in the original video is silent with their lips held in the "o" position during a period of time, no speech is detected in the transcription, and no text is generated. However, during this period, the character in the original video may express strong surprise or exclamation. Therefore, text expressing surprise or exclamation, such as "Oh" or "Wow," is registered when the character's lips remain open in the "o" position. The text generation unit (3) then extracts text that matches the immediately preceding text from the text registered for that period of time. For example, if the immediately preceding text is a woman's speech, the unit selects text that is likely to be spoken by a woman. Alternatively, the text generation unit (3) edits the text registered for that period of time based on the immediately preceding text. For example, if a character has just asked a question or questioned something, the unit can convert it into natural-sounding text, such as "I see!" or "I understand!". The original language text data (31) generated by converting silent portions of the original video data (2) into text may serve as a translation text that more accurately conveys the intent of the video and characters.
[0025] [Translation Department (4)] In the video translator support system (10), a translation application is executed in a translation unit (4) to generate text data (41) in another language from text data (31) in the original language generated in a text generation unit (3).
[0026] Generally, AI functions are also incorporated into the translation application used by the translation unit (4). In this case, the translation application refers to data on the original video, its users (e.g., shops or restaurants), and characters (e.g., actors, store owners), and extracts and arranges translation expressions that match the original video from dictionary data. The dictionary data for translation may be registered in advance in a database (1), or may be dictionary data that can be accessed via the Internet.
[0027] The translation unit (4) can also use the lip movements and other body movements of the people appearing in the original video, just like the text generation unit (3). In this case, the translation unit (4) can run a translation application and an image recognition application in parallel.
[0028] For example, if a period of time is detected in which a character in the original video is silent with their lips held in the "O" position, no speech will be detected in the transcription, and no text will be generated by the text generation unit (3). In this case, it is not possible to generate a translated text corresponding to this period of time. However, during this period of time, the character in the original video may express strong surprise or exclamation. Therefore, translated text (for example, Oh!, Wow!, Oops!, Nooooo! in English) is registered when the character's lips remain open in the "O" position. In this case, translated text will be assigned to the silent periods in the original video.
[0029] In the case of translation, the translation unit (4), like the text generation unit (3), extracts text that matches the immediately preceding text from the text registered for the time period. For example, if the immediately preceding text is text spoken by a woman, the translation unit (4) selects text that is likely to be spoken by a woman. Alternatively, the text generation unit (3) edits the text registered for the time period based on the immediately preceding text. For example, if a character had previously asked a question or expressed a doubt, it can be converted into text that sounds natural in actual conversation, such as "Ahh!" or "Wow! Good job." In this way, the translation unit (4) can also generate translated text for silent (silent) parts of the text generation unit (3). The multilingual text data (41) generated in this way may become translated text that more accurately conveys the intention of the video or characters.
[0030] [Dubbing Department (5)] In the video translator support system (10), a reading application can be executed in the dubbing unit (5) to generate dubbed voice data (51) based on the other language text data (41) generated by the translation unit (4). The reading application used in the dubbing unit (5) is a so-called AI-based voice synthesis application. The voice of the speaker of the original video can be used as the sound source of the reading voice synthesized in the dubbing unit (5). The reading application uses the speaker's voice data extracted from the original video data (2) and the other language text data (41) to generate dubbed voice data (51) as voice data that sounds as if the speaker of the original video is speaking in another language.
[0031] The dubbing unit (5) can also use lip movements and other body movements of the characters appearing in the original video, just like the text generation unit (3) and translation unit (4). In this case, the dubbing unit (5) can run a voice synthesis application and an image recognition application in parallel.
[0032] For example, if the dubbing unit (5) executes an image recognition application and detects a time period in which the lips of a character in an original image are still, the dubbing unit (5) will determine this time period as a "pause" or "silence" and will not assign a reading voice.
[0033] For example, the dubbing unit (5) can also execute an image recognition application to detect and evaluate the degree of change in the image of the lips of a character in the original video. In this case, the degree of change in the image of the lips (for example, the change in the image area corresponding to the lips per unit time) indicates lip movement. The dubbing unit (5) can recognize that the greater the change, the faster the character's speaking rate (speaks faster). In this case, the dubbing unit (5) can cause the speech synthesis application to adjust the reading rate of the other language text data (41) to correspond to the speaking rate recognized by the image recognition application. As a result, it is possible to generate dubbed audio data (51) that sounds like the character in the original video is speaking in a more natural tone.
[0034] [Video editing support data (201), user devices (7)] In this way, the video translator support system (10) generates original language text data (31) and other language text data (41) in addition to the original video data (2) as data that video translators can refer to, and can also generate dubbed audio data (51). After this, in the video translator support system (10), the communication unit (6) transmits video editing support data (201) including the original video data (2), the original language text data (31), and the other language text data (41) to the user terminal (7).
[0035] The "user" in the video translator support system (10) is a person or company that receives assistance in creating dubbed videos and performs dubbing work. Such users are typically workers such as translators contracted by dubbing video editing companies.
[0036] For example, a user who is a dubbing technician can acquire the multilingual text data (41), the dubbed audio data (51), and the original video data (2) on a user terminal (7). The user can display and generate the multilingual text, the dubbed audio, and the original video using the equipment attached to the user terminal (7). The user can modify the multilingual text data (41) to better match the original image.
[0037] The original language text data (31), other language text data (41), and dubbed audio data (51) reproduce the content of the original video with high accuracy given the current state of technology. A user finds a portion where the original language text data (31), other language text data (41), or dubbed audio data (51) does not match the original video, and corrects the other language text data (41) corresponding to this portion. As a result, a single user can complete nearly perfect dubbed video data.
[0038] Typically, the communication unit (6) transmits video editing support data (201) consisting of multilingual text data (41) to the user terminal (7). Preferably, the communication unit (6) transmits the original video data (2) together with the video editing support data (201) to the user terminal (7). In this case, the user, who is a translator, receives the dialogue (speech / conversation) text of the original video that has been translated almost accurately. The user can review and revise the multilingual text data (41) so that it perfectly matches the expression of the original video.
[0039] The video translator support system (10) significantly reduces the labor required to dub original videos. Even translators who do not have the high level of skill required for dubbing movies and TV shows can now be involved in creating dubbed videos. As a result, videos can be dubbed into various languages at low cost and in a short time.
[0040] The "user" in the video translator support system (10) may be an individual who wishes to have an original video dubbed, such as a restaurant, retail store, or individual. In this case, even if the user does not have advanced skills, they can create a high-quality dubbed video by using the video translator support system (10).
[0041] [Dubbing Video Editing System (100)] When editing dubbed videos, there are cases where the original videos have a large number of points. This is the case when the original videos are structured as a series and many related videos are dubbed, such as "Episode 1," "Episode 2," "Episode 3," etc. There are also cases where the original videos are frequently updated, such as when an explanatory video is updated every time a software version is updated. Also, dubbing video editing companies may use many translators to perform the dubbing work. In these cases, an efficient dubbing video editing system (100) using a video translator support system (10) is effective.
[0042] The dubbing video editing system (100) has the following features, similar to the video translator support system (10). The dubbing video editing system (100) includes a database (1), original video data (2), a text generation unit (3), a translation unit (4), a dubbing unit (5), a communication unit (6), a user terminal (7), and a video generation unit (9). The original video data (2) is stored in the database (1). The text generation unit (3) executes a voice recognition application to generate original language text data (31) from audio data (21) included in the original video data (2) extracted from the database (1). The translation unit (4) executes a translation application to generate other language text data (41) from the original language text data (31) generated by the text generation unit (3). The dubbing unit (5) executes a read-aloud application to generate dubbing audio data (51) based on the other language text data (41) generated by the translation unit (4). A communication unit (6) transmits video editing support data (201) including original video data (2), original language text data (31), and other language text data (41) to a user terminal device (7). These features are the same as those described for the video translator support system (10).
[0043] In the dubbing video editing system (100), the updated video data (202) is further transmitted from the user terminal (7) to the communication unit (6). Then, the dubbing video editing system (100) generates dubbing video data (203) based on the updated video data (202).
[0044] [Updated video data (202)] In the dubbing video editing system (100), a user terminal (7) transmits updated video data (202) to a communication unit (6), the updated video data (202) including updated original language text data (32) obtained by processing the original language text data (31) and / or updated other language text data (42) obtained by processing the other language text data (41).
[0045] If the dubbed voice data (51) is included in the video editing support data (201) transmitted from the communication unit (6) to the user terminal device (7), the updated video data (202) transmitted from the user terminal (7) to the communication unit (6) may include updated dubbed voice data (52) obtained by processing the dubbed voice data (51). The original language text data (32), the other language text data (42), and the dubbed voice data (52) are, in short, data for generating more natural and skillful speech.
[0046] When generating the original language text data (32) from the original language text data (31), the data may be corrected automatically or mechanically by executing an application through data input and transmission from the user terminal (7), or the user may input the original language text data (32) into the user terminal (7) and correct the data. The same applies when generating the other language text data (42) and dubbed audio data (52).
[0047] Typically, the revised other-language text data (42) is transmitted from the user terminal (7) to the communication unit (6) as the updated video data (202). In this case, the other-language text data (41), which is a nearly accurate transcription of the dialogue from the original video, is revised to the other-language text data (42) that is translated with even higher accuracy. When a translator independently revise the other-language text data (41) to the other-language text data (42), they may be able to capture the intentions of the characters or the client (user of the original video, person requesting dubbing) that cannot be recognized by machine or automatic translation. In this case, the other-language text data (42) can be said to correspond to the most natural and skillfully translated text.
[0048] [Video Generation Section (8)] In the dubbing video editing system (100), a video generation unit (8) generates dubbing video data (203) based on updated video data (202). Typically, the video generation unit (8) generates the dubbing video data (203) based on other language text data (42).
[0049] Today's "nearly accurate" dubbed dialogue generated mechanically or automatically contains a small number of lines that do not fit the original video (poorly translated text). These small lines sound unnatural to the video viewer, which causes the quality of the dubbed video to decrease. The dubbed video editing system (100) can generate dubbed video data (203) that can generate dubbed videos that sound natural to the viewer by using perfectly translated foreign language text data (42).
[0050] In this case, the user terminal (7) transmits updated video data (202) consisting of multilingual text data (42) to the communication unit (6), and then the dubbing unit (5) executes a read-aloud application to generate dubbed audio data (520) based on the multilingual text data (42) acquired by the communication unit (6). As described above, when the communication unit (6) acquires multilingual text data (42) that can generate the most natural and skillful translation text, it can generate dubbed audio data (520) that can generate the most natural and skillful dubbed lines.
[0051] When generating the dubbed audio data (520), the reading application can refer to the lip movements and / or lip shapes of the characters appearing in the original video. The referencing method is as explained in the video translator support system (10). In this case, a higher quality dubbed audio that more closely matches the movements of the characters in the original video can be generated from the dubbed audio data (520). In this way, dubbed video data (203) that corresponds to the original video with high accuracy is generated.
[0052] The video generation unit (8) can also generate dubbed video data (203) based on the multilingual text data (42) and the dubbed audio data (520). In this case, the dubbed video includes the most natural and effective subtitles based on the multilingual text data (42) and the most natural and effective dialogue based on the dubbed audio data (520).
[0053] The text generation unit (3), translation unit (4), and dubbing unit (5) of the video translator support system (10) and dubbing video editing system (100) can use AI video editing applications provided by, for example, HeyGen.
[0054] The dubbed video data (203) is stored in the database (1) and can be transmitted from the communication unit (6) to the user terminal (7) or other user terminals (70). Here, the other user terminals may be the terminals of the person requesting the dubbed video editing. The dubbed video data (203) can also be uploaded to a database located on a predetermined server on the Internet. The generated dubbed video data (203) is provided to the owner / manager of the original video and can be distributed to viewers along with the original video.
[0055] When providing the dubbing video data (203) generated by the dubbing video editing system (100) to a client, the billing unit (9) can calculate a dubbing video editing fee based on the video display time (T) derived from the original video data (2). The calculated video editing fee is sent to the client requesting the video editing together with the dubbing video data (203). The dubbing video editing system (100) can further include a section for managing customer information and payments. [Example]
[0056] [Example 1] Example 1 is an example of a video translator support system (10). Example 1 is a system that supports translators when dubbing original videos containing Japanese audio into English. Figure 1 is a reference diagram for understanding Example 1. Figure 1 does not depict the actual amount of data or device layout, but omits details and simply shows the relationship between the data handled in the present invention.
[0057] In Example 1, an application running on the Internet is executed on Japanese audio data (21) included in original video data (2), generating Japanese text data (31), English text data (41), and dubbed English audio data (51). Video editing support data (201) is transmitted from a communication unit (6) to a translator's terminal (7) connected to the Internet. The video editing support data (201) includes the original video data (2), Japanese text data (31), and English text data (41).
[0058] The terminal (7) and its associated devices display and play the original video data (2), Japanese text data (31), and English text data (41). The translator can compare the original video played from the original video data (2) with the English text data (41) based on a nearly accurate translation, and correct the English text data (41) as necessary.
[0059] [Example 2] Example 2 is an example in which a video editing company uses a dubbing video editing system (100) to provide a customer with a dubbed video (2030) of an original video (2000). The original video (2000) is published on the homepage of a Japanese restaurant, and in the original video, a female host explains the enlarged dishes.
[0060] 2 and 3 are reference diagrams for understanding Example 2. Figures 2 and 3 do not depict the actual data volume, device layout, or images, but omit details and simply show the relationship between the data handled in the present invention.
[0061] In Example 2, first, the customer's terminal (70) connected to the Internet communicated with the video editing company's server on the Internet and uploaded original video data (2). The original video data (2) was stored in a database (1) managed by the video editing company.
[0062] Next, the dubbing video editing system (100) is operated, and Japanese text data (31), Arabic text data (41), and dubbing Arabic audio data (51) are generated from the Japanese audio data (21) contained in the original video data (2). The Japanese text data (31), the generated Arabic text data (41), and the original video data (2) are sent as video editing support data (201) from the communication unit (6) to the translator's terminal (7) connected to the Internet.
[0063] The translator can download the original video data (2) and play the original video on his / her terminal (7) connected to the Internet. The translator can also display the Japanese text data (31) and the Arabic text based on the Arabic text data (41) on the terminal (7). The translator can refine the Arabic text by referring to the original video. When the translator corrects the Arabic text, the translator transmits the corrected Arabic text data (42) as updated video data (202) from the terminal (7) to the communication unit (6).
[0064] Thereafter, a dubbing unit (5) of the dubbing video editing system (100) generates new dubbed audio data (520) from the more accurate Arabic text data (42). The dubbing unit (5) references the lip movements of the characters in the original video, selects a speech rate and speech rhythm that match the lip movements, and generates the dubbed audio data (520). In this way, the approximately accurate Arabic text data (41) is improved into Arabic text data (42), and the approximately accurate dubbed audio data (51) is improved into dubbed audio data (520) that produces more natural and fluent Arabic speech.
[0065] Furthermore, the video generation unit (8) of the dubbing video editing system (100) generates dubbing video data (203) from data other than Japanese audio contained in the original video data (2), Arabic text data (42), and dubbing audio data (520). The dubbing video data (203) is stored in the database (1) and provided to the customer.
[0066] In addition, the customer terminal (70) is sent a dubbing editing fee calculated based on the playback time of the dubbed video (2030) (the length of the dubbed portion of the original video (2000)).
[0067] When the dubbed video data (203) is uploaded to the customer's store homepage, the dubbed video (2030) can be viewed by an unspecified number of viewers. Figure 3 shows the original video (2000) and the dubbed video (2030) played on a smartphone. If the viewer selects it, subtitles (420) based on Arabic text data (42) are displayed on the dubbed video (2030). [Industrial Applicability]
[0068] The video translator support system (10) and dubbing video editing system (100) of the present invention can significantly improve the efficiency of translation work required for dubbing videos. The dubbing video editing system (100) of the present invention can edit dubbing videos that are more natural and accurately reproduce the intention of the original video by acquiring updated video data (202).
[0069] The present invention can provide high-quality dubbed videos in a short time. The present invention can improve the advertising function of videos distributed over the Internet. The video translator support system (10) and dubbed video editing system (100) of the present invention are particularly useful for providing high-quality dubbed videos at low prices to relatively small stores and restaurants. The present invention is expected to be used in information dissemination in various fields such as the translation industry, retail stores, restaurants, entertainment, and public services. [Explanation of symbols]
[0070] 1 Database 2 Original video data 21 Audio data 3 Text Generation 31 Original language (Japanese) text data 4 Translation Department 41 Text data in other languages (English or Arabic) 42 Updated multilingual (English or Arabic) text data 420 subtitles 5. Dubbing Department 51 dubbed (English or Arabic) audio data 52 updated dubbed (English or Arabic) audio data 6. Communications Department 7 User terminals 70 Customer terminals 8 Video Generation Unit 201 Video editing support data 202 Updated video data 203 dubbed video data 2000 original videos 2030 dubbed video 10 Video translator support system 100 Dubbing Video Editing System
Claims
1. Using a database (1), original video data (2), a text generation unit (3), a translation unit (4), a dubbing unit (5), a communication unit (6), and a user terminal device (7), Store original video data (2) in a database (1), A text generation unit (3) executes a voice recognition application to generate original language text data (31) from the voice data (21) included in the original video data (2) extracted from the database (1); A translation application is executed in a translation unit (4) to generate text data (41) in another language from the original language text data (31) generated in a text generation unit (3); A dubbing unit (5) can execute a reading application to generate dubbed voice data (51) based on the other language text data (41) generated by the translation unit (4); A communication unit (6) transmits video editing support data (201) including original video data (2), original language text data (31), and other language text data (41) to a user terminal device (7); Video translator support system (10).
2. The video translator support system (10) according to claim 1, wherein at least one of the text generation unit (3), the translation unit (4), and the dubbing unit (5) refers to the lip movements and / or lip shapes of characters appearing in the original video.
3. The system includes a database (1), original video data (2), a text generation unit (3), a translation unit (4), a dubbing unit (5), a communication unit (6), a user terminal device (7), and a video generation unit (9), Original video data (2) is stored in a database (1), A text generation unit (3) executes a voice recognition application to generate original language text data (31) from the voice data (21) included in the original video data (2) extracted from the database (1); A translation application is executed in a translation unit (4) to generate text data (41) in another language from the original language text data (31) generated in a text generation unit (3); A dubbing unit (5) can execute a reading application to generate dubbed voice data (51) based on the other language text data (41) generated by the translation unit (4); A communication unit (6) transmits video editing support data (201) including original video data (2), original language text data (31), and other language text data (41) to a user terminal device (7); A user terminal device (7) transmits updated video data (202) including updated original language text data (32) obtained by processing the original language text data (31) and / or updated other language text data (42) obtained by processing the other language text data (41) to a communication unit (6); A video generation unit (8) generates dubbed video data (203) based on the updated video data (202); A dubbing video editing system (100).
4. The user terminal device (7) transmits the updated video data (202) consisting of the multilingual text data (42) to the communication unit (6); Next, the dubbing unit (5) executes a reading application to generate dubbed voice data (520) based on the other language text data (42) acquired by the communication unit (6); Furthermore, a video generation unit (8) generates dubbed video data (203) based on the other language text data (42) and the dubbed audio data (520). The dubbing video editing system (100) of claim 3.
5. The dubbing video editing system (100) according to claim 3, wherein at least one of the text generation unit (3), the translation unit (4), and the dubbing unit (5) refers to the lip movements and / or lip shapes of characters appearing in the original video.
Citation Information
Patent Citations
Web animation culture school system
JP2009302930A
Video editing system
JP2023073184A