Translated content generation system

The translation content generation system addresses the challenge of translating radio broadcast content by using voice data to recreate the original's tone and appeal, resulting in high-quality translated content that expands accessibility and advertising potential.

WO2025120784A1PCT designated stage expired Publication Date: 2025-06-12NIPPON BROADCASTING SYSTEM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/043728
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current technologies struggle to effectively translate radio broadcast content while preserving the unique appeal and individuality of the original performers, limiting the content's accessibility and advertising potential beyond its original language audience.

Method used

A translation content generation system that utilizes voice data of performers to create translated content through speech-to-text conversion, translation, and voice synthesis, ensuring that the translated content retains the original's tone, intonation, and appeal.

Benefits of technology

The system generates high-quality translated content that mimics the original's performance, effectively expanding the content's reach and appeal to non-native speakers, while also enhancing its utility as an advertising medium.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023043728_12062025_PF_FP_ABST
    Figure JP2023043728_12062025_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To improve the quality of translated content and utilize an advertising function of the translated content. [Solution] A translated content generation system (100) includes an input unit (1), a transcription unit (2), a translation unit (3), an audio synthesis unit (4), and an output unit (5). The input unit (1) acquires speech audio (11) of a performer. The transcription unit (2) generates speech text (21) of the performer. The translation unit (3) generates speech text (31) in a language (30) different from the original. The audio synthesis unit (4) synthesizes audio (40) in which the speech text (31) is read in the voice of the performer. The output unit (5) outputs translated content (50) that includes the audio (40). [Selected drawing] Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Translation content generation system

[0001] The present invention relates to a translation content generation system, for example, a system for creating translation content that accurately reproduces the characteristics of original radio broadcast content using audio data of performers in a radio program.

[0002] Radio broadcasting has the longest history of any medium that uses sound to reach an unspecified number of users. The world's first radio broadcast is said to have been the transmission of a Christmas message in the United States in 1906.

[0003] Most radio broadcast content consists mainly of narration by the performer (announcer, presenter, radio personality), music, and audio advertisements. Performers often select the music to broadcast based on the reactions (requests) of viewers (listeners). Therefore, radio broadcast content allows users to feel a sense of familiarity compared to written media such as newspapers and books.

[0004] Radio broadcasts do not have the ability to show viewers real objects or visual data as intended by the creator, as videos, images, or movies do, but they do have the power to stimulate listeners' imaginations. The 1997 film "Radio Time" attracted a great deal of attention for its depiction of how listeners' imaginations are stimulated by short narrations on the radio.

[0005] Currently, internet distribution of video and music content is widespread, as is radio broadcasting. Patent Documents 1, 2, and 3 describe technologies related to the internet distribution of radio broadcast content. However, when comparing users of internet-distributed content, there are significant differences between radio broadcasting and users of television, movies, social media, etc. Radio broadcasting differs from other media in that many users (listeners) use it while doing other things. Many listeners enjoy listening to the host's narration and music while driving, working, studying, or watching sports (listening to radio commentary at the stadium). Recently, internet-distributed radio broadcasts are sometimes used as background music in restaurants. For example, an Italian restaurant might play an Italian radio program to make customers feel as if they were in Italy while dining. On the other hand, while watching content such as movies and videos, users of television, movies, and social media tend to focus their attention on the display of the terminal device on which the content is being distributed.

[0006] Reflecting this difference, radio listeners, unlike television, movie, and social media users, rarely skip over advertising content that is delivered during broadcast programs. This is quite different from the situation of many video and visual content viewers who pay an additional fee to use a feature to skip advertising content. Radio broadcasts can be said to be a method of delivering advertisements that appeal to users, even in the internet age.

[0007] However, because the appeal of radio broadcast content lies in the individuality of the performers themselves, and because radio broadcast content consists of audio and music, it is difficult to modify the original content. For example, as described in Patent Document 4, translated content with subtitles added to movie or video content can be provided to a wide variety of users, but in the case of radio broadcast content, subtitle distribution is of little use to many listeners. Also, for example, the traditionally used "dubbed version" does not reproduce the individuality of the radio broadcast performers and does not appeal to listeners.

[0008] For this reason, translated content has not been a consideration for the Internet distribution of radio broadcast content until now. For example, as described in Patent Documents 5, 6, and 7, methods have been proposed for distributing additional information related to the original content to listeners' terminals, but in none of these cases has the original content been modified.

[0009] JP-T-2017-537487 A JP-A-2018-195939 A JP-A-2022-78725 A JP-A-2019-47391 A JP-A-2008-241943 A JP-A-2017-112411 A JP-A-2021-89675 A

[0010] In this situation, users of radio broadcast content are limited to those who speak the language of the original content. Although radio broadcasts are useful as an advertising medium, the distribution of content is limited. The advertising function of radio broadcast content is not being fully utilized. Furthermore, although radio broadcast content possesses the excellent creativity and copyrightability unique to audio content, only a limited number of users can access the content. Unlike manga or animation, it is virtually impossible to convey the appeal of the original content in other languages ​​through radio broadcasts.

[0011] For radio broadcasters, translated content is a desirable way to reach potential users, but this has not yet been realized. If translated content could be created that retains the appeal of the original radio broadcast content, it would increase the number of users and their options, and create new content that is effective for advertising, entertainment, education, customer service, and more.

[0012] The inventors sought translated content that could be used for radio broadcasting, reproducing the appeal of the original content in another language. As a result, the inventors discovered that translated text-to-speech data using the voice data of the performers in the original content was effective. That is, the present invention is as follows.

[0013] (Invention 1) A translation content generation system (100) including an input unit (1), a transcription unit (2), a translation unit (3), a speech synthesis unit (4), and an output unit (5), wherein the input unit (1) acquires speech (11) of performers from original content (10), the transcription unit (2) transcribes the speech (11) to generate speech text (21) of the performers in an original language (20), the translation unit (3) translates the speech text (21) into a language (30) other than the original language (20) to generate speech text (31) in the language (30), the speech synthesis unit (4) synthesizes speech (40) in which the speech text (31) is read out in the voice of the performer based on the speech (11) of the performers and the speech text (31), and the output unit (5) outputs translation content (50) including the speech (40). (Invention 2) The translation content generation system (100) of Invention 1, wherein the original content (10) is radio broadcast content. (Invention 3) The translation content generation system (100) of Invention 1, further including a subtitle generation unit (6) that generates subtitles (60) based on spoken text (31), and wherein the output unit (5) outputs translation content (50) including audio (40) and / or subtitles (60). (Invention 4) The translation content generation system (100) of Invention 1, wherein the output unit (5) outputs translation content (50) including audio (40) and further advertising content (51).(Invention 5) An input unit (1) acquires speech sounds (11-1, 11-2, ... 11-n) of n (n is an integer equal to or greater than 2) performers from an original content (10), a transcription unit (2) transcribes the speech sounds (11-1, 11-2, ... 11-n) to generate speech texts (21-1, 21-2, ... 21-n) of the n performers in an original language (20), and a translation unit (3) translates the speech texts (21-1, 21-2, ... 21-n) into a language (30) other than the original language (20) to generate speech texts (31-1, 31-2, ... 31-n) in the language (30), A translation content generation system (100) of invention 1, wherein a speech synthesis unit (4) synthesizes speech (40-1, 40-2, ... 40-n) reading out the spoken text (31-1, 31-2, ... 31-n) in the voice of a performer based on the speech (11-1, 11-2, ... 11-n) of the performer and the spoken text (31-1, 31-2, ... 31-n), and an output unit (5) outputs translation content (50) including the speech (40-1, 40-2, ... 40-n). (Invention 6) A device (101) storing data of the translation content (50) according to any one of claims 1 to 4. (Invention 7) A translation content distribution method (102) transmitting the translation content (50) according to any one of claims 1 to 4 to a terminal device via the Internet.

[0014] The translated content (50) generated by the translation content creation system (100) of the present invention (hereinafter referred to as the present system (100)) can provide users with utterances and conversations with the same meaning as those in the original content, using synthesized voice that is indistinguishable from the voices of the performers themselves. The present system (100) can reproduce the characteristics of the original content, such as the performers' unique speaking style and vocal intonation, in the translated content (50).

[0015] The figure shows how the system (100) generates translated radio broadcast content (50) from original radio broadcast content (10). The figure shows an example of how the transcription unit (2) of the system (100) generates spoken text (21). The figure shows an example of how the translation unit (3) of the system (100) generates spoken text (31).

[0016] [Translated Content Generation System (100)] The translated content generation system (100) includes an input unit (1), a transcription unit (2), a translation unit (3), a speech synthesis unit (2), and an output unit (5). The original content (10) handled by this system (100) is not limited as long as it includes the spoken voice of the performers. The original content (10) is, for example, video content with audio such as dialogue and commentary from movies and videos, photo or image content with audio, and content including the spoken voice of specific performers such as conversations, narration, lines, readings, dialogues, and lectures.

[0017] A typical original content (10) is radio broadcast content. Radio broadcast content is content distributed from a radio station to dedicated listening devices (so-called radios) in real time or as recorded rebroadcasts. Radio broadcast content is also content distributed via the Internet to terminal devices such as PCs, tablets, and smartphones, and is content distributed in the form of what are now called "Internet radio" or "podcasts."

[0018] Radio broadcast content includes audio of the performers' conversations and speeches, and may also include audio provided from sources outside the broadcast studio, such as music, musical performances, advertisements, and sound effects. Typical radio broadcast content consists of the narration of the performer who acts as the host and music. The performer who acts as the host is also called an emcee, radio personality, navigator, presenter, host, guest, MC, etc. The performer who acts as the host may also be accompanied by a partner called a guest or partner. In this system (100), the host and partner are collectively referred to as "performers."

[0019] [Input Unit (1)] In this system (100), the input unit (1) acquires the speech sounds (11) of the performers from the original content (10). In practice, the input unit (1) extracts data corresponding to the speech sounds (11) of the performers from the sound source data of the original content (10).

[0020] The method for extracting the speech sound (11) is not limited. For example, known so-called conversation sound source extraction techniques can be used without limitation. In the conversation sound source extraction technique, speech sound is extracted based on, for example, the characteristics of the performer's voice (e.g., identifying male or female voices), the type of language (e.g., extracting sounds that are presumed to be Japanese), and whether or not it has meaning (e.g., excluding inanimate sounds such as explosions, animal cries that cannot be presumed to be words, performer coughs, etc.). Alternatively, the speech sound (11) can be extracted by specifying the time period in which the performer spoke from the sound source of the original content (10) input in chronological order.

[0021] In this case, the start and end times of the speech of the lead actor can also be specified. In radio broadcasts, the time period during which the performers speak is often clearly separated from other time periods such as music and advertising. Therefore, when the original content (10) is radio broadcast content, the input unit (1) can acquire the speech voice (11) by extracting the voice from the end of a jingle (a short piece of music or sound effect inserted at a turning point in a program) to the start of the next jingle, for example.

[0022] [Transcription Unit (2)] In this system (100), the transcription unit (2) transcribes the spoken voice (11) to generate a speech text (21) of the performer in the original language (20).

[0023] A known speech recognition technology can be used to generate the spoken text (21). Recently developed and utilized speech recognition technology is called end-to-end speech recognition, which utilizes deep learning in artificial intelligence (AI) to achieve significantly improved recognition accuracy compared to conventional GMM-HMM speech recognition. The transcription unit (2) preferably generates the spoken text (21) using end-to-end speech recognition. In typical end-to-end speech recognition, the three steps used in DNN-HMM speech recognition—acoustic model step (word segmentation), pronunciation dictionary reference step (word guessing), and language model step (word alignment)—are performed using a series of neural network inferences. The accuracy of neural network inference (deep neural network method, DNN) can be improved by repeatedly performing AI learning (deep learning), which involves loading training data into a processor with AI functionality. As a result, end-to-end speech recognition performs processing similar to the inferences and judgments made by the human brain, making it possible to predict text directly from speech data with high accuracy.

[0024] In this system (100), it is necessary to reflect the utterances and conversational atmosphere of the original content (10) in the speech text (21). Therefore, the speech text (21) may include words that are not converted into text when transcribing minutes or conversations. For example, to reproduce the natural conversation of the participants, utterances such as "um" or "hmm" uttered when hesitant or confused, or "oh" or "wow" uttered when impressed or surprised, may be faithfully transcribed. If the original content (10) is radio broadcast content, the target of the audio to be transcribed (the so-called listening range) must be adjusted depending on the content category (e.g., news broadcast, music program, advice program) and the participants' catchphrases and tone of voice. By using AI learning in the transcription unit (2), it is possible to generate speech text (21) in which the original speech is transcribed more naturally.

[0025] [Translation Unit (3)] In the present system (100), the translation unit (3) translates the spoken text (21) into a language (30) other than the original language (20) to generate a spoken text (31) in the language (30).

[0026] There are no restrictions on the combination of the original language (20) and the language (30). A typical original language (20) is Japanese (20). In a typical system (100), one or more languages ​​(30) that are optimal for the original language (20) are selected, taking into consideration the translation accuracy of the translation unit (3), the characteristics of the original content (10), the distribution of potential users, and advertising content to be distributed together with the translated content.

[0027] The translation unit (3) machine-translates the spoken text (21) in the original language (20) into the spoken text (31) in the language (30). The machine translation method is not limited, but neural machine translation, which has high translation accuracy, is preferable. Neural machine translation also generates highly accurate and natural-looking translations using a deep learning neural network inference (DNN) that uses a large set of original texts and translated texts. The translation unit (3) can also generate spoken text (31) that expresses the original text more naturally by using AI learning.

[0028] The transcription unit (2) and / or translation unit (3) can use document processing technology, which utilizes so-called document processing AI, which has document proofreading, revision, and editing functions, in combination with the speech recognition and translation to correct or edit the actors' speech and lines converted into text. For example, in animated content, phrases can be corrected to match the attributes of the characters (age, gender, role). In news content, proper nouns and abbreviations can be updated to the latest expressions, and so-called fact-checking can be performed using AI technology. These corrections and editing are performed within the scope of not compromising the characteristics of the original content (10) and ensuring natural and fluent translation. Examples of document processing AI technology that can be used include ChatGPT (registered trademark).

[0029] [Speech synthesis unit (4)] In this system (100), the speech synthesis unit (4) synthesizes, as a speech waveform, a speech (40) in which the performer reads out the speech text (31) in his / her voice based on the performer's speech (11) and the speech text (31). The speech (11) is converted to sound as if the performer is speaking in the performer's own language (30), and the speech (40) is synthesized.

[0030] There are no limitations on the speech synthesis method used in the speech synthesis unit (4). The speech synthesis unit (4) preferably uses statistical model-based speech synthesis (DNN speech synthesis) using a DNN. This method is also called voice cloning because it synthesizes speech that is indistinguishable from the input speaker's speech. In DNN speech synthesis, AI learns a speech sample of the person (actual speaker) to acquire acoustic features of the voice. During the DNN process, the speech that would be heard if a text other than the sample was spoken using a speech that exhibits the acquired acoustic features is predicted, and the prediction model is updated to synthesize a more natural speech for the sample speech. By learning the spoken speech (11), the speech synthesis unit (4) predicts the speech waveform that would be heard if a performer spoke the spoken text (31), and generates a speech (40) that sounds as if it were spoken by the performer. The DNN allows the pauses, voice volume, emotions, etc. in the spoken speech (11) to be reproduced in the synthesized speech (40).

[0031] The speech (11) used as the sample speech in the speech synthesis unit (4) does not have to be the same speech as the speech acquired by the input unit (1). Current AI speech synthesis applications can accurately match other text-to-speech speech to the speech of the original speaker, even if the input sample speech is short or a short sentence. Therefore, the speech synthesis unit (4) may use a sample of the voice of a performer in the original content (10) as the speech (11), or may use a portion of the speech (11) acquired by the input unit (1).

[0032] The speech synthesis unit (4) can synchronize elements of the original content (10) with the speech (40). For example, the timing of the speech (40) can be adjusted to match the movements of the actors in the original content (10). As a result, just like in the original content (10), the translated content (50) can also include speech (40) (generally positive lines) corresponding to the actors' movements and emotional expressions, such as laughter, tears, head nodding, and so on.

[0033] Furthermore, for example, if the duration of the voice (40) generated from radio broadcast content in which a song or advertisement begins immediately after a performer finishes speaking (the time it takes for the voice (40) to read the spoken text (31)) is longer than the corresponding speaking time of the original content (10), the duration of the voice (40) can be shortened (the speaking speed can be increased) to synchronize the speech of the voice (40) with the original content (10) and the spoken voice (11). As a result, the output unit (5) described below can insert the song or advertisement in the same configuration as the original radio broadcast content (10).

[0034] [Output Unit (5)] In the present system (100), the output unit (5) outputs translated content (50) including audio (40). The translated content (50) is content in which the original content (10) has been "dubbed" into a language (30) using synthetic audio (40) that is indistinguishable from the voices of the performers. If the original content (10) is video content with audio, such as dialogue or commentary from a movie or video, photo or image content with audio, or content including speech by specific performers in a talk or lecture, the video, photo, and image in the translated content (50) may be identical to the original content (10).

[0035] The system (100) can generate multiple translations (50) for a single original content (10) using multiple languages ​​(30), and can create dubbed versions with more natural and fluent speech.

[0036] However, since users (viewers) often know the races and backgrounds of the actors in the original content (10), there is a possibility that users may feel a certain sense of incongruity, or an impression unique to a "dubbed version," when viewing the translated content (50) in which audio (40) is overlaid on the video or image of the actors included in the original content (10), although not as strong as in a traditional dubbed version. For example, a user who watches the translated content (50) (movie) in which the lines of a star actor born and raised in America have been translated into fluent Japanese may feel uncomfortable hearing the Japanese lines in the actor's voice, knowing that the actor does not actually speak Japanese.

[0037] If the original content (10) is an animated work, there is no need to worry about such incongruity. By using a voice (40) in the translated content (50) that is indistinguishable from the voice of the voice actor himself, the style of the original content (10) can be reproduced in the translated content (50).

[0038] Even when the original content (10) is radio broadcast content, this sense of incongruity specific to dubbed versions is reduced. Neither the original radio broadcast content (10) nor the translated radio broadcast content (50) depicts the radio personality's appearance. Listeners are interested in the fluent narration in the radio personality's voice (or a voice that is almost identical to the original), and rarely notice that the radio personality does not speak a language (30) other than the original language (20). Rather, listeners may find the translated radio broadcast content (50) amusing, imagining what it would be like if the radio personality spoke in a foreign language (30). Therefore, when the original content (10) is radio broadcast content, the translated content (50) may have an appeal that cannot be expected from the original.

[0039] [Subtitle Generation Unit (6)] The system (100) may further include a subtitle generation unit (6). The subtitle generation unit (6) generates subtitles (60) based on the spoken text (31). The subtitles (60) are useful when the original content (10) is video content with audio, such as dialogue or commentary, from a movie or video, or photo or image content with audio.

[0040] When the system (100) includes a subtitle generation unit (6), the output unit (5) can output the translated content (50) including audio (40) and / or subtitles (60). In this case, the user of the translated content (50) can select whether to output and display the audio (40) and / or subtitles (60) on the terminal device.

[0041] When the original content (10) is radio broadcast content, the translated radio broadcast content (50) can be treated as language learning content by distributing and displaying subtitles (60) on the listener's terminal device separately from the translated radio broadcast content (50).

[0042] [Advertising Content (51)] In the present system (100), the output unit (5) can output advertising content (51) in addition to the translated content (50) including the audio (40). The output unit (5) can select and combine advertising content (51) suitable for the translated content (50) from a database (52) attached to the present system (100). The present system (100) can deliver advertising content (51) that could not be applied to the original content (10).

[0043] [Multiple Performers] The system (100) can also be applied to cases where multiple performers speak in the original content (10). In this case, voices (40) are generated in patterns equal to the number of performers.

[0044] That is, in this case, an input unit (1) acquires speech sounds (11-1, 11-2, ... 11-n) of each of n (n is an integer equal to or greater than 2) performers from an original content (10). A transcription unit (2) transcribes the speech sounds (11-1, 11-2, ... 11-n) to generate speech texts (21-1, 21-2, ... 21-n) of each of the n performers in an original language (20). A translation unit (3) translates the speech texts (21-1, 21-2, ... 21-n) into a language (30) other than the original language (20) to generate speech texts (31-1, 31-2, ... 31-n) in the language (30). A speech synthesis unit (4) synthesizes speech (40-1, 40-2, ... 40-n) in which the speech text (31-1, 31-2, ... 31-n) is read out in the voice of the performer based on the speech (11-1, 11-2, ... 11-n) of the performer and the speech text (31-1, 31-2, ... 31-n). An output unit (5) outputs translation content (50) including the speech (40-1, 40-2, ... 40-n).

[0045] For example, for a radio broadcast content (n=2) in which the original content (10) is a radio personality, performer 1, talking in Japanese with guest performer 2, the radio personality's voice reading in English (40-1) and the guest's voice reading in English (40-2) are ultimately synthesized to output a translated radio broadcast content (50) containing the voices of the radio personality and guest talking in English.

[0046] [Device (101)] The device (101) of the present invention stores data of the translated content (50) obtained by the present system (100). The device (101) is, for example, a storage medium that stores data of the translated content (50), a server that distributes the translated content (50), etc. The device (101) may be included in the present system (100).

[0047] [Translated Content Distribution Method (102)] The translated content distribution method (102) of the present invention is a method for transmitting translated content (50) to a terminal device via the Internet. The terminal device is not limited as long as it allows the user to use the translated content (50). The terminal device may be, for example, a PC, a tablet, or a smartphone. The transmission method is not limited. For example, the system (100) can transmit data of the translated content (50) to a terminal device of a user who accesses a distribution site operated and managed by a radio station. Also, for example, a terminal device equipped with a translated content distribution application can download data of the translated content (50) from a distribution server.

[0048] The system (100) may include a communication unit (7) for distributing the translated content (50). In this case, the system (100) may output, store, and distribute the translated content (50) to users.

[0049] [Effects] The system (100) can provide users with high-quality translated content (50). The translated content (50) generated by the system (100) has the realism of a translated version dubbed into another language by the performers themselves. When the original content (10) is radio broadcast content, the translated content (50) is less likely to cause listeners a sense of discomfort, and can instead stimulate the listener's imagination and become attractive content. The translated content (50) is useful as a medium for advertising content (51). Because radio broadcast listeners are unlikely to skip over advertising content (51), translated radio broadcast content (50) generated from the original radio broadcast content (10) is particularly useful as an advertising medium.

[0050] Example 1 FIG. 1 shows how the system (100) generates translated radio broadcast content (50) from original radio broadcast content (10).

[0051] The original radio broadcast content (10) includes a radio personality's narration (11) and a jingle (12) as audio.

[0052] An input unit (1) of this system (100) acquires data of a radio personality's narration voice (11) from original radio broadcast content (10). A transcription unit (2) transcribes the narration voice (11) to generate a radio personality's narration text (21) in Japanese (20). A translation unit (3) translates the narration text (21) to generate a narration text (31) in English (30).

[0053] The speech synthesis unit (4) executes a speech synthesis application that performs DNN inference trained on the radio personality's speech voice (11). The speech synthesis unit (4) synthesizes an English speech voice (40) that reads the speech text (31) in the voice of the radio personality based on the radio personality's speech voice (11) and the spoken text (31).

[0054] An output unit (5) outputs an English radio broadcast content (50) including an English speaking voice (40) and a jingle (12), and an advertisement content (51) including an English voice extracted from an advertisement database (52).

[0055] A content set including English radio broadcast content (50) and an advertisement database (52) is transmitted from a communication unit (7) to a listener's terminal device (not shown).

[0056] However, the handling of the jingle (12) is not limited by this system (1). Common content for each program, such as an opening credit, a title credit, a theme song, an ending song, and a jingle, included in the original content (10) may be placed in the translated content (50) as is, or some or more of these may be changed, edited, or omitted before being placed in the translated content (50). Example 1 is an example in which a common jingle (12) is used between the original radio broadcast content (10) and the English version radio broadcast content (50).

[0057] [Example 2] Figure 2 shows how the transcription unit (2) proofreads the spoken text (21). In the transcription unit (2), first, a speech recognition program (201) generates Japanese text (211) that is faithful to the spoken voice (11). The speech recognition program (201) recognizes gaps (silent periods) between words and syllables as breaths or breaks in speech, and outputs "," or "." Even if there are no breaths or breaks in the actual speech, the listener will welcome the light-hearted style of speech and will be able to clearly understand the meaning of the speech.

[0058] However, the accuracy of the text (212) that is a direct transcription of the utterance decreases in the subsequent translation unit (3). Therefore, next, the AI ​​proofreading program (202) of the transcription unit (2) proofreads the Japanese text (211). As a result, the Japanese text (212) is output. The meaning of the utterance in the Japanese text (212) is the same as that of the Japanese text (211). However, the Japanese text (212) outputs an expression that is more grammatically appropriate than the Japanese text (211) and makes the relationships between words clearer.

[0059] [Example 3] Figure 3 shows how the neural machine translation (301) of the translation unit (3) generates English text (311) from the above-mentioned Japanese text (212). In the English text (311), the Japanese text (212) is accurately translated into rhythmic English expressions.

[0060] The present invention can provide the market with a variety of translated content (50) that is more appealing to users who watch or listen to content. The audio of the translated content (50) cannot be identical to the audio of the original content's performers. In this sense, the translated content (50) has significantly higher quality than conventional dubbed content. The translated content (50) itself is a content product with high added value in fields such as entertainment, education, customer acquisition, and news reporting. Moreover, the translated content (50) has high market value as an advertising medium.

[0061] In particular, when the original content (10) is selected from animation works or radio broadcast content, the generated translated content (50) is likely to be accepted by listeners without any sense of incongruity. Many Japanese animation works are particularly popular overseas. The translated content (50) can contribute to the global expansion of the animation industry. Furthermore, advertising content (51) combined with the translated radio broadcast content (50) is expected to have a strong appeal to consumers. Therefore, the present invention can provide high-value-added program content and advertising content, particularly to the markets related to animation works and radio broadcasts.

[0062] 1 Input unit 10 Original radio broadcast content 11 Radio personality's speech audio 12 Jingle 2 Transcription unit 21 Radio personality's speech text in Japanese 201 Speech recognition program 211 Japanese text 202 AI proofreading program 212 Japanese text 3 Translation unit 31 English speech text 301 Neural machine translation 311 English text 4 Speech synthesis unit 40 Synthesized English speech audio 5 Output unit 50 English version radio broadcast content 51 Advertising content 52 Advertising database 7 Communication unit 100 Example of this system

Claims

1. A translation content generation system (100) including an input unit (1), a speech recognition unit (2), a translation unit (3), a speech synthesis unit (4), and an output unit (5), wherein the input unit (1) acquires the speech (11) of a speaker from the original content (10), the speech recognition unit (2) performs speech recognition on the speech (11) to generate the speech text (21) of the speaker in the original language (20), the translation unit (3) translates the speech text (21) into a language (30) other than the original language (20) to generate the speech text (31) in the language (30), the speech synthesis unit (4) synthesizes the speech (40) of reading the speech text (31) in the voice of the speaker based on the speech (11) of the speaker and the speech text (31), and the output unit (5) outputs the translation content (50) including the speech (40).

2. The translation content generation system (100) according to claim 1, wherein the original content (10) is radio broadcast content.

3. Further including a subtitle generation unit (6) that generates subtitles (60) based on the speech text (31), and the output unit (5) outputs the translation content (50) including the speech (40) and / or the subtitles (60).

4. The translation content generation system (100) according to claim 1, wherein the output unit (5) outputs the translation content (50) including the speech (40) and further the advertisement content (51).

5. The input unit (1) obtains the speech audio (11-1, 11-2,..., 11-n) of each of n (n is an integer of 2 or more) performers from the original content (10), and the speech-to-text conversion unit (2) converts the speech audio (11-1, 11-2,..., 11-n) into text to generate the speech text (21-1, 21-2,..., 21-n) of each of the n performers in the original language (20). The translation unit (3) translates the speech text (21-1, 21-2,..., 21-n) into a language (30) other than the original language (20) to generate the speech text (31-1, 31-2,..., 31-n) in the language (30). The speech synthesis unit (4) synthesizes the speech (40-1, 40-2,..., 40-n) of reading out the speech text (31-1, 31-2,..., 31-n) in the voices of the performers based on the speech audio (11-1, 11-2,..., 11-n) of the performers and the speech text (31-1, 31-2,..., 31-n). The output unit (5) outputs the translated content (50) including the speech (40-1, 40-2,..., 40-n), the translated content generation system (100) according to claim 1.

6. A device (101) that stores the data of the translated content (50) output by the translated content generation system (100) according to any one of claims 1 to 5.

7. A translated content distribution method (102) that transmits the translated content (50) output by the translated content generation system (100) according to any one of claims 1 to 5 to a terminal device via the Internet.

Citation Information

Patent Citations

  • Voice transmission / reception device

    JP1990196372A

  • Device for making audio into multiple languages and medium with program for making audio into multiple languages recorded thereon

    JP2002007396A

  • Content output device

    JP2010033351A

  • Device and method for broadcasting

    JP2017184056A

  • Method, Apparatus and System For Regenerating Voice Intonation In Automatically Dubbed Videos

    US20160021334A1