Audio modification using artificial intelligence
An AI model for voice dubbing adapts translations to match audio-visual content context, addressing labor and resource inefficiencies, and enhancing synchronization and viewer experience.
Patent Information
- Application Number
- PCT/US2024/023927
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-10-16
AI Technical Summary
Traditional voice dubbing techniques are labor-intensive, time-consuming, and often result in low-quality translations that fail to align with the original audio-visual content, leading to viewer distractions and increased computing resource consumption.
An AI model is trained to generate translated audio based on contextual information from prior video frames, including audio and visual elements, to achieve synchronized and high-quality voice dubbing.
This approach reduces manual correction time, conserves computing resources, and enhances viewer experience by ensuring isochrony between audio and video, improving overall efficiency and viewership.
Smart Images

Figure US2024023927_16102025_PF_FP_ABST
Abstract
Description
Attorney Docket No. : 25832.4411 (L4411PCT)AUDIO MODIFICATION USING ARTIFICIAL INTELLIGENCETECHNICAL FIELD
[0001] Aspects and implementations of the present disclosure relate to audio modification using artificial intelligence (Al).BACKGROUND
[0002] A platform (e.g., media content platform, streaming platform, social media platform, etc.) can provide users with media content. The platform can provide some media content in multiple languages. Translating speech in a video from an originally recorded language to another language may involve labor-intensive efforts of video dubbing. Video dubbing refers to combining additional or supplementary speech (dubbed speech, often in a foreign language) with original speech to create the finished soundtrack for the video. However, the dubbed speech may differ from the original speech and may not align with start and end times of the original speech. In addition, the dubbed speech may not visually match the onscreen speaker’s lip movement. As a result, the translated audio may appear out of sync and may not be appealing to viewers.SUMMARY
[0003] The below summary is a simplified summary of the disclosure in order to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is intended neither to identify key or critical elements of the disclosure, nor delineate any scope of the particular implementations of the disclosure or any scope of the claims. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] An aspect of the disclosure provides a computer-implemented method that includes identifying a video content item for translation to a first language to a second language. The video content item includes a set of video frames that each include one or more audio portions and one or more video portions. The method further includes providing, as input to an artificial intelligence (Al) model, contextual information for one or more initial video frames of the set of video frames. The contextual information pertains to an audio portion of the one or more initial video frames and a video portion of the one or more initial video frames. The input also includes translation content representing an audio portion of one or more subsequent video frames of the set of video frames for translation to a first language from a second language. The method further includes extracting, from one or more outputs of the Al model, the audioportion of the one or more subsequent video frames in the first language for the video content item. A speech component of the extracted audio portion in the first language corresponds to a speech component of the audio portion in the second language in view of the provided contextual information.
[0005] In some implementations, the contextual information includes a transcript for the audio portion of the one or more initial video frames.
[0006] In some implementations, the contextual information includes one or more of: an identifier of each speaker associated with the audio portion of the one or more initial video frames, an identifier of a gender for each speaker associated with the audio portion of the one or more initial video frames, an indication of at least one of an environment depicted by the video portion of the one or more initial video frames or one or more objects included in the environment, an indication of one or more motions or actions pertaining to a respective object depicted by the video portion, an indication of one or more additional sound events associated with the audio portion of the one or more initial video frames, or a content summary including a summarized description of at least a portion of content of the one or more initial video frames.
[0007] In some implementations, the method further includes obtaining the contextual information by providing at least the video portion of the one or more initial video frames as input to a video content analysis model that is trained to detect, based on given video content, at least one of an object depicted by the video content, an environment depicted by the video content, or a motion or action performed with respect to an object depicted by the video content. The method may further include extracting, based on one or more outputs of the video content analysis model, a set of labels pertaining to provided video portion. The set of label may include at least one of the identifier of the gender for each speaker associated with the audio portion of the one or more initial video frames, the indication of at least one of the environment or the objects included in the environment depicted by the video portion of the one or more initial video frames, or the indication of the one or more motions or actions pertaining to the respective object depicted by the video portion.
[0008] In some implementations providing the contextual information for the one or more initial video frames of the set of video frames as input to an Al model includes at least one of: providing an audio file including the audio portion of the one or more initial video frames as input to the Al model, or providing a video file including the video portion of the one or more initial video frames as input to the Al model.
[0009] In some implementations, the method further includes determining one or more speech components for the audio portion of the one or more subsequent video frames accordingto the second language and providing the determined one or more speech components as additional input to the Al model. The speech component of the audio portion in the second language may correspond to the speech component of the extracted audio portion in the first language is indicated by the determined one or more speech components.
[0010] In some implementations, the speech component of the extracted audio portion in the first language and the speech component of the audio portion in the second language include one or more of a speech duration, a sequence of lip shapes, or a speech structure.
[0011] In some implementations, the contextual information further includes at least one of an indication of the second language or an indication of a content type or a content genre associated with the video content item.
[0012] In some implementations, providing the contextual information and the translation content as the input to the Al model includes providing a prompt to the Al model. The prompt may include a request for the audio portion of the one or more subsequent video frames in the first language, a reference to the contextual information for the one or more initial frames, and a reference to the translation content representing the audio portion of the one or more subsequent frames.
[0013] In some implementations, the translation content includes an initial translation of the audio portion to the first language from the second language. The audio portion of the one or more subsequent video frames extracted from one or more outputs of the Al model includes an improved translation of the audio portion to the first language from the second language in view of the given contextual information.
[0014] In some implementations, the method further includes responsive to a request from a user for the video content item in the first language, providing, to a client device associated with the user, the extracted audio portion of the one or more subsequent video frames and the video portion of the one or more subsequent video frames for presentation to the user via a user interface (UI) of the client device.
[0015] In some implementations, the method further includes providing, as input to the Al model contextual information for the one or more initial video frames, and translation content representing an audio portion of the one or more initial video frames. The method further includes extracting, from one or more additional outputs of the Al model, the audio portion of the one or more initial video frames in the first language. The speech component of the extracted audio portion of the one or more initial video frames corresponds to the speech component of the audio portion of the one or more initial video frames in the second language.
[0016] In some implementations, a system including a memory and a set of one or more processing devices coupled to the memory is provided. The set of one or more processing devices is to perform operations including identifying a training video content item including a set of video frames that each include one or more audio portions and one or more video portions. The operations further include obtaining contextual information for one or more initial video frames of the set of video frames. The contextual information pertains to an audio portion of the one or more initial video frames and a video portion of the one or more initial video frames. The operations further include obtaining translation content representing a training audio portion of a subsequent video frame of the plurality of video frames. The training audio portion is translated to a first language from a second language. The operations further include performing one or more distortion operations to the translation content to generate distorted translation content including one or more translation errors in the first language. The operations further include generating training data to train an artificial intelligence (Al) model to generate a predicted audio portion of a video frame of an additional video content item in the first language in view of additional contextual information for one or more prior video frames of the additional video content item. The training data includes (i) training input data including the contextual information for the one or more initial video frames and the distorted translation content, and (ii) training output data including the translation content.
[0017] In some implementations, the performing the one or more distortion operations to the translation content includes removing at least one of a word or a phrase from the translation content.
[0018] In some implementations, performing the one or more distortion operations to the translation content includes providing the translation content as an input to first translation engine that is configured to translate content in the first language to content in the second language. The operations may further include providing an output of the first translation engine as an input to a second translation engine, where the second translation engine is configured to translate content in the second language to content in the first language. The operations may further include obtaining one or more outputs of the second translation engine. The one or more outputs may include the distorted translation content.
[0019] Performing the one or more distortion operations to the translation content may include shortening or lengthening a duration of the one or more phrases. In some implementations, the first translation engine or the second translation engine include one or more of a speech-to-speech (S2S) translation engine that is configured to translate an audio signal from an initial language to an audio signal in a target language, a speech-to-text (S2T)translation engine that is configured to translate an audio signal in an initial language to a text string in a target language, or a text-to-text (T2T) translation engine that is configured to translate a text string in an initial language to a text string in a target language.
[0020] In some implementations, performing the one or more distortion operations to the translation content includes determining for one or more phrases of the translation content, at least one of a paraphrased version of the one or more phrases or an elongated version of the one or more phrases. A speech duration of the paraphrased version is shorter than a speech duration of the one or more phrases and a speech duration of the elongated version is longer than the speech duration of the one or more phrases. The distorted translation content includes the at least one of the paraphrased version of the one or more phrases or the elongated version of the one or more phrases.
[0021] In some implementations, performing the one or more distortion operations to the translation content includes obtaining an audio file including the training audio portion of the subsequent video frame. The operations further include updating the audio file to include one or more additional audio signals. The one or more additional audio signals may include one or more of audio noise, background sounds, or additional dialogue or narration. The operations may further include providing the updated audio file as an input to a transcription engine that is configured to generate a set of text strings based on one or more audio signals of a provided audio file. The distorted translation content may include a generated set of text strings included in one or more outputs of the transcription engine.
[0022] In some implementations, the contextual information for the one or more initial frames includes at least one of: a transcript for the audio portion of the one or more initial video frames, an identifier of each speaker associated with the audio portion of the one or more initial video frames, an identifier of a gender for each speaker associated with the audio portion of the one or more initial video frames, an indication of at least one of an environment depicted by the video portion of the one or more initial video frames or one or more objects included in the environment, an indication of one or more motions or actions pertaining to a respective object depicted by the video portion, an indication of one or more speech components for the audio portion of the one or more initial video frames, the one or more speech components including at least one of a speech duration, a sequence of lip shapes, or a speech structure of the audio portion, an indication of one or more additional sound events associated with the audio portion of the one or more initial video frames, or a content summary including a summarized description of at least a portion of content of the one or more initial video frames.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Aspects and implementations of the present disclosure will be understood more fully from the detailed description given below and from the accompanying drawings of various aspects and implementations of the disclosure, which, however, should not be taken to limit the disclosure to the specific aspects or implementations, but are for explanation and understanding only.
[0024] FIG. 1 illustrates an example system architecture, in accordance with implementations of the present disclosure.
[0025] FIG. 2 illustrates an example video dubbing engine, in accordance with some aspects of the present disclosure.
[0026] FIG. 3 is a flow diagram for an example method of translation correction and voice dubbing adaptation, in accordance with some aspects of the present disclosure.
[0027] FIGs. 4A-4C illustrate an example of translation correction and voice dubbing adaptation, in accordance with some aspects of the present disclosure.
[0028] FIG. 5 is a block diagram for an example predictive system in accordance with some aspects of the present disclosure.
[0029] FIG. 6 is a flow diagram for an example method of training an artificial intelligence (Al) used for translation correction and voice dubbing adaptation, in accordance with some aspects of the present disclosure.
[0030] FIGs. 7A-7C illustrate examples of distortion operations applied to translation content, in accordance with some aspects of the present disclosure.
[0031] FIG. 8 is a block diagram illustrating an exemplary computer system, in accordance with implementations of the present disclosure.DETAILED DESCRIPTION
[0032] Aspects of the present disclosure are directed to modification of audio, particularly where modification includes improved translation and voice dubbing adaption using artificial intelligence (Al). Voice dubbing refers to a process where an audio track for an audiovisual media item including dialogue and / or narration in an original language is replaced with a translated version of the dialogue and / or narration in another language. Traditionally, the audio track for the audiovisual media item is transcribed (e.g., converted from an audio format to a textual format) in the original language, translated to the target language (e.g., by a human translator, by a translation engine of a computing system, etc.) and then re-recorded by an actor or actress in the target language. The re-recorded audio track is then overlaid with the videocontent of the audiovisual media item and provided to users requesting access to the audiovisual media item in the target language. As can be seen by the above, traditional voice dubbing techniques are labor-intensive, time-consuming, and, in some instances, require significant linguistic and acting expertise to achieve synchronization between audio and visual elements.
[0033] Some systems employ automated transcription and / or translation techniques to reduce the amount of resources (e.g., human labor, etc.) and / or time it takes to perform voice dubbing for an audiovisual media item. For example, in some systems, an audio signal including dialogue and / or narration in an original language is provided as an input to a transcription engine. The transcription engine converts the audio signal to one or more text strings indicating the dialogue and / or narration in the original language. The text strings output by the transcription engine are then provided as an input to a translation engine that translates the dialogue and / or narration of the text string(s) from the original language to a target language. In some instances, the translated text string(s) are then provided as an input to a text- to-speech engine that generates an audio signal based on the translated text string(s).
[0034] Although some automated transcription and / or translation techniques may reduce the overall amount of time it takes to perform voice dubbing, voice dubbing performed according to such techniques is often very low quality. For example, the transcription engine may not accurately detect a word or phrase in the original language and / or may misinterpret pronunciation or punctuation conventions of the dialogue or narration. If the text strings output by the transcription engine include such errors, the translation engine will not accurately translate the audio signal from the original language to the target language. In some instances, the translation engine can also introduce errors into the translation process, as the translation engine may mis-translate the text string(s) (e.g., by mis-detecting a context of the dialogue or narration) and / or may not determine an accurate or appropriate word or phrase for translation of the word or phrase from the original language into the target language.
[0035] Automated transcription and / or translation can consume a significant amount of computing resources (e.g., processing cycles, memory space, etc.), which are therefore not available to other processes of the system. In some instances, transcribed and / or translated content that is used for the voice dubbing is corrected (e.g., by a human being) in order to address the errors introduced by the transcription engine and / or the translation engine, as described above. This can increase the overall amount of time for performing and / or can increase the overall amount of computing resources consumed for the voice dubbing process. Accordingly, a fewer amount of computing resources available for other processes of the system, which decreases the overall efficiency and increases the overall latency of the system.
[0036] In addition, both manual and automated voice dubbing techniques are frequently unable to match the translated audio of an audiovisual media item with the corresponding video of the audiovisual media item. For example, the translated audio may not align (or approximately align) with the lip movements depicted by the corresponding video. Isochrony refers to the proper alignment of an audio portion of audiovisual content with the video portion of the audiovisual content. Non-isochronic content can be distracting to a viewer because the viewer may see the lips (e.g., of a person or character depicted by the video) moving without hearing a corresponding voice, or vice versa. Viewer distractions can lead to lower enjoyment, intelligibility, and / or viewership of the video content. Manual and automated voice dubbing techniques may attempt to address the non-isochrony of translated audio by speeding up or slowing down the translated dialogue or narration to meet the time constraints of the dialogue in the original language, or allowing for extra time in a visual shot for the translated dialogue or narration to continue after the original language dialogue or narration would have ended. However, sped up speech can be difficult for a viewer or listener to understand and, in some instances, the transition between speech at the original pace to the speech at the sped up pace can be very distracting to the viewer or listener. In addition, slowed down speech may distort the creative intention for the scene, which can also be distracting to the viewer or listener. In addition, it can take an even longer amount of time (and therefore a larger amount of computing resources) to address the above described non-isochrony, which, in some instances, ends up being more distracting to the viewer. Further, such attempts also fail to consider the overall context (e.g., storyline, character development, etc.) of the scene that includes the translated dialogue, which can cause such attempts to impact the story of the content, as intended by the content’s author or creator.
[0037] Implementations of the present disclosure address the above and other deficiencies by providing techniques for translation correction and voice dubbing adaption using artificial intelligence (Al). According to embodiments of the present disclosure, an Al model (e.g., a large language model (LLM)) can be trained to generate, for one or more video frames of a video media item, a predicted audio portion (e.g., an audio signal, a set of one or more text strings representing the translated audio signal, etc.) that is translated from an original language to a target language in view of contextual information for prior video frames of the media item. The contextual information can include information pertaining to audio and video for the prior video frames of the media item. As described herein, contextual information can include, for example, a transcript for the audio of the prior video frames, an identifier of each speaker that provides dialogue or narration of the audio portion, an identifier for a gender of each speakerthat provides dialogue or narration of the audio portion, an indication of an environment and / or one or more objects in the environment depicted by the video for the prior video frames, an indication of a motion and / or an action pertaining to a respective object depicted by the video, an indication of one or more additional sound events (e.g., other than dialogue or narration, associated with the audio portion of the prior video frames, a content summary of the content of the prior video frames, and so forth. In additional or alternative embodiments, the contextual information can indicate one or more speech components (e.g., speech duration, a sequence of lip shapes, a speech structure, etc.) associated with the audio of the prior video frames and / or the subsequent video frames according to the original language.
[0038] As will be described in further detail here, the Al model can be trained based on training data that includes ground truth transcript data associated with one or more video frames of a training video media item, a distorted version of the ground truth transcript data, and contextual information pertaining to the audio and / or the video of one or more prior video frames of the training video media item. The distorted version of the ground truth transcript data can incorporate one or more translation and / or transcription errors into the ground truth transcript data, such as omitted or mistranslated words or phrases, missing or extra punctuation, and so forth. Upon training of the model, a platform (e.g., a content sharing platform, etc.) can provide translation content (e g., audio content and / or text strings indicating the audio that is to be translated) associated with one or more video frames of a video media item as an input to the Al model with contextual information for one or more prior video frames of the video media item. The Al model can generate predicted translation content in the target language, where a speech component (e.g., speech duration, sequence of lip shapes, speech structure, etc.) of the translation content in the target language matches (or substantially matches) a speech component of the audio in the original language, in view of the provided contextual information.
[0039] Aspects of the present disclosure provide techniques for enabling systems to perform isochronic video dubbing based on Al translation and video dubbing adaptation techniques. As described herein, the Al model is trained to predict translated content of a video item in a target language based on contextual information for prior video frames of the media item. The contextual information can include information that impacts the isochrony of the translated audio with the corresponding video, which can inform the Al model of optimal translation conventions, when applied to incoming translation content. In an illustrative example, the contextual information can indicate that the speaker that provided the dialogue to be translated has a female gender. The output of the Al model can include the translated content(e.g., a translated audio signal, translated text strings, etc.) including proper gendered language, in view of the target language, and that also matches (or at least partially matches) the speech pattern of the dialogue to the lip movements in the corresponding video. In another illustrative example, the contextual information can indicate that, in a prior scene of the video media item, a character lost an object. The output of the Al model can include a reference to the object in the translated audio data, so to elongate (or paraphrase) the dialogue in the target language to match the speech duration of the lip movements in the corresponding video.
[0040] As seen above, embodiments of the present disclosure enable a platform to provide corrected and / or optimized voice dubbing for video media items, based on one or more outputs of an Al model. Accordingly, manual correction of the voice dubbing is not performed (e.g., by a human being), and the media item in the target language can be more quickly available to users of the platform. As correction of the voice dubbing is not performed, fewer computing resources (e.g., processing cycles, memory space, etc.) are consumed by the overall system, which improves an efficiency and / or latency of the overall system. Further, as the Al model can provide translated content that has speech components that match (or substantially match) the speech components of the corresponding video, an isochrony of the voice dubbed video media item is improved, which can improve the viewer’s overall experience with the media item, and can improve viewership of the media item.
[0041] FIG. 1 illustrates an example system architecture 100, in accordance with implementations of the present disclosure. The system architecture 100 (also referred to as “system” herein) includes client devices 102A-N, a data store 110, a platform 120 (e.g., a content sharing platform), and / or one or more server machines (e.g., server machine 150), each connected to a network 104. In some embodiments, system 100 can additionally or alternatively include a predictive system 180. In implementations, network 104 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., Ethernet network), a wireless network (e.g., an 802.11 network or a Wi-Fi network), a cellular network (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or a combination thereof.
[0042] In some implementations, data store 110 can be a persistent storage capable of storing data as well as data structures to tag, organize, and index the data. A data item can include video content 121, in accordance with implementations described herein. Data store 110 can be hosted by one or more storage devices, such as main memory, magnetic or optical storage based disks, tapes, or hard drives, and so forth. In some implementations, data store 110 can be a network-attached file server or some other type of persistent storage such as anobject-oriented database, a relational database, and so forth, that may be hosted by platform 120 or one or more different machines coupled to the platform 120 via network 104.
[0043] The client devices 102A-N (collectively and individually referred to as client device(s) 102 herein) may each include computing devices such as personal computers (PCs), laptops, mobile phones, smart phones, tablet computers, netbook computers, network- connected televisions, etc. In some implementations, client devices 102A-N may also be referred to as “user devices.” Each client device may include a content viewer. In some implementations, a content viewer may be an application that provides a user interface (UI) for users to view or upload content, such as images, video items, web pages, documents, etc. For example, the content viewer may be a web browser that can access, retrieve, present, and / or navigate content (e.g., web pages such as Hyper Text Markup Language (HTML) pages, digital media items, etc.) served by a web server. The content viewer may render, display, and / or present the content to a user. The content viewer may also include an embedded media player (e.g., a Flash® player or an HTML5 player) that is embedded in a web page (e.g., a web page that may provide information about a product sold by an online merchant). In another example, the content viewer may be a standalone application (e.g., a mobile application or app) that allows users to view digital media items (e.g., digital video items, digital images, electronic books, etc.) and / or allows users to view, edit, and / or create video content 122. In some implementations, the content viewer can be a media content platform application for users to view, generate, edit, and / or upload the video content 122 to platform 120. As such, the media content viewers can be provided to the client devices 102 by platform 120.
[0044] In some implementations, platform 120 and / or server machine 130 can be one or more computing devices (such as a rackmount server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, etc.), data stores (e.g., hard disks, memories, databases), networks, software components, and / or hardware components that may be used to provide a user with access to video content 122 and / or provide the video content 122 to the user. For example, platform 120 can be a content sharing platform. The content sharing platform can allow a user to create, edit, access, or share with other users, video content stored at data store 110. Platform 120 can also include a website (e.g., a webpage) or application back-end software that can be used to provide a user with access to video content 122. In some embodiments, a client device 102 can request a particular video content 122 from platform 120. Platform 120 can identify the video content 122 (e.g., stored in data store 110) and can determine how to present the client device 102 withthe requested video content 122. In some implementations, platform 120 can provide the client device 122 with access to the file via the GUI of the media content viewer, as described above.
[0045] As described herein, video content 122 of platform 120 can be made up of a sequence of video frames, where each video frame has a video portion (e.g., depicting a scene) and a corresponding audio portion. In some instances, the audio portion can include dialogue (e g., a conversation between two or more characters of a scene) and / or narration (e g., spoken words of a narrator or other such character of the scene) provided by one or more people or characters depicted (or otherwise associated with) a scene of the video portion. For example, the video portion can include one or more video signals depicting two or more characters engaging in a dialogue. The corresponding audio portion can include audio signals representing the dialogue. In another example, the video portion can include one or more video signals depicting a scene (e.g., with or without characters) and the corresponding audio portion can include audio signals representing a narration of the scene. The audio portion for video content 122 can include other types of audio content, in some instances. For example, the audio portion for video content 122 can include sound effects associated with the scene depicted by a video portion, background music associated with the scene, background dialogue (e.g., from other characters or objects depicted in the scene), and so forth. In some instances, the audio portion for the video content 122 can include noise (e.g., static, cross talk, etc.) or other such audial components that impact the quality of the audio portion. Such noise or audial components may be included in the audio portion of the video content 122 based on a type of device used to generate the video content 122, or based on other such factors.
[0046] In some embodiments, video content 122 of platform 120 can be associated with one or more original languages. Users of platform 120, may want to access the video content 122 in other languages (e g., other than the original languages). In some embodiments, platform 120 can provide users with access to dubbed video content 124 which, in some instances, includes the video portion of the original video content 122 with the audio portion translated to a target language. Voice dubbing engine 152 may generate or otherwise provide dubbed video content 124 to users of platform 120 (e.g., in accordance with a received request from a client device 102). Further details regarding voice dubbing engine 152 are described herein with respect to FIGs. 2 and 3.
[0047] As illustrated in FIG. 1, system 100 can include a predictive system 180, in some embodiments. Predictive system 180 can implement one or more artificial intelligence (Al) and / or machine learning (ML) techniques to generate or otherwise obtain dubbed video content 124. In some embodiments, predictive system 180 can train an Al model (e.g., a large languagemodel (LLM)) to generate predicted translated audio for video content 122 of platform 120 based on contextual information obtained for an audio portion and / or a video portion of one or more prior frames of the video content 122. Further details regarding predictive system 180 and the trained Al model are provided herein with respect to FIGs. 5 and 6.
[0048] It should be noted that although FIG. 1 illustrates video dubbing engine 152 as part of platform 120, in additional or alternative embodiments, video dubbing engine 152 can reside on one or more server machines that are remote from platform 120. For example, video dubbing engine 152 can reside at server machine 150. In other or similar embodiments, video dubbing engine 152 can reside on one or more client devices 102. Further, although FIG. 1 illustrates predictive system 180 as remote from platform 120, in additional or alternative embodiments, predictive system 180 can reside on platform 120, server machine(s) 150, client device 102, and / or any other component of system 100. It should be noted that in some other implementations, the functions of platform 120, server machine 150, and / or predictive system(s) 180 can be provided by more or a fewer number of machines. For example, in some implementations, components and / or modules of platform 120, server machine 150, and / or predictive system(s) 180 may be integrated into a single machine, while in other implementations components and / or modules of any of platform 120, server machine 150, and / or predictive system(s) 180 may be integrated into multiple machines. In addition, in some implementations, components and / or modules of server machine 150, and / or predictive system(s) 180 into platform 120.
[0049] In general, functions described in implementations as being performed by platform 120, server machine 150, and / or predictive system(s) 180 can also be performed on the client device 102 in other implementations. In addition, the functionality attributed to a particular component can be performed by different or multiple components operating together. Platform 120 can also be accessed as a service provided to other systems or devices through appropriate application programming interfaces, and thus is not limited to use in websites.
[0050] In implementations of the disclosure, a “user” can be represented as a single individual. However, other implementations of the disclosure encompass a “user” being an entity controlled by a set of users and / or an automated source. For example, a set of individual users federated as a community in a social network can be considered a “user.” Further to the descriptions above, a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content orcommunications from a server. In addition, certain data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity can be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location can be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user can have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0051] FIG. 2 illustrates an example video dubbing engine 152, in accordance with some aspects of the present disclosure. As described above, platform 120 may be a content sharing platform that provides users with access to content (e.g., video content 122). For instance, upon receiving a request for video content 122 from a client device 102, platform 120 can identify the requested video content 122 (e.g., from data store 110, from another memory of system 100) and can provide the requested video content 122 to the client device 102 (e.g., via network 104). Video content 122 can include an audio portion 252 and a video portion 254, as described above. In some embodiments, the audio portion 252 can include dialogue and / or narration in an original language. Platform 120 may receive a request from a client device 102 for access to the video content 122 in a different language than the original language. As described herein, voice dubbing engine 152 can generate or otherwise obtain dubbed video content 124, which includes the audio portion 252 translated to the target language for presentation with the video portion 254 of the original video content 122.
[0052] As illustrated in FIG. 2, voice dubbing engine 152 can include at least a context module 212, an audio translation module 214, a speech component module 216, and / or a video dubbing module 218. Further details regarding voice dubbing engine 152 and the one or more modules of voice dubbing engine 152 are provided below, with respect to FIGs. 2 and 3.
[0053] In some embodiments, voice dubbing engine 152 can include or be otherwise connected to one or more translation engines 220 that are configured to translate content from an original language to a target language. In some embodiments, a translation engine 220 can be or otherwise include a speech-to-speech (S2S) translation engine that takes an audio signal in an original language as an input and provides an audio signal in the target language as an output. In other or similar embodiments, a translation engine 220 can be or otherwise include a speech-to-text (S2T) translation engine, that takes an audio signal in an original language as an input and provides a set of text strings representing the audio signal in the target language as an output. In yet other or similar embodiments, the translation engine 220 can be or otherwise include a text-to-speech (T2S) translation engine (e.g., which takes text string(s) in an originallanguage as an input and outputs an audio signal of the text string(s) in the target language) or a text-to-text (T2T) translation engine (e.g., which takes text string(s) in the original language as an input and outputs text string(s) in the target language). In some embodiments, the translation engine 220 can have a cascade architecture, where input data is fed to a cascade of modules (not shown) each associated with a respective function or capability corresponding to the translation of the data from the original language to the target language. For example, translation engine 220 can include an automatic speech recognition (ASR) module that converts speech of an audio signal into one or more text strings, a text-to-text machine translation (MT) module that translates text in an original language to a target language, a text- to-speech (TTS) synthesis module that converts text to an audio signal, and so forth. It should be noted that embodiments and examples for translation engine 220 are provided for purposes of example and explanation only. Any type of translation engine 220 implementing any type of modules can be applied, in accordance with embodiments of the present disclosure.
[0054] In additional or alternative embodiments, voice dubbing engine 152 can include or otherwise be connected to one or more transcription engines 222 that is configured to convert an audio signal in an original language to a set of text strings in the original language. In some embodiments, voice dubbing engine 152 can provide the audio portion 252 of video content 122 to transcription engine 222 and transcription engine 222 can segment the audio signal in to two or more portions (e.g., in accordance with a protocol of transcription engine 222). Transcription engine 222 can then associate speech included in the segmented audio signal with sound units and convert the sound units to one or more text strings. In some embodiments, transcription engine 222 can employ one or more Al models to determine a word or phrase associated with each sound unit and / or generate the corresponding text string.
[0055] It should be noted that although some embodiments of the present disclosure provide that translation engine(s) 220 and / or transcription engine(s) 222 are part of platform 120, in other or similar embodiments, translation engine(s) 220 and / or transcription engine(s) 222 may not be part of platform 120. For example, in some embodiments, translation engine(s) 220 and / or transcription engine(s) 222 may be a part of another platform (e.g., of system 100 or of another system) that is accessible to platform 120.
[0056] As illustrated in FIG. 2, platform 120, voice dubbing engine 152, and / or translation engine(s) 220 can be connected to or otherwise have access to memory 250. In some embodiments, memory 250 can include one or more portions of data store 110. In other or similar embodiments, memory 250 can include any memory of system 100, a component ordevice of system 100, and / or a component or device connected to system 100 (e.g., via network 104 and / or another network).
[0057] FIG. 3 is a flow diagram for an example method 300 of audio translation and voice dubbing adaptation, in accordance with some aspects of the present disclosure. Method 300 can be performed by processing logic that may include hardware (circuitry, dedicated logic, etc ), software (e g , instructions run on a processing device), or a combination thereof. In one implementation, some or all the operations of method 300 can be performed by one or more components of system 100 of FIG. 1 and / or one or more components of FIG. 2. In some embodiments, one or more operations of method 300 may be performed by one or more components of voice dubbing engine 152, as described herein.
[0058] At block 302, processing logic identifies a video content item for translation to a first language (e.g., a target language) from a second language (e g., an initial language or an original language), the video content item including a set of video frames that each include one or more audio portions and one or more video portions. The video content item can include or otherwise correspond to video content 122, in some embodiments. The one or more audio portions can be or otherwise include audio portion 252 and the one or more video portions can be or otherwise include video portion 254. As described above, video content 122 can be made up of sequence of video frames that each include respective audio portions 252 and video portions 254. In some instances, a subset of the sequence of video frames can correspond to a scene of video content 122. As described herein, embodiments and examples of the present disclosure may refer to video frames of a scene of video content 122, where each scene includes one or more initial video frames and one or more subsequent video frames (e.g., of the sequence of video frames for the video content 122). It should be noted that such embodiments can apply to any video frames of video content 122. For example, initial video frames, as described herein, can correspond to ending of a first scene and subsequent video frames can correspond to a beginning of a second scene.
[0059] In some embodiments, voice dubbing engine 152 can identify the video content 122 for voice dubbing based on a request from a user of platform 120. For example, platform 120 can provide users with access to video content 122 in an original language. One or more users may request (e g., via respective client devices 102) access to the video content 122 in another language (i.e., a target language). Platform 120 can receive the request from client devices 102 and, in some embodiments, can identify such video content 122 for voice dubbing based on the received requests. In some embodiments, platform 120 and / or voice dubbing engine 152 can identify such video content 122 for voice dubbing upon determining that a threshold numberof voice dubbing requests have been received for the video content 122. In other or similar embodiments, a user that provided the video content 122 to platform 120 (e.g., a creator of video content 122) can provide the request (e g., via a client device 102) to generate the dubbed video content 124 in the target language. Platform 120 and / or voice dubbing engine 152 can identify the video content 122 for voice dubbing upon receiving the request from such user.
[0060] Referring back to FIG. 3, at block 304, processing logic provides one or more inputs to an artificial intelligence (Al) model. The one or more inputs can include, as indicated by block 306, contextual information for one or more initial video frames of the set of video frames. The contextual information can pertain to an audio portion of the one or more initial frames and a video portion of the one or more initial video frames, in some embodiments. The one or more inputs can also include, as indicated by block 308, translation content representing an audio portion of one or more subsequent video frames of the set of video frames. Further details regarding the inputs of block 306 and block 308 are described below. In some embodiments, the Al model can be a large language model (LLM) and / or a generative Al model that is trained to generate predicted audio translated in a target language based at least on contextual information pertaining to one or more prior video frames of the video content 122. Further details regarding the training of the Al model are described herein. As illustrated by FIG. 2, memory 250 can include one or more Al models that, in some embodiments, may be trained to perform different or distinct tasks, as described herein. For purposes of example and illustration only, the Al model 260 that is trained to generate the predicted translated audio is referred to as a translation model 260. Other models are described below.
[0061] As described above, one or more inputs to the translation Al model 260 can include contextual information for one or more initial video frames of the set of video frames of video content 122. In some embodiments, contextual information 256 can include information that indicates prior dialogue, prior narration, surrounding visuals, the genre, and / or the purpose (e.g., as intended by the video content creator) of the portion of the scene depicted by the initial video frames of the set of video frames. As described herein, the contextual information 256 for the prior video frames of the video content 122 can indicate and inform the overall narrative of the video content 122, which allows for more accurate translation and voice dubbing adaptation, as described herein. Examples and embodiments regarding the type of information that can be included as contextual information 256 are provided below.
[0062] In some embodiments, context module 212 of voice dubbing engine 152 can obtain contextual information 256, as described herein. In other or similar embodiments, voice dubbing engine 152 may provide the audio portion 252 and the video portion 254 of the initialvideo frames as an input to the translation Al model 260 and the translation Al model 260 may predict or otherwise obtain contextual information 256 based on the provided audio portion 252 and the provided video portion 254.
[0063] It should be noted that the use of the phrases “initial video frames” and “subsequent video frames” are provided for purposes of example and illustration only and is not intended to be limiting. For instance, “initial video frames” as described herein, refers to any video frames that are located earlier in the sequence of video frames of video content 122 than “subsequent video frames” of video content 122. For example, a sequence of video frames of video content 122 can include video frame 0 through video frame X. The subsequent video frame, as described herein, can include video frame X and the initial video frame can include video frame X-l . In another example, the initial video frame can include video frame 0 and the subsequent video frame can include video frame 1. While embodiments of the present disclosure can be applied on a frame-by-frame basis, embodiments of the present disclosure can also be applied to sets of frames. For instance, in accordance with the prior example, the subsequent video frames of video content 122 can include video frame X-l 00 through video frame X, while the initial video frames can include video frame X-200 through video frame X- 101. It should also be noted that initial video frames may not be directly adjacent to subsequent video frames. For example, the subsequent video frames can include video frame X-100 through video frame X, while the initial video frames include video frame 0 through video frame 99.
[0064] In some embodiments, the contextual information 256 can include a transcript for the audio portion 252 of the one or more initial video frames. The transcript can include a set of text strings indicating at least one of the dialogue or the narration of the audio portion 252 for the one or more initial video frames. In some embodiments, the transcript can be in the target language. In some embodiments, context module 212 can provide the audio portion 252 of the initial video frames to transcription engine 222 and can obtain, as an output of transcription engine 222, a set of text strings indicating the dialogue or narration of the audio portion 252 for the initial video frames. In other or similar embodiment, context module 212 can obtain the set of text strings for the audio portion 252 of the initial video frames based on the translation of the audio portion 252 from the original language to the target language, as described herein. In such embodiments, upon obtaining the translated audio portion 252 (e.g., as an output of translation Al model 260), context module 212 can update contextual information 256 at memory 250 to include the text string(s) indicating the translated audio content.
[0065] In other or similar embodiments, contextual information 256 can indicate a context of one or more visual components of the scene corresponding to the initial video frames. For example, contextual information 256 can include, an identifier of each speaker of the audio portion 252 of the initial video frames, an identifier of a gender for each speaker associated with audio portion 252 of the initial video frames, an indication of an environment depicted by the video portion 254 of the initial video frames and / or one or more objects included in the depicted environment, an indication of one or more motions or actions pertaining to a respective object depicted the video portion 254 of the initial video frames, an indication of one or more additional sound events (e g., other than dialogue or narration) associated with the audio portion 252 of the initial video frames, a content summary including a summarized description of at least a portion of the content of the initial video frames, and so forth.
[0066] In some embodiments, a creator or developer associated with video content 122 can provide the contextual information 256 pertaining to the visual components of the scene to platform 120. For example, such contextual information 256 can be indicated by a script for the scene. The creator associated with video content 122 can provide a file for an electronic document including the scene script with the video content 122, in some embodiments. Context module 122 can compare the dialogue or narration included in the audio portion 252 for the initial video frames to the dialogue or narration indicated by the provided script and can identify contextual information 256 corresponding to the initial video frames based on the comparison in some embodiments.
[0067] In other or similar embodiments, context module 212 can obtain the contextual information 256 pertaining to such video components based on one or more outputs of an Al model 260. In some embodiments, an Al model 260 (e.g., stored at memory 250 or otherwise accessible to platform 120) can be trained to detect, based on video content, at least one of an object depicted by the video content, an environment depicted by the video content, or a motion or action performed with respect to an object depicted by the video content. Such Al model 260 is referred to as a video content analysis model 260B herein. In some embodiments, video content analysis model 260B can be trained or otherwise provided by predictive system 180. In other or similar embodiments, platform 120 can access video content analysis model 260B based on resources or tools provided by another platform or system.
[0068] Context module 212 can provide the video portion 254 of the initial video frames as an input to the video content analysis model 260B and can obtain one or more outputs of the video content analysis model 260B. The one or more outputs can indicate at least one of the gender for each speaker associated with the audio portion 252, the environment or the objectsincluded in the environment depicted by the provided video portion 254, and / or a motion or action pertaining to the respective object depicted by the provided video portion 254. Context module 212 can extract a set of labels associated with such visual information indicated by the one or more outputs, and can store the extracted set of labels at memory 250 as contextual information 258.
[0069] In some embodiments, Al model(s) 260 can include multiple video content analysis models 260B that are each trained to perform a particular task pertaining to the analysis of the visual components of video portion 254, as described above. For example, Al model(s) 260 can include a speaker diarization model that is trained to assign an identifier (e g., a character identifier, a numerical identifier, etc.) to each speaker associated with audio portion 252, a gender estimation model that is trained to predict a gender associated with a speaker (e.g., regardless of whether the speaker is depicted by the video portion 254 or not), an object detection model that is trained to detect objects in an environment depicted by a video portion 254 of one or more video frames, a content summarization model that is trained to predict a summary of the video content 122), and so forth. In some embodiments, the content summarization model can implement unimodal approaches (e.g., that analyze a visual modality of video content 122 for feature extraction) and / or multimodal approaches (e.g., that analyze textual metadata) for predicting the summary of the video content 122. Context module 212 can provide the audio portion 252 and / or the video portion 254 of the initial video frames to such models 260, as described above, and can obtain the contextual information 256 based on one or more outputs of such models, as described above.
[0070] In some embodiments, contextual information 256 can include information regarding other audio signals of audio portion 252 (e g., that do not include or correspond to dialogue or narration). For example, context module 212 can determine, based on the transcript obtained for the audio portion 252 of the initial video frames and / or the data obtained from the one or more video content analysis models 260B, that one or more sound events (e.g., a car horn, a knock on the door, etc.) are included in audio portion 252 for the initial video frames. Context module 212 can include an indication of the sound event with contextual information 256, in some embodiments. In other or similar embodiments, context module 212 can obtain information pertaining to the sound events and / or other sound included in the audio portion 252 (e.g., background music, background dialogue that does not pertain to the storyline of video content 122) according to other techniques.
[0071] Contextual information 256 can additionally or alternatively include other information pertaining to the video content 122. For example, contextual information 256 caninclude an indication of the original language of video content 122, a genre associated with video content 122, and so forth.
[0072] FIG. 4A illustrates example contextual information 400 for initial video frames of video content 122, as described above. Example contextual information 400 can be obtained according to any of the techniques described above. As illustrated by FIG. 4A, example contextual information 400 can indicate a source language of the video content 122 (e.g., Tamil) and / or one or more genres of the video content 122 (e.g., action, science fiction fantasy, comedy, etc.). Example contextual information 400 can also indicate, for one or more frame subsets of the initial frames, an indication of a scene and / or one more objects of the scene, dialogue of the scene, and / or one or more sound events of the scene. For example, as illustrated by FIG. 4A, contextual information 400 can indicate that in a first subset of video frames, the video portion 254 depicts a wide view of an arena. Contextual information 400 can also indicate a change in the environment between video frame subsets (e.g., in a second subset of video frames, the video portion 254 depicts a crowd on its feed and spotlights change to bright stadium lights). Contextual information 400 further indicates dialogue or narration from one or more speakers of the audio portion 252 of the initial frames. For example, contextual information 400 indicates at frame subset 3 that a male speaker provides the dialogue of “He’s undefeated. He’s the reigning... He’s the defending... Ladies and gentlemen... I give you...” Other contextual information 256 obtained by context module 212 is included in example contextual information 400 (e.g., an indication of sound events), as illustrated by FIG. 4A.
[0073] As described above, in some embodiments, context module 212 of voice dubbing engine 152 can obtain the contextual information 256 for the initial video frames, as described above, and can provide the obtained contextual information 256 as the input to voice dubbing Al model 260A, as described above. In other or similar embodiments, context module 212 (or another module of voice dubbing engine 152) can provide the audio portion 252 and / or the video portion 254 as an input to the voice dubbing Al model 260A and the voice dubbing Al model 260A can predict or otherwise obtain the contextual information 256.
[0074] Referring back to FIG. 3, as indicated above, the one or more inputs provided to voice dubbing Al model 260A can include translation content representing an audio portion of one or more subsequent video frames of the set of video frames to be translated to the target language. In some embodiments, the translation content can include an indication of the audio portion 252 of the subsequent video frames that is to be translated. For example, the translation content can include an audio signal that includes the dialogue and / or narration of the audio portion 252 that is to be translated. In other or similar embodiments, the translation content caninclude a set of text strings (e.g., obtained from transcription engine(s) 222) that includes a textual version of the dialogue and / or narration for translation.
[0075] In other or similar embodiments, the translation content provided as input to the voice dubbing Al model 260A can include an initial translation of the audio portion 252 of the subsequent video frames from the original language to the target language. For example, in some embodiments, audio translation module 214 of voice dubbing engine 152 can provide the audio portion 252 (e.g., as an audio signal, as a set of text strings, etc.) to translation engine 220 and translation engine 220 can translate the provided audio portion 252 from the original language to the target language, as described above Audio translation module 214 can obtain one or more outputs from translation engine(s) 220 which can include an audio signal and / or a set of text strings indicating the audio portion 252 of the subsequent video frames in the target language. Audio translation module 214 can store the translated audio portion at memory 250 as translated audio portion 258, in some embodiments. As described above, audio translation module 214 (and / or another module of voice dubbing engine 152) can provide the translated audio portion 258 as an input to the voice dubbing Al model 260A.
[0076] FIG. 4B illustrates an example translated audio portion 402, in accordance with aspects of the present disclosure. FIG. 4C illustrates the target translated audio portion 404, in accordance with aspects of the present disclosure. Target translated audio portion 404 can include a correct and adapted translation for the voice dubbing of video content 122, as described herein. As can be seen with respect to FIG. 4C, translated audio portion 402 can include one or more translation errors. For example, translated audio portion 402 includes a punctuation error between the phrases “where have you been” and “So much has happened since I last saw you.” In another example, translated audio portion 402 can include a mistranslation pertaining to the phrase “Like yesterday, it’s still beautiful,” where the target translated audio portion 404 includes the phrase “Like, yesterday, so that’s still pretty fresh.”
[0077] Referring back to FIGs. 2 and 3, in additional or alternative embodiments, voice dubbing engine 152 can provide additional information pertaining to the initial video frames and / or the subsequent video frames as an input to the voice dubbing Al model. For example, speech component module 216 can obtain information associated with one or more speech components of the audio portion 252 of the initial video frames and / or the subsequent video frames in the original language, in some embodiments. A speech component can include, but is not limited to, a duration of speech by a speaker for one or more words or phrases (referred to as speech duration), a sequence of lip shapes made by a speaker that provided the words or phrases, and / or a speech structure (e.g., phonemes, word ordering, prosodic pauses, etc.) of thespoken words or phrases. The duration of speech by a speaker can indicate the amount of time that the speaker provided words or phases of particular speech, in some embodiments. In additional or alternative embodiments, the duration of speech can indicate the amount of time that the speaker provided the words or phrases and pauses between the words or phrases. The provided speech and / or pauses between speech may be represented by viseme codes, as described below.
[0078] In some embodiments, speech component module 216 can obtain the speech components for the initial video frames and / or the subsequent video frames by providing the video portion 254 for the initial video frames and / or the subsequent video frames as an input to an Al model 260 that is trained to generate predicted speech component data for speakers depicted by given video content 124. Such model is referred to herein as a speech component Al model 260C. In some embodiments, speech component Al model 260C can be trained or otherwise provided by predictive system 180. In other or similar embodiments, speech component Al model 260C can be provided by another platform or system. Upon providing the video portion(s) 254 as an input to the speech component Al model 260C, speech component module 216 can obtain one or more outputs of the speech component Al model 216C, which indicates the speech component data for the speakers of the audio portion(s). Speech component module 216 can include the speech component data of the one or more outputs with contextual data 256, in some embodiments.
[0079] In some embodiments, the speech component data can be based on or otherwise indicated by visemes associated with the original language and / or the target language. A viseme refers to a facial expression and / or lip shape that is made by a speaker when the speaker utters a particular sound. In some embodiments, the speech component AU model 260C and / or speech component module 216 can compare the facial shapes of the speakers that provide dialogue and / or narration, as depicted by video portion(s) 254, and identify a corresponding viseme. The speech component data obtained by the speech component module 216 can include a series of viseme codes, which each correspond to a detected viseme, indicate that no viseme is detected or can be matched, for the speech provided by the speaker, and / or indicate that the speaker took a pause while providing the speech. Accordingly, such viseme codes can, in some instances, indicate the duration and / or structure of the speech provided by the speaker. As described above, the viseme codes can be included with contextual data 256, in some embodiments. In other or similar embodiments, the viseme codes can be obtained based on one or more outputs of the video content analysis model 256B, as described above.
[0080] As described above, context module 212 can obtain contextual information 256 for one or more prior video frames of video content 122, which is provided as an input to the Al model 260A to obtain translation content for a subsequent set of frames, in some embodiments. In other or similar embodiments, the frames for which translation content is sought may be the first video frames (e.g., video frame 0) of video content 122, and no prior video frames may be available for video content 122. In such embodiments, context module 212 may obtain contextual information associated with the video frames for which translation content is sought and provide such obtained contextual information as an input to the trained Al model 206A. For example, context module 212 can obtain metadata (e g , the original language, genre, etc.) associated with the video frames and include the obtained metadata in the contextual information. In another example, context module 212 can provide video content 122 for the video frames as input to the content summarization model and can obtain a content summary for the video frames based on one or more outputs of the content summarization model, as described above. The obtained content summary can be included in the contextual information 256 for such video frames. In yet other or similar embodiments, context module 212 can identify one or more other media items associated with the creator of the video content 122 for which voice dubbing is to be performed and can obtain contextual information 256 for the identified one or more other media items, as described herein.
[0081] It should be noted that although some embodiments of the present disclosure provide that contextual information 256 for prior video frames is provided as input to Al model 260A, contextual information 256 can also be obtained for subsequent video frames for which translation content is sought, as described above, and provided as additional or alternative input to the Al model 260 A, in some embodiments.
[0082] Referring back to FIG. 3, at block 310, processing logic extracts, from one or more outputs of the Al model, the audio portion of the one or more subsequent video frames in the first language. The one or more outputs of the Al model can include, in some embodiments, the translated content (e.g., an audio signal, a set of text strings, etc.) of the one or more subsequent video frames in the first language. As described above, the voice dubbing Al model 260A can predict a translated audio portion for the one or more subsequent video frames in the target language. One or more speech components of the translated audio portion 262 can correspond to (e.g., match or substantial match within a particular margin) speech components of the audio portion in the original language in view of the provided contextual information 256. Voice dubbing module 218 of voice dubbing engine 152 can extract the translated audio portion 262 from the one or more outputs of the Al model 260A. In some embodiments, thetranslated audio portion 262 can include a set of text strings that indicates the translated dialogue and / or narration in the target language. In such embodiments, voice dubbing module 218 can obtain an audio signal based on the set of text strings (e.g., by providing the set of text strings as an input to a text-to-speech module of translation engine(s) 220 or according to other techniques). In other or similar embodiments, the translated audio portion 262 can include an audio signal that includes the translated dialogue and / or narration.
[0083] As indicated above, FIG. 4C illustrates the target translated audio portion 404. Target translated audio portion 404 can be included or otherwise correspond to the translated audio portion 262 that is output by voice dubbing Al model 260 A, in accordance with previous embodiments and examples.
[0084] At block 312, processing logic provides the audio portion of the one or more subsequent video frames in the first language with the video portion of the one or more subsequent video frames to one or more client devices. As illustrated in FIG. 2, dubbed video content 124 can include the translated audio portion 262 and the video portion 254 of the original video content 122. In some embodiments, video dubbing module 218 can replace the original audio portion 252 of the original audio portion 252 with the translated audio portion 254 for dubbed video content 124. In some embodiments, voice dubbing module 218 can provide the dubbed video content 124 to a client device 124 (e.g., in accordance with a request from a user of client device 124).
[0085] In other or similar embodiments, voice dubbing Al model 260A can generate dubbed video content 124, which includes the translated audio portion 262. In such embodiments, voice dubbing module 218 can obtain the dubbed video content 124 based on one or more outputs of the voice dubbing Al model 260A and can provide the obtained dubbed video content 124 to the client device 102, as described above.
[0086] FIG. 5 is a block diagram for an example predictive system 180 in accordance with some aspects of the present disclosure. In some embodiments, predictive system 180 can train one or more Al models 260, as described herein. For example, predictive system 180 can train one or more of voice dubbing Al model 260A, video content analysis model 260B, and / or speech component Al model 260C. Details regarding training of voice dubbing Al model 260A are provided with respect to FIGs. 6 and FIGs. 7A-7C below.
[0087] As illustrated in FIG. 5, predictive system 180 can include a training set generator 512 (e.g., residing at server machine 510), a training engine 522, a validation engine 524, a selection engine 526, and / or a testing engine 528 (e.g., each residing at server machine 520), and / or a predictive component 552 (e.g., residing at server machine 550). In accordance withembodiments described herein, predictive component 552 can be a component of or otherwise accessible to platform 120 and / or client devices 102 of system 100. In some embodiments, predictive component 552 can include or otherwise be accessible by voice dubbing engine 152, described with respect to FIG. 2. Training set generator 512 may be capable of generating training data (e.g., a set of training inputs and a set of target outputs) to train model(s) 260.
[0088] As mentioned above, training set generator 512 can generate training data for training model(s) 260. Training set generator 512 obtain training data for training model(s) 260 and can organize or otherwise group the training data for training model(s) 260 (e.g., according to the purpose of the model(s)). Details regarding generating training data for training a voice dubbing Al model 260A are provide below with respect to FIG. 6.
[0089] FIG. 6 is a flow diagram for an example method 600 for training an artificial intelligence (Al) used for audio translation and voice dubbing adaptation, in accordance with some aspects of the present disclosure. Method 600 can be performed by processing logic that may include hardware (circuitry, dedicated logic, etc.), software (e.g., instructions run on a processing device), or a combination thereof. In one implementation, some or all the operations of method 600 can be performed by one or more components of system 100 of FIG. 1 and / or predictive system 180 of FIG. 5. In some embodiments, one or more operations of method 600 may be performed by training set generator 512.
[0090] At block 602, processing logic identifies a training video content item including a set of video frames that each include one or more audio portions and one or more video portions. In accordance with previously described embodiments, a training video content item can include one or more video content items 122 of platform 120. In some embodiments, one or more video content items 122 of platform 122 can be designated (e.g., by developers or operators of platform 120, by users of platform 120, etc.) to be used for generating training data to train one or more Al model(s) 260. Such video content 122 can be included in a training corpus (e.g., at data store 110 or at another memory of system 100) that stores such video content 122 and / or corresponding data to be used for training. In some embodiments, training set generator 512 can identify the training video content item from the training corpus.
[0091] In accordance with previously described embodiments, the training video content item can be made of a sequence of video frames that each have an audio portion 252 and a corresponding video portion 254. In some embodiments, training set generator 512 can identify a set of video frames of the training video content item for which training data is to be generated, as described below. Training set generator 512 can identify such set of video framesrandomly and / or according to a training protocol for platform 120 (e.g., as provided by a developer or operator of platform 120).
[0092] At block 604, processing logic obtains contextual information for one or more initial video frames of the set of video frames. The one or more initial video frames of the set of video frames for the training video content item can include any set of video frames that is prior to the identified set of video frames described above. As will be seen below, the identified set of video frames are referred to as “subsequent video frames” and the initial video frames are additionally or alternatively referred to as “prior video frames,” in accordance with previously described embodiments. As described above, contextual information can include information that indicates prior dialogue, prior narration, surrounding visuals, the genre, and / or the purpose (e.g., as intended by the video content creator) of the portion of the scene depicted by the initial video frames. Training set generator 512 can obtain contextual information for the initial video frames in accordance with embodiments described with respect to context module 212 and / or speech component 216 above.
[0093] At block 606, processing logic obtains translation content representing a training audio portion of a subsequent video frame of the set of video frames, the training audio portion translated to a first language from a second language. As described above, the training corpus can include data to be used for training using the designated training video content items. In some embodiments, the included data can include a transcript of a dialogue and / or a narration of one or more audio portions 252 of the training video content item. The included transcript may include an indication or notification that such transcript has been reviewed and validated (e.g., for correctness) by a language authority. Accordingly, such transcript may be a ground truth transcript. Training set generator 512 can identify the ground truth transcript for the training video content item identified at block 602 from the media corpus and can identify the section of the ground truth transcript that corresponds to the subsequent video frames.
[0094] In additional of alternative embodiments, training set generator 512 can obtain the ground truth transcript according to other techniques. For example, in some embodiments, training set generator 512 can provide the audio portion 252 for the subsequent video frames of the training video content item to one or more transcription engine(s) 222 and can obtain one or more outputs of transcription engine(s) 222, which includes a transcript of the provided audio portion 252. In some embodiments, training set generator 512 can provide the transcript to a client device 102 associated with a user (e.g., a developer, an operator, etc.) of predictive system 180 and can request validation of the transcript. Upon receiving a response to therequest that indicates that the transcript is validated (e.g., is correct), training set generator 512 can designate the transcript as a ground truth transcript.
[0095] In yet other or similar embodiments, training set generator 512 can provide a transcript obtained from one or more transcription engine(s) 222 as an input to a transcript validation model, which is trained to predict a validation score for a given transcript, the validation score indicating a level or degree of correctness in view of grammar and / or language conventions for the transcript. Training set generator 512 can obtain one or more outputs of the transcript validation model, which can indicate a validation score for the given transcript. Upon determining that the validation score indicated by the one or more outputs satisfies one or more validation criteria (e.g., meets or exceeds a threshold validation score), training set generator 512 can designate the transcript as a ground truth transcript.
[0096] At block 608, processing logic performs one or more distortion operations to the translation content to generate distorted translation content including one or more translation errors in the first language. The one or more distortion operations can include any type of operations that introduce one or more translation errors that may be found in translated data (e.g., obtained from one or more outputs of a translation engine 220). FIGs. 7A-7C below illustrate example techniques for introducing translation errors into the ground truth transcript, in accordance with embodiments of the present disclosure.
[0097] As illustrated in FIG. 7A, transcript 700 is a ground truth transcript (e.g., as designated by training set generator 512 described above). In some embodiments, training set generator 512 can perform one or more distortion operations to transcript 700 to remove one or more words or phrases from the ground truth transcript 700. As illustrated in FIG. 7A, one or more distortion operations have been applied to transcript 700 to cause the distorted translation content 702 to be missing words and punctuation that was included in the ground truth transcript 700. For instance, the ground truth transcript 700 includes the phrase “Where have you been? So much has happened since I last saw you. I lost my hammer,” while the distorted transcript 702 includes the phrase “Where have you been much has happened since I last saw I lost my.”
[0098] In other or similar embodiments, training set generator 512 can perform one or more distortion operations to paraphrase or elongate words or phrases of the ground truth transcript. Paraphrased words or phrases may have a shorter speech duration than the original words or phrases of the ground truth transcript 700, while elongated words or phrases may have a longer speech duration than the original words or phrases of the ground truth transcript 700. As illustrated by FIG. 7B, training set generator 512 can perform one or more distortion operationsto elongate the words or phrases of the ground truth transcript 700. For example, “So much has happened since I last saw you” of transcript 700 can be elongated to “So many things have happened since I last saw you,” as illustrated by transcript 704. Further, “so that’s pretty fresh” of transcript 700 can be elongated to “which is still pretty fresh in my mind” in transcript 704. As also illustrated by FIG. 7B, training set generator 512 can perform one or more distortion operations to paraphrase the words or phrases of the ground truth transcript 700. For example, “So much has happened since I last saw you. I lost my hammer. Like, yesterday, so that’s still pretty fresh” of transcript 700 can be translated to “So much has happened since I last saw you, like I lost my hammer. That’s still fresh,” as illustrated by transcript 706.
[0099] In yet other or similar embodiments, training set generator 512 can perform one or more distortion operations that apply round-trip translation to the words or phrases of the ground tmth transcript 700. For example, training set generator 512 can perform one or more operations to provide the ground truth transcript 700 as initial content to a first translation engine 220A. The first translation engine 220A can translate the initial content 700 to translated content 708 in a second language. Training set generator 512 can provide the translated content 708 in the second language (e.g., obtained from one or more outputs of the first translation engine 220A) as an input to a second translation engine 220B that can translate the translated content 708 from the second language back to the first language. The one or more outputs of the second translation engine 220B can include one or more translation errors applied by the first translation engine 220A and / or the second translation engine 220B, and therefore, the output of the second translation engine 220B can include degraded initial content 710 in the first language.
[0100] Training set generator 512 can generate the distorted translation content according to other techniques, in other or similar embodiments. For example, training set generator 512 can obtain an audio file including the audio portion 252 of the subsequent set of video frames and can update the audio file to include one or more additional audio signals, which include audio noise, background sounds, additional dialogue or narration, and so forth. In other or similar embodiments, training set generator 512 can remove a portion (e.g., a chunk) of the audio file that includes one or more words or phrases spoke by speakers of the audio portion 252. In yet other or similar embodiments, training set generator 512 can apply one or more sound separation operations to the audio file to remove the background noise and / or can introduce background noise and subsequently apply the one or more sound separation techniques. Training set generator 512 can provide the updated audio file as input to one or more transcription engine(s) 522 and can obtain a transcript of the updated audio file based onone or more outputs of the transcription engine(s) 522. The obtained transcript can be a distorted version of the ground truth transcript due to the updates made to the original audio file as described above.
[0101] In other or similar embodiments, training set generator 512 can generated the distorted transcript according to other techniques, including but not limited to, removing the beginning or end of a sentence of the ground truth transcript, including a portion of a previous utterance or the next utterance before and / or after a word or phrase of the audio portion 252, replacing one or more words or phrases of the ground truth transcript with synonyms or homonyms of the words or phrases and so forth.
[0102] In some embodiments, training set generator 512 can apply one or more types of distortion operations to the ground truth transcript to generate the distorted transcript, as described above. For instance, training set generator 512 can apply multiple types of distortion operations to the ground truth transcript to generate a single distorted transcript. In other or similar embodiments, training set generator 512 may apply multiple types of distortion operations to the ground truth transcript to generate multiple distorted transcripts (e.g., each corresponding to a respective operation type). In such embodiments, training set generator 512 may generate a subset of training data (e.g., including a respective training input and corresponding training output) for each generated distorted transcript, as described herein.
[0103] Referring back to FIG. 6, at block 610, processing logic generates training data to train an Al model to generate a predicted audio portion of a video frame of an additional video content item in the first language in view of additional contextual information for one or more prior video frames of the additional video content item. The training data can include (i) training input data including the contextual information for the one or more initial video frames and the distorted translation content, and (ii) training output data including the translation content. In some embodiments, the training data can include an input / output mapping, where the input includes the contextual information of the one or more initial video frames and the distorted translation content (e.g., the distorted transcript) obtained of the subsequent video frames, and the output includes the ground truth transcript. Upon generating the training data, training set generator 512 can include the generated training data in a training data set, which includes training data for one or more video frame sets of one or more training video content items. In some embodiments, the training data for one or more video frame sets may not include distorted translation content and / or the distorted translation content may match, or substantially match, the ground truth transcript for such frame sets.
[0104] Referring back to FIG. 5, as described above, training set generator 512 can generate training data for training a video content analysis model 260B and / or a speech component Al model 260C, in some embodiments. In some embodiments, the training data for the video content analysis model 260B can include a video frame depicting and environment, one or more objects in the environment, and / or one or more motions or actions performed with respect to the one or more objects and labels indicating the environment, the one or more objects, and the one or more motions or actions. In some embodiments, the training data for the speech component Al model 260C can include a video frame depicting a facial movement of a speaker and a label indicating a speech component (e.g., a viseme) associated with the facial movement.
[0105] Training engine 522 can train Al model(s) 260 using the training data (e.g., training set T) from training set generator 512, as described above. The Al model (s) 260 can refer to the model artifact that is created by the training engine 522 using the training data that includes training inputs and / or corresponding target outputs (correct answers for respective training inputs). The training engine 522 can find patterns in the training data that map the training input to the target output (the answer to be predicted), and provide the Al model(s) 260 that captures these patterns. The Al model(s) 260 can be composed of, e.g., a single level of linear or nonlinear operations (e.g., a support vector machine (SVM or may be a deep network, i.e., Al model that is composed of multiple levels of non-linear operations). An example of a deep network is a neural network with one or more hidden layers, and such an Al model may be trained by, for example, adjusting weights of a neural network in accordance with a backpropagation learning algorithm or the like. In other or similar embodiments, Al model(s) 260 can include one or more large language models (LLMs). Such Al model(s) 260 can be trained, for example, according to LLM and / or generative Al training techniques.
[0106] Validation engine 524 may be capable of validating a trained Al model(s) 260 using a corresponding set of features of a validation set from training set generator 512. The validation engine 524 may determine an accuracy of each of the trained Al model(s) 260 based on the corresponding sets of features of the validation set. The validation engine 524 may discard a trained Al model(s) 260 that has an accuracy that does not meet a threshold accuracy. In some embodiments, the selection engine 526 may be capable of selecting a trained Al model 260 that has an accuracy that meets a threshold accuracy. In some embodiments, the selection engine 526 may be capable of selecting the trained Al model 260 that has the highest accuracy of the trained Al models 560.
[0107] The testing engine 528 may be capable of testing a trained Al model 260 using a corresponding set of features of a testing set from training set generator 512. For example, a first trained Al model 260 that was trained using a first set of features of the training set may be tested using the first set of features of the testing set. The testing engine 528 may determine a trained Al model 260 that has the highest accuracy of all of the trained Al models based on the testing sets
[0108] Predictive component 552 of server machine 550 may be configured to feed data as input to model 560 and obtain one or more outputs. As described above, a predictive component 552 residing at server machine 550 can be included at or otherwise accessible to encryption engine 152. Predictive component 552 can provide data as input to model 260 and obtain one or more outputs. In some embodiments, predictive component 552 can include voice dubbing engine 152. Voice dubbing engine 152 can provide input data that includes translation content for an audio portion of one or more video frames of a video content item (e.g., video content 122), as described above.
[0109] In accordance with embodiments described herein, predictive component 552 can provide a prompt as input to one or more of model(s) 260. For example, predictive component 552 can provide a prompt as an input to voice dubbing Al model 260A that includes a request for the audio portion of one or more video frames of video content 122 in a first language, a reference to the contextual information 256 for one or more prior video frames of the video content 122, and / or a reference to the translation content 258 representing the audio portion of the one or video frames.
[0110] In some embodiments, one or more portions of model(s) 260 can reside at a client device 102 (e.g., owner device 102A). In such embodiments, owner device 102A may obtain the private key based on one or more outputs of model 560, as described above. The portion of model(s) 260 residing at a client device 102 may be discrete and / or remote form other portions of the model(s) 260 residing at other client devices 102 and / or a memory of predictive system 180. In such embodiments, data that is provided as an input and / or obtained as an output of model(s) 260 may not be provided to predictive system 180 and / or other client devices 102. Accordingly, private keys generated for a respective client device 102 may remain at the client device 102 and therefore are not accessible to any other component of system 100.
[0111] FIG. 8 is a block diagram illustrating an exemplary computer system 800, in accordance with implementations of the present disclosure. The computer system 800 can be the server machine 130-140 or client devices 102A-N in FIG. 1. The machine can operate in the capacity of a server or an endpoint machine in endpoint-server network environment, or asa peer machine in a peer-to-peer (or distributed) network environment. The machine can be a television, a personal computer (PC), a tablet PC, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0112] The example computer system 800 includes a processing device (processor) 802, a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), double data rate (DDR SDRAM), or DRAM (RDRAM), etc.), a static memory 806 (e.g., flash memory, static random access memory (SRAM), etc ), and a data storage device 818, which communicate with each other via a bus 840.
[0113] Processor (processing device) 802 represents one or more general-purpose processing devices such as a microprocessor, central processing unit, or the like. More particularly, the processor 802 can be a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or processors implementing a combination of instruction sets. The processor 802 can also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. The processor 802 is configured to execute instructions 805 for performing the operations discussed herein.
[0114] The computer system 800 can further include a network interface device 808. The computer system 800 also can include a video display unit 810 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an input device 812 (e.g., a keyboard, and alphanumeric keyboard, a motion sensing input device, touch screen), a cursor control device 814 (e.g., a mouse), and a signal generation device 820 (e.g., a speaker).
[0115] The data storage device 818 can include a non-transitory machine-readable storage medium 824 (also computer-readable storage medium) on which is stored one or more sets of instructions 805 embodying any one or more of the methodologies or functions described herein. The instructions can also reside, completely or at least partially, within the main memory 804 and / or within the processor 802 during execution thereof by the computer system800, the main memory 804 and the processor 802 also constituting machine-readable storage media. The instructions can further be transmitted or received over a network 830 via the network interface device 808.
[0116] In one implementation, the instructions 805 include instructions for predicting channel lineup viewership. While the computer-readable storage medium 824 (machine- readable storage medium) is shown in an exemplary implementation to be a single medium, the terms “computer-readable storage medium” and “machine-readable storage medium” should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store the one or more sets of instructions. The terms “computer-readable storage medium” and “machine-readable storage medium” shall also be taken to include any medium that is capable of storing, encoding or carrying a set of instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present disclosure. The terms “computer-readable storage medium” and “machine-readable storage medium” shall accordingly be taken to include, but not be limited to, solid-state memories, optical media, and magnetic media.
[0117] Reference throughout this specification to “one implementation,” or “an implementation,” means that a particular feature, structure, or characteristic described in connection with the implementation is included in at least one implementation. Thus, the appearances of the phrase “in one implementation,” or “in an implementation,” in various places throughout this specification can, but are not necessarily, referring to the same implementation, depending on the circumstances. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more implementations.
[0118] To the extent that the terms “includes,” “including,” “has,” “contains,” variants thereof, and other similar words are used in either the detailed description or the claims, these terms are intended to be inclusive in a manner similar to the term “comprising” as an open transition word without precluding any additional or other elements.
[0119] As used in this application, the terms “component,” “module,” “system,” or the like are generally intended to refer to a computer-related entity, either hardware (e.g., a circuit), software, a combination of hardware and software, or an entity related to an operational machine with one or more specific functionalities. For example, a component may be, but is not limited to being, a process running on a processor (e g., digital signal processor), a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a controller and the controller can be acomponent. One or more components may reside within a process and / or thread of execution and a component may be localized on one computer and / or distributed between two or more computers. Further, a “device” can come in the form of specially designed hardware; generalized hardware made specialized by the execution of software thereon that enables hardware to perform specific functions (e.g., generating interest points and / or descriptors); software on a computer readable medium; or a combination thereof.
[0120] The aforementioned systems, circuits, modules, and so on have been described with respect to interact between several components and / or blocks. It can be appreciated that such systems, circuits, components, blocks, and so forth can include those components or specified sub-components, some of the specified components or sub-components, and / or additional components, and according to various permutations and combinations of the foregoing. Subcomponents can also be implemented as components communicatively coupled to other components rather than included within parent components (hierarchical). Additionally, it should be noted that one or more components may be combined into a single component providing aggregate functionality or divided into several separate sub-components, and any one or more middle layers, such as a management layer, may be provided to communicatively couple to such sub-components in order to provide integrated functionality. Any components described herein may also interact with one or more other components not specifically described herein but known by those of skill in the art.
[0121] Moreover, the words “example” or “exemplary” are used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the words “example” or “exemplary” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X employs A or B” is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then “X employs A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form.
[0122] Finally, implementations described herein include collection of data describing a user and / or activities of a user. In one implementation, such data is only collected upon the user providing consent to the collection of this data. In some implementations, a user is prompted to explicitly allow data collection. Further, the user may opt-in or opt-out of participating insuch data collection activities. In one implementation, the collect data is anonymized prior to performing any analysis to obtain any statistical patterns so that the identity of the user cannot be determined from the collected data.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: identifying a video content item to be translated to a first language from a second language, the video content item comprising a plurality of video frames that each include one or more audio portions and one or more video portions; providing, as input to an artificial intelligence (Al) model: contextual information for one or more initial video frames of the plurality of video frames, wherein the contextual information pertains to an audio portion of the one or more initial video frames and a video portion of the one or more initial video frames, and translation content representing an audio portion of one or more subsequent video frames of the plurality of video frames; and extracting, from one or more outputs of the Al model, the audio portion of the one or more subsequent video frames in the first language, wherein a speech component of the extracted audio portion in the first language corresponds to a speech component of the audio portion in the second language in view of the provided contextual information.
2. The method of claim 1, wherein the contextual information comprises a transcript for the audio portion of the one or more initial video frames.
3. The method of claim 1, wherein the contextual information comprises one or more of: an identifier of each speaker associated with the audio portion of the one or more initial video frames, an identifier of a gender for each speaker associated with the audio portion of the one or more initial video frames, an indication of at least one of an environment depicted by the video portion of the one or more initial video frames or one or more objects included in the environment, an indication of one or more motions or actions pertaining to a respective object depicted by the video portion, an indication of one or more additional sound events associated with the audio portion of the one or more initial video frames, ora content summary comprising a summarized description of at least a portion of content of the one or more initial video frames.
4. The method of claim 3, further comprising obtaining the contextual information by: providing at least the video portion of the one or more initial video frames as input to a video content analysis model that is trained to detect, based on given video content, at least one of an object depicted by the video content, an environment depicted by the video content, or a motion or action performed with respect to an object depicted by the video content; and extracting, based on one or more outputs of the video content analysis model, a set of labels pertaining to provided video portion, the set of labels comprising at least one of the identifier of the gender for each speaker associated with the audio portion of the one or more initial video frames, the indication of at least one of the environment or the objects included in the environment depicted by the video portion of the one or more initial video frames, or the indication of the one or more motions or actions pertaining to the respective object depicted by the video portion.
5. The method of claim 1, wherein providing the contextual information for the one or more initial video frames of the plurality of video frames as input to an Al model comprises at least one of: providing an audio file comprising the audio portion of the one or more initial video frames as input to the Al model, or providing a video file comprising the video portion of the one or more initial video frames as input to the Al mode.
6. The method of claim 1, further comprising: determining one or more speech components for the audio portion of the one or more subsequent video frames according to the second language; and providing the determined one or more speech components as additional input to the Al model, wherein the speech component of the audio portion in the second language corresponding to the speech component of the extracted audio portion in the first language is indicated by the determined one or more speech components.
7. The method of claim 1, wherein the speech component of the extracted audio portion in the first language and the speech component of the audio portion in the second language comprise one or more of a speech duration, a sequence of lip shapes, or a speech structure.
8. The method of claim 1, wherein the contextual information further comprises at least one of an indication of the second language or an indication of a content type or a content genre associated with the video content item.
9. The method of claim 1, wherein providing the contextual information and the translation content as the input to the Al model comprises: providing a prompt to the Al model, wherein the prompt comprises: a request for the audio portion of the one or more subsequent video frames in the first language, a reference to the contextual information for the one or more initial frames, and a reference to the translation content representing the audio portion of the one or more subsequent frames.
10. The method of claim 1, wherein the translation content comprises an initial translation of the audio portion to the first language from the second language, and wherein audio portion of the one or more subsequent video frames extracted from one or more outputs of the Al model comprises an improved translation of the audio portion to the first language from the second language in view of the given contextual information.
11. The method of claim 1, further comprising: responsive to a request from a user for the video content item in the first language, providing, to a client device associated with the user, the extracted audio portion of the one or more subsequent video frames and the video portion of the one or more subsequent video frames for presentation to the user via a user interface (UI) of the client device.
12. The method of claim 1, further comprising: providing, as input to the Al model: contextual information for the one or more initial video frames, and translation content representing an audio portion of the one or more initial video frames; andextracting, from one or more additional outputs of the Al model, the audio portion of the one or more initial video frames in the first language, wherein the speech component of the extracted audio portion of the one or more initial video frames corresponds to the speech component of the audio portion of the one or more initial video frames in the second language.
13. A system comprising: a memory; and a set of one or more processing devices coupled to the memory, wherein the set of one or more processing devices is to perform operations comprising: identifying a training video content item comprising a plurality of video frames that each include one or more audio portions and one or more video portions; obtaining contextual information for one or more initial video frames of the plurality of video frames, wherein the contextual information pertains to an audio portion of the one or more initial video frames and a video portion of the one or more initial video frames; obtaining translation content representing a training audio portion of a subsequent video frame of the plurality of video frames, wherein the training audio portion is translated to a first language from a second language; performing one or more distortion operations to the translation content to generate distorted translation content comprising one or more translation errors in the first language; and generating training data to train an artificial intelligence (Al) model to generate a predicted audio portion of a video frame of an additional video content item in the first language in view of additional contextual information for one or more prior video frames of the additional video content item, wherein the training data comprises (i) training input data comprising the contextual information for the one or more initial video frames and the distorted translation content, and (ii) training output data comprising the translation content.
14. The system of claim 13, wherein performing the one or more distortion operations to the translation content comprises: removing at least one of a word or a phrase from the translation content.
15. The system of claim 13, wherein performing the one or more distortion operations to the translation content comprises: providing the translation content as an input to first translation engine that is configured to translate content in the first language to content in the second language; providing an output of the first translation engine as an input to a second translation engine, wherein the second translation engine is configured to translate content in the second language to content in the first language; and obtaining one or more outputs of the second translation engine, wherein the one or more outputs comprise the distorted translation content.
16. The system of claim 15, wherein the first translation engine or the second translation engine comprise one or more of a speech-to-speech (S2S) translation engine that is configured to translate an audio signal from an initial language to an audio signal in a target language, a speech-to-text (S2T) translation engine that is configured to translate an audio signal in an initial language to a text string in a target language, or a text-to-text (T2T) translation engine that is configured to translate a text string in an initial language to a text string in a target language.
17. The system of claim 13, wherein performing the one or more distortion operations to the translation content comprises: determining for one or more phrases of the translation content, at least one of a paraphrased version of the one or more phrases or an elongated version of the one or more phrases, wherein a speech duration of the paraphrased version is shorter than a speech duration of the one or more phrases, and wherein a speech duration of the elongated version is longer than the speech duration of the one or more phrases, wherein the distorted translation content includes the at least one of the paraphrased version of the one or more phrases or the elongated version of the one or more phrases.
18. The system of claim 13, wherein performing the one or more distortion operations to the translation content comprises: obtaining an audio file comprising the training audio portion of the subsequent video frame;updating the audio file to include one or more additional audio signals, wherein the one or more additional audio signals comprise one or more of audio noise, background sounds, or additional dialogue or narration; and providing the updated audio file as an input to a transcription engine that is configured to generate a set of text strings based on one or more audio signals of a provided audio file, wherein the distorted translation content comprises a generated set of text strings included in one or more outputs of the transcription engine.
19. The system of claim 13, wherein the contextual information for the one or more initial frames comprises at least one of: a transcript for the audio portion of the one or more initial video frames, an identifier of each speaker associated with the audio portion of the one or more initial video frames, an identifier of a gender for each speaker associated with the audio portion of the one or more initial video frames, an indication of at least one of an environment depicted by the video portion of the one or more initial video frames or one or more objects included in the environment, an indication of one or more motions or actions pertaining to a respective object depicted by the video portion, an indication of one or more speech components for the audio portion of the one or more initial video frames, the one or more speech components comprising at least one of a speech duration, a sequence of lip shapes, or a speech structure of the audio portion, an indication of one or more additional sound events associated with the audio portion of the one or more initial video frames, or a content summary comprising a summarized description of at least a portion of content of the one or more initial video frames.
20. A non-transitory computer readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform operations comprising: identifying a video content item for translation to a first language from a second language, the video content item comprising a plurality of video frames that each include one or more audio portions and one or more video portions; providing, as input to an artificial intelligence (Al) model:contextual information for one or more initial video frames of the plurality of video frames, wherein the contextual information pertains an audio portion of the one or more initial video frames and a video portion of the one or more initial video frames, and translation content representing an audio portion of one or more subsequent video frames of the plurality of frames; and extracting, from one or more outputs of the Al model, the audio portion of the one or more subsequent video frames in the first language for the video content item, wherein a speech component of the extracted audio portion in the first language corresponds to a speech component of the audio portion in the second language in view of the provided contextual information.
21. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of claim 1.
Citation Information
Patent Citations
Audio and video translator
US20230088322A1