Generative text replacement in videos
The generative neural network system addresses the inefficiencies of conventional text modification in videos by applying context-aware text edits efficiently, using fine-tuned neural networks and human feedback to produce high-quality, resource-saving outputs.
Patent Information
- Application Number
- PCT/US2024/018191
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-09-04
AI Technical Summary
Conventional techniques for modifying text in videos often result in low-quality outputs due to handling text independently of the surrounding visual context, requiring manual user input and consuming excessive computational resources, and are inefficient in applying seamless and natural text modifications.
A system utilizing a generative neural network fine-tuned on a training dataset for text-editing tasks, which processes input videos and queries to apply modifications like text replacement or removal while preserving the visual context, using discriminative and generative operations to determine and modify text pixels without requiring additional user input.
The system generates modified videos with natural and seamless text modifications, saving computational resources by focusing on impacted pixels and incorporating human feedback through reward models, resulting in high-quality outputs that preserve the underlying visual content and text appearance.
Smart Images

Figure US2024018191_04092025_PF_FP_ABST
Abstract
Description
[0001] GENERATIVE TEXT REPLACEMENT IN VIDEOS
[0002] BACKGROUND
[0003] This specification relates to generating outputs conditioned on inputs using machine learning models.
[0004] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.
[0005] Some machine learning models are deep models that employ multiple layers of operations to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
[0006] SUMMARY
[0007] This specification describes a system implemented as computer programs on one or more computers in one or more locations that generates a modified video conditioned on an input video and an input query that specifies a text-editing task to be performed on the input video.
[0008] According to a first aspect there is provided a method comprising receiving an input video; receiving an input query specifying a text-editing task to be performed on the input video, the text-editing task requiring a modification to text displayed in one or more video frames of the input video; and processing, using a generative neural network, the input video and the input query to generate an output comprising a modified video, with the modification applied to the text displayed in the one or more video frames.
[0009] In some implementations, receiving an input video comprises receiving the input video from a user.
[0010] In some implementations, receiving an input query comprises receiving the input query from a user.
[0011] In some implementations, the output further comprises a status indicator for the modification.
[0012] In some implementations, the status indicator indicates whether the modification was successfully applied, or a location of the text displayed in one or more video frames of the input video. In some implementations, the method further comprises providing the modified video for display to a user.
[0013] In some implementations, the text displayed in one or more video frames of the input video comprises any one or more of: overlay text, in-scene text, or animated text.
[0014] In some implementations, the text-editing task comprises any one or more of: modifying one or more visual features of the text, removing the text, or replacing the text with alternative text that shares visual features with the text.
[0015] In some implementations, the input query specifies i) text to be replaced and ii) alternative text.
[0016] In some implementations, the input query specifies a different color for the text.
[0017] In some implementations, the input query specifies i) text to be displayed in a different color and ii) the different color.
[0018] In some implementations, the input query specifies text to be removed.
[0019] In some implementations, the input query7specifies i) text to be displayed in a different size and ii) a definition representing a change in size for the specified text.
[0020] In some implementations, the generative neural network has been pre-trained to generate at least a video conditioned on a video and text.
[0021] In some implementations, the generative neural network has been fine-tuned on a training dataset for a plurality of text-editing tasks that each require modifying pixels that depict text in input videos.
[0022] In some implementations, the generative neural network has been fine-tuned sequentially on each of the plurality of text-editing tasks of the training dataset.
[0023] In some implementations, the training dataset comprises a plurality of training examples, each comprising a training input comprising a training video and a training query specifying a text-editing task from the plurality of text-editing tasks, and a corresponding training output comprising at least an output video.
[0024] In some implementations, the generative neural network has been further fine-tuned to minimize a loss function based on an aggregate reward value.
[0025] In some implementations, the aggregate reward value is derived from one or more reward values, wherein each reward value is generated by a corresponding rew ard model.
[0026] In some implementations, each corresponding reward model is configured to generate a reward value representing an assessment of a given video according to a corresponding comparison prompt and an original video. In some implementations, the corresponding comparison prompt comprises a question about a visual similarity, a layout, a style, a font, a weight, or a color of text displayed in one or more video frames of the given video compared to text displayed in one or more video frames of the original video.
[0027] In some implementations, each corresponding reward model is trained to generate the reward value for the corresponding comparison prompt.
[0028] According to a further aspect, there is also described a system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the methods described herein.
[0029] According to a further aspect, there is also described one or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the methods described herein.
[0030] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0031] The system described in this specification can perform a variety of text-editing tasks for onscreen text replacement or modification. For example, the system can receive an input video and an input query' that specifies a text-editing task to be performed on the input video. The text-editing task requires a modification to text displayed in one or more video frames of the input video. The system can generate a modified video with the modification required by the text-editing task by processing the input video and the input query using a generative neural network. The system can generate a modified video with modified text using the generative neural network while preserving the underlying visual content of the input video, as well as preserving the look and feel of the text displayed in the input video relative to the underlying visual content of the input video. The text can be, for example, overlay text, inscene text, and / or animated text. The system can provide for seamless and natural modifications of different types of text displayed in videos.
[0032] For example, the system can perform text-editing tasks such as erasing text. The system can provide the input video and the input query to a generative neural network that has been fine-tuned on a training dataset for text-editing tasks that includes training examples that demonstrate erasing text. For example, the input query' can identify the text of the input video to be erased. The system can identify the pixels representing the identified text, and perform in-filling or inpainting for the pixels with the content of the underlying visual content or the background. The underlying visual content or background may be potentially non- uniform across video frames. The output video can thus include a reconstruction of the pixels of the visual content of the input video that are exposed as a result of erasing the pixels of the text displayed in the input video.
[0033] As another example, the system can perform text-editing tasks such as onscreen text replacement. The system can provide the input video and the input query to a generative neural network that has been fine-tuned on a training dataset for text-editing tasks that includes training examples that demonstrate replacing text. The input query can identify the text of the input video to be replaced, and the alternative text with which to replace the text of the input video.
[0034] The system can identify the original text to be replaced, perform in-filling or inpainting to erase the pixels of the original text, and generate pixels representing the alternative text in the output video. The output video can display the alternative text, and the alternative text can share visual features such as font, color, shading, position, and / or style with the text displayed in the input video.
[0035] In some examples, the system may change the layout of the text relative to the underlying visual content to accommodate shorter or longer alternative text. The system can also apply visual features such as geometry or lighting distortions of the original text to the alternative text.
[0036] In some examples, the system can perform onscreen text replacement to localize text displayed in videos and images. For example, the system can replace text that is displayed in English in the input video so that the modified video displays text in Spanish. The input query’ can identify the English text to be replaced, and the alternative Spanish text with which to replace the English text. The sy stem can recreate the text style of the original text into a potentially different alphabet with different characters. Onscreen text replacement can also be used to correct mistakes in text, or alter sensitive content, of input videos.
[0037] In some examples, the system can perform modifications of animated text by animating movement and transformation patterns over time across video frames according to the movement and transformation patterns of the input video. The system can provide the input video and the input query to a generative neural network that has been fine-tuned on a training dataset for text-editing tasks that includes training examples that display animated text. The input query' can identify the text of the input video to be modified. The system can identify the pixels representing the identified text, and perform onscreen text replacement while retaining features of the animation of the identified text in the input video. Conventional techniques for modifying text displayed in videos may not result in output videos with seamless and natural modifications of text. For example, some conventional techniques may handle text independently of the surrounding visual context or the underlying visual content. For example, when replacing text displayed within a street sign with alternative text, the alternative text may be longer than the original text. Conventional techniques may ignore the bounds of the street sign, resulting in a video of low visual qualify’ and coherence.
[0038] The system described in this specification can replace the original text with the alternative text while preserving the look and feel of the original text relative to the underlying visual content. For example, the system can replace the original text with the alternative text, with the original font style and color, but in a different size and / or in a different layout so that the alternative text is confined within the bounds of the street sign. The system can use a generative neural network that has learned, through fine-tuning on a training dataset for text-editing tasks, to enforce visual and spatial constraints such as displaying text within a street sign or displaying text within the video frame.
[0039] Conventional techniques for modifying text displayed in videos have limitations which make them difficult or impractical to use. For example, some techniques require the user to manually identify and select pixels representing text displayed in video frames, and to manually fill in the pixels representing text, or modify the pixels representing text, or modify- other pixels to represent text.
[0040] The system can apply different types of modifications specified by an input query to an input video without further user input. The system can determine which pixels represent text, and which subset of pixels represent the text specified in the input query, without user input. The system can also determine which pixels to modify to apply the modification specified in the input query, such as which pixels to in-fill, or which pixels to modify- to represent text, without user input. The system does not require any further inputs other than the input query and the input video, resulting in a more efficient user experience.
[0041] Some conventional techniques require generating an entirely new video that includes modified text, which may require the user to manually describe all aspects of the desired video, and may require more computational resources to generate. Other techniques require a sequence of machine learning models that may result in lower qualify' output videos.
[0042] The system described in this specification can apply modifications specified by an input query to an input video by using a generative neural network that performs discriminative and generative operations. For example, the generative neural network can have been fine-tuned to perform multiple text-editing tasks that require discriminating pixels that include text and / or generating output videos with modifications to the pixels that include text. The system can thus generate a video with modifications applied to the text without generating an entirely new video based on a text prompt that describes the video and the text, saving computing time and resources. The system can further save computing time and resources by focusing computation on pixels that are impacted by the modification of the text-editing task. The system can also generate the video with modifications applied to the text using a single generative neural network, resulting in a more seamless modified video than a modified video generated using a sequence of machine learning models.
[0043] Furthermore, in some implementations, the system described in this specification can generate modified videos with natural and seamless modifications based on human feedback. For example, the generative neural network can be fine-tuned by a training system to generate modified videos that represent human preferences for multiple aspects, such as layout, sty le, color, or other visual features. The generative neural network can be fine-tuned to minimize a loss function based on reward values. The reward values can be generated by reward models that are trained on human feedback. Generating reward values from reward models can be more efficient than obtaining reward values from human raters. In addition, the reward models can predict human preference nuances that may not be represented in the training dataset. For example, particularly for in-scene text, it may be computationally expensive to generate a large number of high-quality training outputs to be used in training examples. For example, each training example includes a training video with in-scene text with properties such as lighting, orientation, skew deformation, and obstruction. The training output for each training example would need to have the modification specified by a training query applied, while retaining the properties of the in-scene text displayed in the training video, which may be computationally expensive to generate. Generating a high- quality training output may also be infeasible. Furthermore, the training dataset may not include exhaustive examples of complex properties and / or animation patterns. Thus, the system can use reward models to predict human preferences for the outputs of the generative neural network, and thus improve the performance of the generative neural network without needing high-quality training outputs for the training dataset.
[0044] The training system can also fine-tune the generative neural network in a computationally efficient manner. For example, the training system can fine-tune a pretrained generative neural network on a training dataset for text-editing tasks. The generative neural network can have been pre-trained to generate a video conditioned on a video and text. The training system can also generate the reward values for further fine-tuning based on human feedback in a computationally efficient manner. For example, the training system can derive the reward models from the fine-tuned generative neural network.
[0045] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0046] BRIEF DESCRIPTION OF THE DRAWINGS
[0047] FIG. 1 is a diagram of an example process for generating a modified video.
[0048] FIG. 2 is a diagram of an example system for generating a modified video.
[0049] FIGS. 3A-3C show example frames of input videos and modified videos for example text-editing tasks.
[0050] FIG. 4 is a flow diagram of an example process for generating a modified video.
[0051] FIG. 5 is a flow diagram of an example process for fine-tuning a generative neural network.
[0052] FIG. 6 shows an example presentation of a user interface for collecting human feedback.
[0053] FIG. 7 is a diagram of an example process for fine-tuning a generative neural network.
[0054] Like reference numbers and designations in the various drawings indicate like elements.
[0055] DETAILED DESCRIPTION
[0056] FIG. 1 is a diagram of an example process 100 for generating a modified video. For convenience, the process 100 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system for generating a modified video, e.g.. the system 200 of FIG. 2, appropriately programmed in accordance with this specification, can perform the process 100.
[0057] The system 200 can receive an input video 104. The input video 104 can display original text as well as other underlying visual content such as a scene or a background. FIG. 1 shows a video frame of the example input video 104. In the example of FIG. 1, the input video 104 displays the text “Foundation models are cool!” The input video 104 can display the text with certain visual features, such as style, color, font, etc. In some examples, the input video 104 can display no text, additional text, or other text.
[0058] The system 200 can receive an input query 106. The input query 106 can specify the ty pe of text-editing task and / or the text of the input video 104 to be modified. In the example of FIG. 1, the input query 106 includes “Replace the text in the video: “Foundation models are cool!'’ with the text: “jLos modelos de base son geniales!'”' The input query 106 specifies the type of text-editing task, text replacement. The input query 106 also specifies the text of the input video 104 to be modified, “Foundation models are cool!” The input query 106 also specifies alternative text, “jLos modelos de base son geniales!” In some examples, the input query 106 can specify other types of text-editing tasks as described below with reference to FIGS. 3A-3C.
[0059] The system 200 can generate the modified video 180 and the status indicator 112 as described with reference to FIG. 2. FIG. 1 shows a video frame of the example modified video 180. In the example of FIG. 1, the system 200 can apply the modification of the input query 106 to the modified video 180. The modified video 180 of FIG. 1 thus displays “jLos modelos de base son geniales!” in the place of “Foundation models are cool!” The modified video 180 can display the alternative text with the same visual features as the text of the input video 104. In some examples, because the alternative text includes a different number of characters, a different number of words, or words with different numbers of characters, the modified video 180 may display the alternative text in a different layout or size compared to the original text of the input video 104.
[0060] The status indicator 112 can indicate whether the modification was successfully applied. In the example of FIG. 2. the status indicator 112 can include text such as “Text found.” In some examples, the input video 104 may not include the text “Foundation models are cool!” The status indicator 1 12 may include text such as “Text not found.”
[0061] The status indicator 112 can also indicate a location of the text displayed in one or more video frames of the input video 104. For example, the location of the text can be represented by a bounding box. In the example of FIG. 2, the status indicator 112 can include text such as “Text at: (10, 20) to (100. 400).” In some examples, the location of the text may change between video frames of the input video 104. The location of the text can be represented by the bounding box for the first video frame of the input video 104 in which the text is displayed. As another example, the location of the text can be represented by data representing a fine-grained mask of pixels that display the text. As described below with reference to FIG. 5, a training system such as the training system 140 described with reference to FIG. 1 can use the status indicator 112 for each training example as part of calculating a loss value for the training example.
[0062] In some examples, the status indicator 112 can also indicate a time that the text is displayed in one or more video frames of the input video. For example, the time that the text is displayed can be represented by a timestamp that indicates the time elapsed from the beginning of the input video to the point that text was displayed. In the example of FIG. I, the status indicator 112 can include text such as “Text at 9:59 seconds.” In some examples, the text may be displayed over a time interval of the input video 104. The time that the text is displayed can be represented by the timestamp for the first video frame of the input video 104 in which the text is displayed.
[0063] FIG. 2 is a diagram of an example system 200 for generating a modified video 280 that includes multiple video frames. The system 200 includes multiple components, such as a generative neural network 210. The system 200 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0064] To generate the modified video 280, the system 200 receives an input video 204 and an input query 206. In some implementations, the system 200 can receive the input video 204 and / or the input query 206 from a user.
[0065] The input video 204 includes multiple video frames. Each video frame is an image that includes multiple pixels that each has one or more intensity values. Each video frame can display visual context, such as a scene or a background.
[0066] Some of the multiple video frames can also display text. The text can be overlay text, in-scene text, and / or animated text.
[0067] Overlay text can include text that was added to the video frame that was not part of the scene. For example, the overlay text can have been added to the video frame after the video of the scene was recorded or generated. As an example, the video frame can display the scene with the overlay text displayed over the scene. As another example, the video frame can display the scene with the overlay text in a split-screen. For example, the video frame can display the text next to or around the scene. Examples of overlay text are shown below in FIGS. 3A-3B and FIG. 4.
[0068] In-scene text can include text that appears within the scene. For example, the text can be displayed in the environment in which the video was recorded. An example of in-scene text is shown below in FIG. 3A. Animated text can include text that moves and / or transforms between video frames. For example, animated text can have been added to the video frames after the video was recorded or generated. Examples of transformations for animated text include overlay text or in-scene text that changes position, size, appearance, color, transparency, and / or orientation across multiple video frames.
[0069] The input query’ 206 specifies a text-editing task to be performed on the input video 204. For example, the input query 206 can specify the type of text-editing task. In some examples, the input query 206 can specify the text of the input video 204 to be modified. Example text-editing tasks such as modifying visual features of the text, removing the text, and / or replacing the text with alternative text, are described below with reference to FIGS. 3A-3C.
[0070] The system 200 generates the modified video 280 and a status indicator 212 for the modified video 280 using a generative neural network 210, as described in further detail below.
[0071] The system 200 can then provide the modified video 280 and / or the status indicator 212 for display to a user, for example, through a user interface.
[0072] The status indicator 212 includes text that indicates whether the modification specified by the input query 206 was successfully applied. As an example, the system 200 may provide the input video 204 and the input query 206 as input to the generative neural network 210 to generate a modified video 280, but the generative neural network 210 may output a status indicator 212 that indicates the text specified by the input query 206 was not detected in the input video 204. For example, the status indicator 212 can indicate that the text was not found. The system 200 can provide the status indicator 212 for display to a user, for example, through a user interface. In some implementations, the system 200 may not output a modified video 280. In some implementations, the system 200 can output the input video 204 as the modified video 280.
[0073] As another example, the system 200 may provide the input video 204 and the input query 206 as input to the generative neural network 210 to generate a modified video 280, and the generative neural network 210 may output a status indicator 212 that indicates the text specified by the input query 206 was detected in the input video 204. The generative neural network 210 can also output the modified video 280 with the modification applied to the text.
[0074] For example, the status indicator 212 can indicate that the text was found. In some examples, the status indicator 212 can also indicate a location of the text displayed in one or more video frames of the input video. During training of the generative neural network 210, the training system 240 described below can use the location of the text displayed for each training example to calculate a loss for the training example. In some examples, the status indicator 212 can also indicate a time that the text is displayed in one or more video frames of the input video. For example, the status indicator 212 can include a timestamp that indicates the time elapsed from the beginning of the input video to the point that text was first displayed.
[0075] The generative neural network 210 is a generative neural network that is configured to generate at least a video conditioned on a video and text. The generative neural network 210 can have been pre-trained to generate a video conditioned on a video and text, and fine-tuned on a training dataset for multiple text-editing tasks that each require modifying pixels that depict text in input videos.
[0076] The generative neural network 210 can have any appropriate architecture that allows the generative neural network 210 to generate at least a video conditioned on a video and text. For example, the generative neural network 210 can be a diffusion model, a Transformer-based model, or an autoregressive model.
[0077] The generative neural network 210 is pre-trained, e.g., by the training system 240 or by one or more other systems. For example, the training system 240 can pre-train the generative neural network 210 on a conditional video generation task. As a particular example, the generative neural network 210 can be pre-trained to reverse a forward diffusion process on a dataset of videos and corresponding text.
[0078] A training system 240 within the system 200 or another training system can fine-tune the generative neural network 210. That is, the training system 240 can obtain the pretrained generative neural network 210 and can then fine-tune the generative neural network 210 starting from pre-trained values of the network parameters of the generative neural network 210.
[0079] The training system 240 can fine-tune the generative neural network 210 on a training dataset 244. The training dataset 244 can include multiple training examples that correspond to different text-editing tasks.
[0080] Alternatively or in addition, the training system 240 can fine-tune the generative neural network 210 to minimize a loss function based on an aggregate reward value generated by a set of reward models 242. Each of the reward models 242 can be configured to generate a reward value representing an assessment of a modified video 280 generated by the generative neural network 210 compared to the input video 204. By using a pre-trained generative neural network 210, the system 200 can generate a modified video 280 with natural and coherent modifications of text. For example, the pretrained generative neural network 210 can have learned to associate the objects depicted over video frames. For example, the input video 204 may display a bag of chips moving over multiple video frames, with distortions over reflections and other visual changes over the multiple video frames. The pre-trained generative neural network 210 can be trained to understand that the same bag of chips is moving over the video frames, rather than treating pixels for each video frame independently. By fine-tuning the pre-trained generative neural network 210, the system 200 can use the generative neural network 210 to generate a modified video 280 with modifications to the text that appear natural and fit the context of the scene.
[0081] As another example, the pre-trained generative neural network 210 can have learned to identify faces, or underlying visual content such as signs. The system 200 can fine-tune the pre-trained generative neural network 210 to follow visual constraints such as not positioning text over a character’s face, positioning text relative to underlying visual content, or ensunng that all of the text content is displayed within the video frame.
[0082] Training is described in more detail below with reference to FIGS. 5-7.
[0083] FIGS. 3A-3C show example frames 310-360 of input videos and modified videos for example text-editing tasks. For example, referring to FIG. 3A, frames 310a and 310b depict a text-editing task that is modifying a visual feature of text, in particular, modifying the color of the text. Frame 310a is an example frame of an input video that depicts a scene with six movie characters. Frame 310a displays text 311 “This is from a movie.” The text 311 is displayed in black. The text 311 is an example of overlay text that is not part of the scene.
[0084] The system can generate a modified video that includes the frame 310b by receiving an input video that includes the frame 310a and an input query, and providing the input video and the input query to a generative neural network as described above with reference to FIGS. 1-2. The input query can specify the different color for the text of the input video. In the example of the frames 310a and 310b, the input query may include “Change the color of the text in the video to gray.”
[0085] Frame 310b is an example frame of the modified video generated by the system. The frame 310b depicts the same scene as frame 310a, with the modification specified by the input query above, applied to the text 311. The frame 310b displays text 312 that has the same content as the text 311 , but is displayed in a different color than the text 311. For example, the text 312 of the example of FIG. 3A is displayed in gray. Frames 320a and 320b depict a text-editing task that is modifying a visual feature of text, in particular, modifying the color of specified text. Frame 320a is an example frame of an input video that depicts a scene with two movie characters. Frame 320a displays text 321 “Scene from a film.” The text 321 is displayed in black. The text 321 is an example of overlay text that is not part of the scene. Frame 320a also displays text 325 “STOP” on a street sign in the background. The text 325 is an example of in-scene text.
[0086] The system can generate a modified video that includes the frame 320b by receiving an input video that includes the frame 320a and an input query, and providing the input video and the input query' to a generative neural network as described above. The input query can specify the text to be displayed in a different color and the different color. In the example of the frames 320a and 320b. the input query may include “Change the color of the text in the video: “a film” to gray.” In some examples, the capitalization of the specified text of the input query' may differ from the capitalization of the text 322. For example, the input query may include “Change the color of the text in the video: “A Film” to gray.”
[0087] Frame 320b is an example frame of the modified video generated by the system. The frame 320b depicts the same scene as frame 320a. with the modification specified by the input query' above, applied to the text 321. The frame 320b displays text 322 that has the same content as the text 321, but some of the text 322 is displayed in a different color than the text 311. For example, “a film” of the text 322 is displayed in gray. The other words of the text 322, “Scene from,” are the same color as the text 321. The text 325 remains the same in frame 320b as in frame 320a.
[0088] Another example of modifying visual features of text include changing the size of the text, as discussed below with reference to frames 340a and 340b. Other examples of modifying visual features of text include changing the weight, font, or thickness of the text.
[0089] Referring to FIG. 3B, frames 330a and 330b depict a text-editing task that is erasing specified text. Frame 330a is an example frame of an input video that depicts a scene with two movie characters. Frame 330a displays text 331 “Something to delete.” The text 331 is an example of overlay text that is not part of the scene. Frame 330a also displays text 335 “Other text to ignore.” The text 335 is also an example of overlay text.
[0090] The system can generate a modified video that includes the frame 330b by receiving an input video that includes the frame 330a and an input query', and providing the input video and the input query' to a generative neural network as described above. The input query can specify the text to be erased or removed. In the example of the frames 330a and 330b, the input query may include “Erase the text in the video: “Something to delete”.” In some examples, the capitalization of the specified text of the input query may differ from the capitalization of the text 331. For example, the input query may include “Erase the text in the video: “something to delete”.”
[0091] Frame 330b is an example frame of the modified video generated by the system. Frame 330b depicts the same scene as frame 330a, with the modification specified by the input query above, applied to the text 331. The frame 330b does not display any text corresponding to text 331 or any alternative text in its place. The pixels of frame 330b that correspond to the pixels of the text 331 of frame 330a have been in-filled or inpainted according to the underlying scene. The text 335 remains the same in frame 330b as in frame 330a.
[0092] Frames 340a and 324b depict a text-editing task that is modifying a visual feature of text, in particular, modifying the size of specified text. Frame 340a is an example frame of an input video that depicts a scene with one movie character. Frame 340a displays text 341 “A scene from a movie.” The text 341 is an example of overlay text that is not part of the scene. Frame 340a also displays the text 341 as following certain visual constraints, such as not covering the movie character’s face, and fitting within the frame 340a.
[0093] The system can generate a modified video that includes the frame 340b by receiving an input video that includes the frame 340a and an input uery, and providing the input video and the input query to a generative neural network as described above. The input query can specify the text to be displayed in a different size and a definition for the change in size. For example, the change in size can be defined by a type of change, such as an increase in size or a decrease in size, and a magnitude of the change, such as a percentage or a proportion compared to the text 341. In the example of the frames 340a and 340b, the input query' may include “Increase the size of the text in the video: “A scene from a movie” by 10%.” In some examples, the capitalization of the specified text of the input query may differ from the capitalization of the text 341. For example, the input uery may include “Increase the size of the text in the video: “A SCENE from A movie” by 10%.”
[0094] Frame 340b is an example frame of the modified video generated by the system. Frame 340b depicts the same scene as frame 340a, with the modification specified by the input query' above, applied to the text 341. The frame 340b displays text 342 that has the same content and sty le as the text 341, but displayed in a different size. For example, the text 342 is displayed in a larger size than the text 341. In addition, because the text 342 is larger, the text 342 is also displayed with added pixel detail compared to the text 341. The frame 340b also displays the text 342 as following the same visual constraints as the text 341. That is, frame 340b also displays the text 342 as not covering the movie character's face and as fitting within the frame 340b.
[0095] Furthermore, the text 342 has a different layout. The text 341 is divided between two lines, “A scene,” and “from a movie.” The text 342 is divided between three lines, “A scene,” “from,” and “a movie."’ If the text 342 had a larger size but had the same layout and followed the same visual constraint of not covering the movie character’s face as the text 341, the text 342 may not fit in the frame 340b. This may violate the other visual constraint of fitting the text within the frame.
[0096] Frames 350a and 350b depict a text-editing task that is replacing text with alternative text. Frame 350a is an example frame of an input video that depicts a scene with an observatory and shapes in the sky. Frame 350a displays text 351 “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY.” The text 351 is displayed in black with a gray outline. The text 351 is an example of overlay text that is not part of the scene.
[0097] The system can generate a modified video that includes the frame 350b by receiving an input video that includes the frame 350a and an input query, and providing the input video and the input query to a generative neural network as described above. The input query can specify the text to be replaced and the alternative text. In the example of the frames 350a and 350b, the input query may include “Replace the text “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY” with the text “UBERALL MERKWURDIGE DINGE AU HIMMEL”.” In some examples, the capitalization of the specified text of the input query may differ from the capitalization of the text 351 . For example, the input query may include “Replace the text “everywhere you look weird things in the sky” with the text “UBERALL MERKWURDIGE DINGE AU HIMMEL”.”
[0098] Frame 350b is an example frame of the modified video generated by the system. Frame 350b depicts the same scene as frame 350a, with the modification specified by the input query above, applied to the text 351. The frame 350b displays text 352 that has the same visual features, such as style and placement, as the text 351, but with different content. For example, the text 352 is displayed in black with a gray outline. The text 352 includes “UBERALL MERKWURDIGE DINGE AU HIMMEL” in place of “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY.” In this example, the text 352 includes characters such as “U” that are not found in the text 351.
[0099] Referring to FIG. 3C, frames 360a and 360b are another example of the text-editing task of replacing text with alternative text. Frame 360a is an example frame of an input video that depicts a scene with four characters and a raccoon. Frame 360a displays text 361 “There was a noisy raccoon in our hotel!’' The text 361 is an example of overlay text that is not part of the scene. The text 361 is displayed as following certain visual constraints, such as fitting within a text box 365.
[0100] The system can generate a modified video that includes the frame 360b by receiving an input video that includes the frame 360a and an input query, and providing the input video and the input query’ to a generative neural network as described above. The input query can specify the text to be replaced and the alternative text. In the example of the frames 360a and 360b, the input query' may include “Replace the text “There was a noisy raccoon in our hotel!” with the text “jHabia un mapache ruidoso en nuestro hotel!”.” In some examples, the capitalization of the specified text of the input query’ may differ from the capitalization of the text 361. For example, the input query may include “Replace the text “there was a noisy raccoon in our hotel!” with the text “jHabia un mapache ruidoso en nuestro hotel!”.”
[0101] Frame 360b is an example frame of the modified video generated by the system. Frame 360b depicts the same scene as the frame 360a, with the modification specified by the input query above, applied to the text 361. The frame 360b displays text 362 that has the same visual features, such as style and following the visual constraints, as the text 361. but with different content and typographic elements. For example, the text 362 is also displayed within the text box 365. The text 362 includes “jHabia un mapache ruidoso en nuestro hotel! ” in place of “There was a noisy raccoon in our hotel! ” In this example, the text 362 includes characters such as “j” that are not found in the original text 361.
[0102] FIG. 4 is a flow diagram of an example process for generating a modified video. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system for generating a modified video, e.g., the system 200 of FIG. 2, appropriately programmed in accordance with this specification, can perform the process 400.
[0103] The system receives an input video (step 410). For example, the system can receive the input video from a user. The input video can include multiple video frames. One or more of the video frames may display text.
[0104] The system receives an input query (step 420). The input query can specify a textediting task to be performed on the input video. In some implementations, the system can receive the input query’ from a user.
[0105] For example, the input query can include a natural language description of how text displayed in the input video should be modified. In some examples, the input query can also specify the text to be modified and / or alternative text for replacing the text to be modified. In some examples, the input query can specify a text-editing task that modifies all text displayed in the input video. For example, the input query can include a description of the modification to be made, followed by '‘the text in the video” without identifying specific text to change.
[0106] In some examples, the input query can specify a text-editing task that modifies specific text displayed in the input video. The input query’ can identify the text displayed in one or more video frames of the input video to be modified. For example, the input query can identify the text using specific characters or symbols. For example, the text to be modified can be enclosed within quotation marks. As an example, the input query’ can include a description of the modification to be made, follow ed by “the text: “noisy raccoon”.” In some examples, the input query can specify the text to be modified using different capitalization than the capitalization of the text displayed in one or more video frames of the input video to be modified. For example, the input quety can include a description of the modification to be made, followed by “the text: “Noisy raccoon”.”
[0107] Alternatively or in addition, the input query can identify the text using the time that the text is displayed or the time interval that the text is displayed in the input video. For example, the text to be modified can be displayed in the input video at one or more timestamps or during a time interval. As an example, the input query’ can include “From 10:03 to 10: 17 seconds in the video,” follow ed by a description of the modification to be made to the text displayed in the input video. As another example, the input query can include “Starting at 10:03 seconds in the video,” followed by a description of the modification to be made to the text displayed in the video.
[0108] As another example, the input query can identify’ the text using quotation marks and using the time interval that the text is displayed in the input video. For example, the input query can include “From 10:03 to 10: 17 seconds in the video,” followed by a description of the modification to be made, and further followed by “the text: “noisy raccoon”.”
[0109] In some examples, the input query can specify a text-editing task that replaces specific text displayed in the input video with alternative text. The input query can also identify the alternative text using specific characters or symbols. For example, the alternative text can be enclosed within quotation marks. As an example, the input query can include a description of replacing “the text: “noisy raccoon” with the text: “un mapache ruidoso”.” In some examples, the input query can specify the text to be replaced using different capitalization than the capitalization of the text display ed in one or more video frames of the input video to be replaced. For example, an alternative input query to the example input query above can include a description of replacing ‘‘the text: “Noisy raccoon” with the text: “un mapache ruidoso”.”
[0110] Alternatively or in addition, as described above, the input query can identify the text using the time that the text is displayed or the time interval that the text is displayed in the input video. For example, the text to be replaced can be displayed in the input video at one or more timestamps or during a time interval. As an example, the input query can include “From 10:03 to 10: 17 seconds in the video,” followed by a description of replacing specified text with alternative text. As another example, the input query can include “Starting at 10:03 seconds in the video,” followed by a description of replacing specified text with alternative text.
[0111] As another example, the input query can identify the text using quotation marks and using the time interval that the text is displayed in the input video. For example, the input uery can include “From 10:03 to 10: 17 seconds in the video,” followed by a description of replacing “the text: “Noisy raccoon” with the text: “un mapache ruidoso”.”
[0112] In some implementations, the system can generate the input query’ based on inputs from the user. For example, the system can provide a user interface that allows the user to play and pause the input video. The user interface can also allow the user to input a natural language description of how text displayed in the input video should be modified. For example, the user can input a natural language description such as “Replace the text with “Foundation models are cool!”.” The user may pause the input video at a timestamp of 10:05 seconds. The system can determine that at 10:05 seconds into the input video, the text displayed is “jLos modelos de base son geniales!”
[0113] The system can identify the time interval that the text “jLos modelos de base son geniales!” is displayed in the input video and that includes the timestamp of 10:05. The system can identify the time interval using a machine learning model, for example, that is configured to identify a time interval that a given video displays given text, where the time interval includes a given timestamp. For example, the system can determine that the text “Foundation models are cool!” that is displayed in the input video at 10:05 is displayed in the input video from 10:02 to 10: 19. The system can include the time interval in the input query. For example, the system can generate the input query to include “Replace the text in the video from 10:02 to 10: 19 seconds with “jLos modelos de base son geniales!”.”
[0114] The text-editing task can require a modification to text displayed in one or more video frames of the input video. The text can be, for example, overlay text, in-scene text, or animated text. Example text-editing tasks are described below. For example, the text-editing task can include replacing the text with alternative text that shares visual features with the text. The input query can specify the text to be replaced and the alternative text. Examples of replacing text with alternative text are described above with reference to frames 350a and 350b of FIG. 3B, and frames 360a and 360b of FIG. 3C.
[0115] As another example, the text-editing task can include modifying one or more visual features of the text. For example, visual features can include color, size, font, and weight. The text-editing task can include modifying visual features of all of the text displayed in the input video. For example, the input query can specify' a different color for the text. An example of modifying the color of all of the text displayed in the input video is described above with reference to frames 310a and 310b of FIG. 3 A.
[0116] As another example, the text-editing task can include changing the size of all of the text displayed in the input video. The input query can specify a definition representing a change in size for the text displayed in the input video.
[0117] The text-editing task can also include selectively modifying visual features of the text displayed in the input video. For example, the input query can specify text to be displayed in a different color and the different color. An example of selectively modifying the color of text displayed in the input video is described above with reference to frames 320a and 320b of FIG. 3 A.
[0118] As another example, the text-editing task can include selectively changing the size of the text displayed in the input video. For example, the input query can specify text to be displayed in a different size and a definition representing a change in size for the specified text. An example of selectively changing the size of text displayed in the input video is described above with reference to frames 340a and 340b of FIG. 3B.
[0119] As another example, the text-editing task can include removing or erasing the text. For example, the input query can specify text to be removed. An example of selectively removing the text displayed in the input video is described above with reference to frames 330a and 330b of FIG. 3B.
[0120] As another example, the text-editing task can include removing or erasing all of the text displayed in the input video. The input query can specify that the text of the video should be removed or erased.
[0121] The system processes, using a generative neural network, the input video and the input query to generate an output that includes a modified video (step 430). The modified video can include the modification applied to the text displayed in the one or more video frames. For example, the generative neural network can process the input query, e.g., using a text encoder neural network, to generate an encoded representation of the input query. For example, the encoded representation can include tokens for the input query. An ‘'encoded representation” or “embedding” as used in this specification is a vector of numeric values, e.g., floating point values or other values, having a pre-determined dimensionality7. The space of possible vectors having the pre-determined dimensionality is referred to as the “embedding space.” The generative neural network can process the input video, e.g., using an encoder neural network, to generate an encoded representation of the input video. For example, the encoder neural network can process the intensity values of the pixels of the input video. For example, the encoded representation can include tokens for the input video.
[0122] The generative neural network can generate one or more video frames of the output video conditioned on the encoded representation of the input query and the encoded representation of the input video. For example, the generative neural network can generate a sequence of input tokens from the encoded representation of the input query7and the encoded representation of the input video. The generative neural network can process the sequence of input tokens to generate an output that includes the modified video. For example, the generative neural network can generate the one or more video frames of the modified video autoregressively. That is, the generative neural network can auto-regressively generate an output sequence of tokens by generating each token in the output sequence conditioned on a current input sequence that includes (i) the sequence of input tokens followed by (ii) any tokens that precede the particular token in the output sequence. The generative neural network can generate the one or more video frames from the output sequence of tokens, e.g., using a decoder neural network. As a particular example, the generative neural network can include an auto-regressive Transformer-based neural network that includes a plurality of layers that each apply a self-attention operation. The neural network can have any of a variety of Transformer-based neural network architectures.
[0123] As another example, the generative neural network can generate the one or more video frames of the modified video by reversing a forward diffusion process conditioned on a representation of the input query and a representation of the input video.
[0124] In some implementations, the output can also include a status indicator for the modification. The status indicator can indicate whether the modification was successfully applied. For example, the generative neural network may not detect any text or the text specified by the input query in the input video. The generative neural network can generate an output that includes a status indicator that indicates the modification was not successfully applied. For example, the output can include ‘‘Text not found.’' The status indicator can be used during fine-tuning of the generative neural network as described below with reference to FIG. 5.
[0125] The status indicator can also indicate a location of the text displayed in one or more video frames of the input video. For example, the location of the text can be represented by a bounding box. The bounding box can include a rectangle that defines the text displayed in one or more video frames of the input video and can be defined by pixel coordinates. For example, the bounding box can be defined by the coordinates of the top left comer of the rectangle and the bottom right comer of the rectangle.
[0126] In some examples, the location of the text differs across video frames. For example, the text may be animated text that is displayed in different locations, sizes, rotations, transparencies, colors, etc. in different video frames. The location of the text can be represented by a bounding box for the first video frame that the text appears in.
[0127] In some examples, the location of the text can be represented by data representing a fine-grained mask of pixels. For example, the data representing a fine-grained mask of pixels can include data representing the locations of pixels that display the text.
[0128] In some implementations, the system can further provide the modified video for display to a user. For example, the system can provide data representing the modified video to a user interface for display to the user.
[0129] The generative neural network can have been pre-trained to generate at least a video conditioned on a video and a text. In some implementations, the generative neural network has been fine-tuned on a training dataset for multiple text-editing tasks. Each of the textediting tasks can require modifying pixels that depict text in input videos. In some examples, the generative neural network has been fine-tuned sequentially on each of the multiple textediting tasks of the training dataset. Fine-tuning on the training dataset is described in more detail below with reference to FIG. 5.
[0130] In some implementations, the generative neural netw ork has been further fine-tuned to minimize a loss function based on an aggregate reward value. Fine-tuning based on the aggregate reward value is described in more detail below with reference to FIGS. 5-7.
[0131] FIG. 5 is a flow diagram of an example process 500 for fine-tuning a generative neural network. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system, e.g., the training system 240 of FIG. 2, appropriately programmed in accordance with this specification, can perform the process 500 to train the generative neural network 210 of FIG. 2.
[0132] The training system can obtain data representing a pre-trained generative neural network (step 510). For example, the generative neural network can have been pre-trained to generate at least a video conditioned on a video and text. For example, the training system can clone a version of a generative neural network that has been pre-trained to generate a video conditioned on a video and text. Examples of a generative neural network that has been pre-trained to generate a video conditioned on a video and text include VideoPoet, Gen- 1, and Gen-2.
[0133] The training system can obtain a training dataset (step 520). For example, the training dataset can be the training dataset 244 described with reference to FIG. 2. The training dataset can include multiple training examples. Each training example can include a training video and a training query that specifies a text-editing task from the multiple text-editing tasks. Each training example can also include a corresponding training output that includes at least an output video. In some implementations, the corresponding training output can also include a status indicator. For example, the status indicator can indicate whether there is any text displayed in the training video, or whether the text specified by the training query is displayed in the training video. The status indicator can also indicate the location of text displayed in the training video.
[0134] In some implementations, the training system can generate the training dataset. For example, the training system can receive or generate seed videos that each include one or more video frames. The training system can generate synthetic training videos and / or training outputs by adding text to the seed videos.
[0135] The training system can generate training videos and corresponding output videos that follow certain visual constraints. For example, visual constraints may include not positioning text over a character’s face, positioning text relative to underlying visual content, or ensuring that all of the text content is displayed within the video frame.
[0136] The training system can also generate different training videos and corresponding output videos that display text with different visual features such as color, size, style, font, orientation, geometric distortion, lighting, pattern, position, layout, and animation.
[0137] The training system can also generate different training videos and output videos with different types of text characters. For example, the training system can generate different training videos and output videos that display different words and characters, for example, in different alphabets for different languages. The training system can also generate training videos that do not display text or do not display the specified text of the training query. The training system can generate expected status indicators that indicate the text is not found.
[0138] As an example, the text-editing task for a particular training example can be modifying a visual feature of all text in an input video. For example, the training query' can specify a color such as yellow. The training video can include a video that displays text in a first color such as pink. The training output can include an output video that displays the same text content and other visual features as the text displayed in the training video, but in a second color, yellow.
[0139] The training system can generate the particular training example by obtaining a seed video that does not display text. The training system can generate the training video by creating a copy of the seed video and adding text in pink and with certain other visual features. The training system can generate the output video by creating a copy of the seed video and adding the same text content in yellow and with the same other visual features. The training system can generate the training query to include, for example, "‘Change the color of the text in the video to yellow.” The training system can also generate an additional training example. For example, the training system can use the training video of the particular training example as the output video of the additional training example, and the output video of the particular training example as the training video of the additional training example. The training system can generate the training query of the additional training example to include “Change the color of the text in the video to pink.”
[0140] The training system can generate different training examples for the same text. For example, the training system can generate different training videos that display the text in different colors. In particular, the training system can generate training videos that display the text in colors that are found in or similar to colors that are found in the scene of the training video. The training system can thus fine-tune the generative neural network to determine which pixels correspond to text, rather than learning to replace all pixels of a certain color.
[0141] As another example, the text-editing task for a particular training example can be modifying a visual feature of specified text in an input video. For example, the training query can specify text to be displayed in a different color, and a color such as yellow. The training video can include a video that displays text in a first color such as pink. The training output can include an output video that displays the same text content and other visual features as the text displayed in the training video, but with the specified text displayed in a second color, yellow.
[0142] The training system can generate the particular training example by obtaining a seed video that may or may not display text. The training system can generate the training video by creating a copy of the seed video and adding text in pink and with certain other visual features. The training system can generate the output video by creating a copy of the seed video and adding some of the text content in yellow, and some of the text content in pink, and with the same other visual features. The training system can generate the training query to include, for example, “Change the color of the text “a film” in the video to yellow.” The training system can also generate an additional training example. For example, the training system can use the training video of the particular training example as the output video of the additional training example, and the output video of the particular training example as the training video of the additional training example. The training system can generate the training query7of the additional training example to include “Change the color of the text “a film” in the video to pink.”
[0143] As another example, the text-editing task for a particular training example can be erasing text in an input video. For example, the training query can specify text to be erased. The training video can include a video that displays the specified text. The training output can include an output video that does not display the specified text.
[0144] The training system can generate the particular training example by obtaining a seed video that may or may not display text. The training system can generate the training video by creating a copy of the seed video and adding text such as “Something to delete.” The training system can generate the output video by creating a copy of the seed video. The training system can generate the training query to include, for example. “Erase the text in the video: “Something to delete”.”
[0145] As another example, the text-editing task for a particular training example can be changing the size of text in an input video. For example, the training query can specify text to be displayed in a different size and a definition for the change in size. The training video can include a video that displays text. The training output can include an output video that displays the same text content and other visual features as the text displayed in the training video, but in a different size.
[0146] The training system can generate the particular training example by obtaining a seed video that may or may not display text. The training system can generate the training video by creating a copy of the seed video and adding text with certain visual features, such as following certain visual constraints. The training system can generate the output video by creating a copy of the seed video and adding the same text content in a different size, with the same or similar other visual features as the text of the training video. The training system can generate the training query to include, for example, “Increase the size of the text in the video: “A scene from a movie” by 10%.” The training system can also generate an additional training example. For example, the training system can use the training video of the particular training example as the output video of the additional training example, and the output video of the particular training example as the training video of the additional training example. The training system can generate the training query' of the additional training example to include “Decrease the size of the text in the video: “A scene from a movie” by 10%.”
[0147] In some examples, the output videos may display text with different amounts of pixel detail than the training video. For example, when displayed in a smaller size, the text may have less pixel detail. When displayed in a larger size, the text may have more pixel detail.
[0148] In some examples, the training system may generate corresponding output videos that display the text content in different positions or layouts. The corresponding output video should display the text content in approximately the same position as the training video, but in potentially a different layout, and follow the visual constraints. For example, when displayed in a smaller size, some words of the text may be combined into a single line. When displayed in a larger size, some words of the text may be separated into different lines.
[0149] As another example, the text-editing task for a particular training example can be replacing specified text with alternative text. For example, the training query can specify text to be replaced and alternative text. The training video can include a video that displays the specified text. The training output can include an output video that does not display the specified text, and displays the alternative text with similar visual features as the specified text in the training video. The alternative text can include one or more different words or characters. For example, the alternative text can be text of a different language than the specified text.
[0150] The training system can generate the particular training example by obtaining a seed video that may or may not display text. The training system can generate the training video by creating a copy of the seed video and adding text with certain visual features. The training system can generate the output video by creating a copy of the seed video and adding alternative text with the same or similar visual features as the text of the training video. The training system can generate the training query to include, for example, “Replace the text “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY’’ with the text “UBERALL MERKWURDIGE DINGE AU HIMMEL”.” The training system can generate an additional training example. For example, the training system can use the training video of the particular training example as the output video of the additional training example, and the output video of the particular training example as the training video of the additional training example. The training system can generate the training query of the additional training example to include “Replace the text “UBERALL MERKWURDIGE DINGE AU HIMMEL” with the text “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY”.”
[0151] In some examples, the alternative text may be longer or shorter than the text of the training video. In order to follow the visual constraints, the corresponding output videos maydisplay text in a different size than the training video. The output videos may display text with different amounts of pixel detail than the training video, or in different positions or layouts as described above.
[0152] The training system can fine-tune the generative neural network on the training dataset (step 530). For example, the training system can process each training video using the pre-trained generative neural network to determine an update to the parameters of the generative neural network, e g., using backpropagation. For example, the training system can fine-tune the generative neural network to minimize a based on a difference between the video generated by the generative neural network for the training video and the output video of the corresponding training output, also referred to as the expected video. The loss can be measured per-pixel, for example. For example, the loss can be a mean squared error (MSE) loss, a mean absolute error (MAE) loss, or a structural similarity index measure (SSIM).
[0153] In some implementations, the loss can also be measured between bounding boxes represented by the text of the status indicators. That is, the loss can be a combined loss based on the difference between the generated video and the expected video, as well as the difference between the bounding box of the status indicator generated by the generative neural netw ork, also referred to as the generated bounding box, and the bounding box of the status indicator of the training output, also referred to as the expected bounding box. For example, the loss can be based on a difference in location and / or size between the generated bounding box and the expected bounding box. For example, if the generated bounding box is not located in the same region of the video frame as the expected bounding box, the loss value can be higher than if the generated bounding box is located in the same region as the expected bounding box. As another example, the loss can also be based on an overlap between the generated bounding box and the expected bounding box. For example, the loss based on the difference between the generated bounding box and the expected bounding box can be an MSE loss or an MAE loss.
[0154] In some implementations, the loss can also be measured between masks represented by the text of the status indicators. That is, the loss can be a combined loss based on the difference between the generated video and the expected video, as well as the difference between a fine-grained mask of pixels of the status indicator generated by the generative neural network, and the fine-grained mask of pixels of the status indicator of the training output. For example, the loss based on the difference between the masks of pixels can be a pixel-wise cross entropy loss, or a boundary loss.
[0155] In some implementations, the loss can also be based on the difference between the text of the status indicators. That is. the loss can be a combined loss based on the difference between the generated video and the expected video, as well as the difference between the text of the status indicators. For example, for some training examples, the expected status indicator may indicate that no text is displayed in the training video, or the specified text is not displayed in the training video. If the generated status indicator indicates that text was found, the combined loss value can be higher, compared to a loss that is based only on the difference between the generated video and the expected video, because of the difference between the text of the status indicators.
[0156] As another example, the expected status indicator may indicate that the specified text is first displayed at 9:00 seconds. If the generated status indicator indicates that the specified text was first displayed at 11 :00 seconds, the combined loss value can be higher, compared to a loss that is based only on the difference between the generated video and the expected video, because of the difference between the text of the status indicators.
[0157] In some implementations, the training system can train the generative neural network on the training dataset using curriculum learning. For example, the training system can train the generative neural network sequentially on training examples for different text-editing tasks. For example, the training system can fine-tune the generative neural network sequentially on each of the multiple text-editing tasks of the training dataset.
[0158] Each of the text-editing tasks can have an associated complexity level. The textediting tasks can be ordered by increasing levels of complexity in a sequence of text-editing tasks. For example, the sequence of text-editing tasks can start with tasks of lower complexity. For example, tasks of lower complexity can include simpler tasks or tasks that require fewer steps or subtasks to accomplish than tasks of higher complexity. The complexity of subsequent tasks of the sequence can increase gradually. For example, tasks of higher complexity may require subtasks that require more steps or subtasks to accomplish than tasks of lower complexity. In some examples, the subtasks can include one or more preceding tasks of the sequence.
[0159] The training system can train the generative neural network on training examples for the same text-editing task for multiple consecutive training stages before training the generative neural network on training examples for the next text-editing task in the sequence of text-editing tasks. For example, the training system can train the generative neural network on training examples for a particular text-editing task until the training system has trained the generative neural network on all of the training examples for the particular textediting task in the training dataset. As another example, the training system can train the generative neural network on training examples for a particular text-editing task for a predetermined number of training stages. As another example, the training system can train the generative neural network on training examples for a particular text-editing task until a predetermined performance (e g., a training or validation accuracy) of the generative neural network is achieved. As another example, the training system can train the generative neural network on training examples for a particular text-editing task until a marginal improvement in the performance (e g., in the training or validation accuracy) of the generative neural network between respective training stages drops below a predetermined threshold.
[0160] As an example, the sequence of text-editing tasks can include i) modifying a visual feature of all text in an input video, ii) modifying a visual feature of specified text in an input video, iii) erasing text in an input video, iv) changing the size of text in an input video, and v) replacing specified text with alternative text. In this example, modifying a visual feature of all text in an input video is the least complex task. Replacing specified text with alternative text is the most complex task. For example, replacing specified text with alternative text may require subtasks such as erasing text in an input video.
[0161] In some implementations, the training system can further fine-tune the generative neural network to minimize a loss function based on an aggregate reward value (step 540). The training system can derive the aggregate reward value from one or more reward values. Each reward value can be generated by a corresponding reward model. The reward models can have been trained to generate the reward values for corresponding comparison prompts.
[0162] For example, to determine the aggregate reward values, the training system can receive an ultimate dataset. The ultimate dataset can include examples that each include an original video that displays text and an input query that specifies a text-editing task to be performed. For each example, the training system can provide the original video to the generative neural network to generate a modified video. The training system can provide the generated video to the reward models to generate the one or more reward values. The training system can derive the aggregate reward values from the one or more reward values. Fine-tuning the generative neural network based on the aggregate reward value is described in further detail below with respect to FIG. 7.
[0163] In some implementations, the training system can further fine-tune the generative neural network to minimize a loss function based on an aggregate reward value at each of multiple fine-tuning iterations. For example, for each generated video for each iteration, the training system can generate one or more reward values for the generated video using the reward models. The training system can derive the aggregate reward value for the generated video from the one or more reward values.
[0164] In some implementations, the training system can derive the reward models from the fine-tuned generative neural network. For example, the training system can obtain a copy of data representing the generative neural network. The training system can modify the copy of data representing the generative neural network to obtain a reward model. For example, the training system can replace one or more last layers of the generative neural network with one or more layers that produce a reward value.
[0165] The training system can train each reward model to generate a reward value for a corresponding comparison prompt. The training system can train each reward model on reward training examples that each include a reward training input and a reward training output. The reward training input can include an input video of an evaluation example, the generated output video for the input query' of the evaluation example, and a comparison prompt. The reward training output can include the ratings data for the generated output video for the aspect evaluated by the comparison prompt. The training system can train the reward models to learn rater preference with respect to comparison prompts that evaluate different aspects of videos.
[0166] The training system can obtain the reward training examples for the reward models from an evaluation dataset and ratings data. The evaluation dataset can include evaluation examples that each include an input video that displays text, an input query that specifies a text-editing task to be performed, and a comparison video. In some examples, the comparison video is professionally generated. In some examples, the comparison video is an output video generated by a different version of the generative neural network. For example, the different version of the generative neural network can be an earlier checkpoint for the generative neural network. The training system can provide the input video of each evaluation example as input to the generative neural network. The training system can obtain each generated output video from the generative neural network. The training system can generate ratings data for each generated output video compared to the comparison video for a particular aspect. For example, the training system can derive the ratings data from preference data received from one or more human raters.
[0167] The ratings data for the generated output video can represent a preference between the generated output video and the comparison video for a particular aspect. The training system can receive the preference data from a rater in response to presenting the generated output video and the comparison video to the rater.
[0168] After the reward models have been trained, the training system can use the reward models to generate reward values for each generated video. The training system can combine the reward values to obtain an aggregate reward value for the generated video, as described in further detail below with reference to FIG. 7.
[0169] An example presentation 600 of a user interface for collecting human feedback is depicted in FIG. 6. The training system can obtain preference data through the presentation 600. The presentation 600 can be displayed on a user device of a rater, for example. The user device can be any type of computer or computing device that has a display and is configured to receive an input from the rater. For example, the user device can be a computer, laptop, tablet, or mobile phone that has a display and can receive user inputs through interfaces such as a keyboard, mouse, touchpad, or touchscreen.
[0170] The presentation 600 can include the input video 610 of the evaluation example, the generated output video 620, and the comparison video 630. The presentation 600 can also include a description of the editing task between the input video 610, and the generated output video 620 and the comparison video 630. In the example of FIG. 6, the editing task is “Replaced “There was a noisy raccoon in our hotel!” with “jHabia un mapache ruidoso en nuestro hotel!”.”
[0171] The presentation 600 can also include prompts 640 for the rater to respond to. Each prompt can include a query or question about assessing the generated output video 620 and the comparison video 630 according to a particular aspect. For example, prompt 640a includes “Which video matches the look of the original text better?” Prompt 640b includes “Which video matches the layout of the original text better?”
[0172] In the example of FIG. 6, a rater may indicate, through user input to the presentation 600, that the generated output video 620 matches the look of the original text (the text of the input video 610) better, and the comparison video 630 matches the layout of the original text better. In some examples, the rater may indicate that the videos are indistinguishable or that they are ‘'not sure.”
[0173] The training system can receive the preference data indicating the rater’s preferences for each of the prompts 640. In some examples, the training system can combine the preference data for the rater with preference data for other raters.
[0174] The training system can assign a value for ratings data for the generated output video 620 based on the preference data. For example, a higher value for the generated output video 620 can represent a higher preference for the generated output video 620 for the particular aspect. A lower value for the generated output video 620 can represent a lower preference for the generated output video 620 for the particular aspect.
[0175] In some examples where the comparison video 630 is an output video generated by a previous version of the generative neural network, the training system can assign values for the generated output video 620 based on preference data indicating preferences relative to comparison videos generated by previous versions of the generative neural network. For example, the training system can provide a series of presentations such as the presentation 600 to the rater. Each presentation can include a comparison between an output video generated by a particular version of the generative neural network, and an output video 620 generated by a different version of the generative neural network. The training system can receive the preference data for each prompt for each comparison.
[0176] The training system can determine an order for how the output videos of each version are preferred relative to output videos of other versions for each particular aspect. For example, the training system can determine that the output video for version three was more preferred than the output video for version seven, which was more preferred than the output video for version two, for a particular aspect. In some examples, the training system can determine the order by comparing the generated output videos for each version to the output videos for each other version. In some examples, the training system can assume the transitive property to determine the order in 0(1 og n) comparisons.
[0177] Based on the order, the training system can assign a value for the ratings data for the output video generated by the current version. For example, the output videos can be ordered in a decreasing preference order. Each output video in the sequence can correspond to a value at a predetermined interval between 1 and -1. For example, the first output video can correspond to a value of 1. the second output video can correspond to a value of 0.75, and the third output video can correspond to a value of 0.5. If the output video for the current version is the output video with the lowest preference, or the last output video, the training system can assign a value of -1 for the ratings data. If the output video for the cunent version is the output video with the highest preference, or the first output video, the training system can assign a value of 1.0 for the ratings data. If the output video for the current version is an intermediate video, the training system can assign the corresponding value.
[0178] In some examples, the preference data may indicate that two or more output videos are indistinguishable for a particular aspect. For example, the rater may have indicated that they were “not sure” which video was better according to the prompt. As another example, there may have been a split in preference between the output videos among multiple raters. The two or more output videos can correspond to the same value.
[0179] FIG. 7 is a diagram of an example process 700 for fine-tuning a generative neural network based on an aggregate reward value 710. For example, the generative neural network can be the generative neural network 210 of FIG. 2. The process 700 can be performed by the training system 240 of FIG. 2 as part of step 540 of FIG. 5. The training system can perform the process 700 at each of multiple fine-tuning iterations.
[0180] The training system can provide an original video of an example from the ultimate dataset, Ui, to the generative neural network 210. In the example of FIG. 7, the current version of the generative neural network 210 is version m.
[0181] The training system can provide the input query of the example, “replacement prompt;” to the generative neural network 21 . The generative neural network 210 can generate an output video Gi,m.
[0182] The training system can provide the output video Gi,m and the original video Ui to one or more reward models 242a-k, such as reward model 242a, to generate reward values 750a- k. Each of the reward models 242 can be configured to generate a reward value 750, ai.m.k, for a particular comparison prompt k, 740, given a video. The reward value 750 can represent an assessment of a generated video Gi,m and an original video Ui according to the corresponding comparison prompt 740.
[0183] For example, the reward model 242a can receive a video that includes an original video, Ui. from the ultimate dataset, and a generated video, Gi.m, for the original video Ui. The video can be a concatenation of Ui + Gi,m.
[0184] The comparison prompt 740a can correspond to the reward model 242a. The comparison prompt 740a can include a question or query about the same aspect for which the rew ard model 242a was trained to predict rater preferences. The comparison prompt 740a can also indicate the text displayed in each video of Ui and Gi,m on which to focus the assessment for a particular aspect. For example, the original video Ui may display “textl.” The input query may include “Replace “textl” in the video with “text2” .” The generated video Gi,mmay display “text2.” The comparison prompt 740a can thus include “Assess “textl” and “text2” for visual similarity .”
[0185] Other comparison prompts for the other reward models 242b-k of FIG. 7 can include questions or queries about aspects such as look, feel, layout, style, font, weight, or color of text displayed in one or more video frames of the generated video Gi,mcompared to text displayed in one or more video frames of the original video Ui.
[0186] The training system can receive the reward values 750 from the reward models 242. The training system can generate an aggregate reward value 710 from the reward values 750. In the example of FIG. 8, the aggregate reward value 710 is an average of the reward values 750. In some implementations, each of the reward values 750 can be multiplied by a corresponding weight. For example, the training system may place more emphasis on layout than font by assigning a higher corresponding weight to the reward value for layout.
[0187] The training system can then fine-tune the generative neural network 110 to minimize the loss function 720. The loss function 720 is based on the aggregate reward value 710. In the example of FIG. 7, the loss function 720 is a sum of the aggregate reward value 710 and a constant 1.0, divided by a constant 2.0. The training system can thus fine-tune the generative neural network 210 based on reward values, which are predicted by reward models that are configured to predict human preferences.
[0188] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0189] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry7, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of. data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0190] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0191] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0192] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0193] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0194] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry7. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data. e.g.. magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0195] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory7, media and memory7devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory7devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0196] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g.. a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0197] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
[0198] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0199] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0200] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0201] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0202] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0203] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0204] What is claimed is:
Claims
CLAIMS1. A computer-implemented method comprising: receiving an input video; receiving an input query specifying a text-editing task to be performed on the input video, the text-editing task requiring a modification to text displayed in one or more video frames of the input video; and processing, using a generative neural network, the input video and the input query to generate an output comprising a modified video, with the modification applied to the text displayed in the one or more video frames.
2. The method of claim 1, wherein receiving an input video comprises receiving the input video from a user.
3. The method of any preceding claim, wherein receiving an input query comprises receiving the input query from a user.
4. The method of any preceding claim, wherein the output further comprises a status indicator for the modification.
5. The method of claim 4, wherein the status indicator indicates whether the modification was successfully applied, or a location of the text displayed in one or more video frames of the input video.
6. The method of any preceding claim, further comprising providing the modified video for display to a user.
7. The method of any preceding claim, wherein the text displayed in one or more video frames of the input video comprises any one or more of: overlay text, in-scene text, or animated text.
8. The method of any preceding claim, wherein the text-editing task comprises any one or more of: modifying one or more visual features of the text, removing the text, or replacing the text with alternative text that shares visual features with the text.
9. The method of any preceding claim, wherein the input query specifies i) text to be replaced and ii) alternative text.
10. The method of any preceding claim, wherein the input query' specifies a different color for the text.
11. The method of any preceding claim, wherein the input query specifies i) text to be displayed in a different color and ii) the different color.
12. The method of any preceding claim, wherein the input query specifies text to be removed.
13. The method of any preceding claim, wherein the input query' specifies i) text to be displayed in a different size and ii) a definition representing a change in size for the specified text.
14. The method of any preceding claim, wherein the generative neural network has been pre-trained to generate at least a video conditioned on a video and text.
15. The method of any preceding claim, wherein the generative neural network has been fine-tuned on a training dataset for a plurality of text-editing tasks that each require modifying pixels that depict text in input videos.
16. The method of claim 15, wherein the generative neural network has been fine-tuned sequentially on each of the plurality of text-editing tasks of the training dataset.
17. The method of claim 15, wherein the training dataset comprises a plurality of training examples, each comprising a training input comprising a training video and a training query specifying a text-editing task from the plurality of text-editing tasks, and a corresponding training output comprising at least an output video.
18. The method of any of claims 14-17, wherein the generative neural network has been further fine-tuned to minimize a loss function based on an aggregate reward value.
19. The method of claim 18, wherein the aggregate reward value is derived from one or more reward values, wherein each reward value is generated by a corresponding reward model.
20. The method of claim 19, wherein each corresponding reward model is configured to generate a reward value representing an assessment of a given video according to a corresponding comparison prompt and an original video.
21. The method of claim 20, wherein the corresponding comparison prompt comprises a question about a visual similarity, a layout, a style, a font, a weight, or a color of text displayed in one or more video frames of the given video compared to text displayed in one or more video frames of the original video.
22. The method of claim 21, wherein each corresponding reward model is trained to generate the reward value for the corresponding comparison prompt.
23. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising: receiving an input video; receiving an input query specifying a text-editing task to be performed on the input video, the text-editing task requiring a modification to text displayed in one or more video frames of the input video; and processing, using a generative neural network, the input video and the input query to generate an output comprising a modified video, with the modification applied to the text displayed in the one or more video frames.
24. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations comprising: receiving an input video; receiving an input query specifying a text-editing task to be performed on the input video, the text-editing task requiring a modification to text displayed in one or more videoframes of the input video; and processing, using a generative neural network, the input video and the input query to generate an output comprising a modified video, with the modification applied to the text displayed in the one or more video frames.
Citation Information
Patent Citations
Facilitating Text Identification and Editing in Images
US20170124417A1
Systems and methods for generating personalized videos with customized text messages
US20200234483A1
Image generation using one or more neural networks
US20220114698A1
Text-driven editor for audio and video assembly
US20220130427A1
Method and system for removing scene text from images
US20220189088A1