Generative text replacement in video

CN122847718APending Publication Date: 2026-09-29GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480088968.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2026-09-29

Smart Images

  • Figure CN122847718A_ABST
    Figure CN122847718A_ABST
Patent Text Reader

Abstract

Methods, systems, and devices, including computer programs encoded on computer storage media, for generating modified video conditioned on an input video and an input query. In one aspect, a method includes receiving an input video, receiving an input query that specifies a text editing task to be performed on the input video, the text editing task requiring a modification to text displayed in one or more video frames of the input video, and processing the input video and the input query using a generative neural network to generate an output that includes a modified video having the modification applied to the text displayed in the one or more video frames.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This specification relates to using machine learning models to generate outputs based on inputs.

[0002] Machine learning models receive input and generate outputs, such as predicted outputs, based on the received input. Some machine learning models are parametric models, and they generate outputs based on the received input and the values ​​of the model's parameters.

[0003] Some machine learning models are deep models, which employ multiple layers of operations to generate outputs from the received inputs. For example, deep neural networks are deep machine learning models that include an output layer and one or more hidden layers, each of which applies a nonlinear transformation to the received input to generate an output. Summary of the Invention

[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations, which generates a modified video based on an input video and an input query specifying a text editing task to be performed on the input video.

[0005] According to a first aspect, a method is provided, the method comprising: receiving an input video; receiving an input query specifying a text editing task to be performed on the input video, the text editing task requiring modification of text displayed in one or more video frames of the input video; and using a generative neural network to process the input video and the input query to generate an output including a modified video having modifications applied to the text displayed in the one or more video frames.

[0006] In some implementations, receiving input video includes receiving input video from the user.

[0007] In some implementations, receiving input queries includes receiving input queries from the user.

[0008] In some implementations, the output further includes a status indicator for the modification.

[0009] In some implementations, a status indicator indicates whether the modification was successfully applied, or the position of the text displayed in one or more video frames of the input video.

[0010] In some implementations, the method further includes providing a modified video for display to the user.

[0011] In some implementations, the text displayed in one or more video frames of the input video includes any or more of the following: overlaid text, in-scene text, or animated text.

[0012] In some implementations, text editing tasks include one or more of the following: modifying one or more visual features of the text, removing the text, or replacing the text with alternative text that shares visual features with the original text.

[0013] In some implementations, the input query specifies i) the text to be replaced and ii) the alternative text.

[0014] In some implementations, different colors are used to specify the text to be entered as a query.

[0015] In some implementations, the input query specifies i) the text to be displayed in a different color and ii) that different color.

[0016] In some implementations, the input query specifies the text to be removed.

[0017] In some implementations, the input query specifies i) the text to be displayed at different sizes and ii) a definition representing the size change of the specified text.

[0018] In some implementations, the generative neural network has been pre-trained to generate videos conditioned on at least video and text.

[0019] In some implementations, the generative neural network has been fine-tuned on a training dataset used for multiple text editing tasks, each of which requires modifying the pixels depicting text in the input video.

[0020] In some implementations, the generative neural network has been sequentially fine-tuned on each of the multiple text editing tasks in the training dataset.

[0021] In some implementations, the training dataset includes multiple training examples, each including training input and corresponding training output. The training input includes training videos and training queries specifying text editing tasks from multiple text editing tasks, and the corresponding training output includes at least the output video.

[0022] In some implementations, the generative neural network has been further fine-tuned to minimize the loss function based on the aggregate reward value.

[0023] In some implementations, the aggregated reward value is derived from one or more reward values, each of which is generated by a corresponding reward model.

[0024] In some implementations, each corresponding reward model is configured to generate a reward value representing the evaluation of a given video based on the corresponding comparison cue and the original video.

[0025] In some implementations, the corresponding comparison prompts include questions about the visual similarity, layout, style, font, weight, or color of text displayed in one or more video frames of a given video compared to text displayed in one or more video frames of the original video.

[0026] In some implementations, each corresponding reward model is trained to generate a reward value for the corresponding comparison cue.

[0027] According to another aspect, a system is also described, comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the methods described herein.

[0028] According to another aspect, one or more computer-readable storage media are also described, which store instructions that, when executed by one or more computers, cause one or more computers to perform the methods described herein.

[0029] Specific embodiments of the subject matter described in this specification may be implemented in order to achieve one or more of the following advantages.

[0030] The system described in this specification can perform various text editing tasks for replacing or modifying text on screen. For example, the system can receive an input video and an input query specifying a text editing task to be performed on the input video. The text editing task requires modification of text displayed in one or more video frames of the input video. The system can generate a modified video with the modifications required by the text editing task by processing the input video and the input query using a generative neural network. The system can use a generative neural network to generate a modified video with modified text while preserving the underlying visual content of the input video, as well as the appearance of the text displayed in the input video relative to the underlying visual content of the input video. The text can be, for example, overlaid text, in-scene text, and / or animated text. The system can achieve seamless and natural modification of different types of text displayed in the video.

[0031] For example, the system can perform text editing tasks such as erasing text. The system can feed an input video and an input query to a generative neural network, fine-tuned on a training dataset for text editing tasks, which includes training examples demonstrating text erasure. For instance, the input query can identify the text to be erased in the input video. The system can identify the pixels representing the identified text and perform fill or repair on the pixels using the content of the underlying visual content or background. The underlying visual content or background may potentially be uneven across video frames. Therefore, the output video can include a reconstruction of the pixels of the visual content of the input video, which are exposed due to the erasure of the pixels of the text displayed in the input video.

[0032] As another example, the system can perform text editing tasks such as on-screen text replacement. The system can feed an input video and an input query to a generative neural network, which has been fine-tuned on a training dataset for text editing tasks, including training examples displaying the replacement text. The input query can identify the text to be replaced in the input video, as well as alternative text to replace the text in the input video.

[0033] The system can identify the original text to be replaced, perform padding or repair to erase pixels of the original text, and generate pixels representing the alternative text in the output video. The output video can display the alternative text, and the alternative text can share visual features such as font, color, shadow, position, and / or style with the text displayed in the input video.

[0034] In some examples, the system can alter the layout of text relative to the underlying visual content to accommodate shorter or longer alternative text. The system can also apply visual features of the original text, such as geometric or lighting distortion, to the alternative text.

[0035] In some examples, the system can perform on-screen text replacement to localize text displayed in videos and images. For instance, the system can replace text displayed in English in an input video so that the modified video displays the text in Spanish. The input query can identify the English text to be replaced, as well as the alternative Spanish text to replace the English text. The system can recreate the text style of the original text using potentially different alphabets with different characters. On-screen text replacement can also be used to correct errors in the text of an input video or to change sensitive content in the input video.

[0036] In some examples, the system can perform animated text modification by animing the movement and transformation patterns of the input video over time across video frames, based on these patterns. The system can feed the input video and an input query to a generative neural network fine-tuned on a training dataset for text editing tasks, which includes training examples displaying animated text. The input query identifies the text to be modified in the input video. The system can identify the pixels representing the identified text and perform on-screen text replacement while preserving the animated features of the identified text in the input video.

[0037] Conventional techniques for modifying text displayed in a video may fail to produce output videos with seamless and natural text modification. For example, some conventional techniques can handle text independently of the surrounding visual context or underlying visual content. For instance, when replacing text displayed within a street sign with alternative text, the alternative text may be longer than the original. Conventional techniques may ignore the boundaries of street signs, resulting in videos with low visual quality and inconsistency.

[0038] The system described in this specification can replace original text with alternative text while preserving the appearance of the original text relative to the underlying visual content. For example, the system can replace original text with alternative text that has the original font style and color but is of a different size and / or layout, such that the alternative text is confined within the boundaries of a street sign. The system can use a generative neural network that has learned to enforce visual and spatial constraints, such as displaying text within a street sign or within a video frame, through fine-tuning on a training dataset used for text editing tasks.

[0039] Conventional techniques for modifying text displayed in a video have limitations that make them difficult or impractical to use. For example, some techniques require users to manually identify and select the pixels representing the text displayed in a video frame, and to manually fill in or modify the pixels representing the text, or modify other pixels to represent the text.

[0040] The system can apply different types of modifications specified by the input query to the input video without requiring further user input. The system can determine which pixels represent text, and which subset of pixels represents the text specified in the input query, without user input. The system can also determine which pixels to modify to apply the modifications specified in the input query, such as which pixels to fill or which pixels to modify to represent text, without user input. The system requires no further input beyond the input query and the input video, resulting in a more efficient user experience.

[0041] Some conventional techniques require generating entirely new videos that include modified text, which may require users to manually describe all aspects of the desired video and could potentially require more computational resources for generation. Other techniques require sequences of machine learning models that could result in lower-quality output videos.

[0042] The system described in this specification can apply modifications specified by an input query to an input video using a generative neural network that performs discriminative and generative operations. For example, the generative neural network can be fine-tuned to perform multiple text editing tasks that require discriminating pixels containing text and / or generating an output video with modifications to those pixels. Therefore, the system can generate a video with modifications applied to the text without generating a completely new video based on text prompts describing the video and text, thus saving computation time and resources. The system can further save computation time and resources by focusing computation on the pixels affected by the modifications of the text editing tasks. The system can also use a single generative neural network to generate a video with modifications applied to the text, resulting in a more seamless modified video compared to modified videos generated using sequences of machine learning models.

[0043] Furthermore, in some implementations, the system described in this specification can generate modified videos with natural and seamless modifications based on human feedback. For example, a generative neural network can be fine-tuned by a training system to generate modified videos representing human preferences for multiple aspects, such as layout, style, color, or other visual features. The generative neural network can be fine-tuned to minimize a loss function based on reward values. Reward values ​​can be generated by a reward model trained on human feedback. Generating reward values ​​from a reward model can be more efficient than obtaining reward values ​​from human raters. Furthermore, the reward model can predict subtle nuances of human preferences that may not be represented in the training dataset. For example, generating a large number of high-quality training outputs to be used in training examples, particularly for in-scene text, can be computationally expensive. For example, each training example includes a training video with in-scene text having attributes such as lighting, orientation, tilt, distortion, and occlusion. The training output for each training example would require applying modifications specified by the training query while maintaining the attributes of the in-scene text displayed in the training video, which could be computationally expensive in terms of generation. Generating high-quality training outputs may also be infeasible. Furthermore, the training dataset may not include exhaustive examples of complex attributes and / or animation patterns. Therefore, the system can use a reward model to predict human preferences for the output of the generative neural network, and thus improve the performance of the generative neural network without requiring high-quality training outputs for the training dataset.

[0044] The training system can also fine-tune the generative neural network in a computationally efficient manner. For example, it can fine-tune a pre-trained generative neural network on a training dataset used for text editing tasks. The generative neural network may have been pre-trained to generate videos conditionally based on both video and text. The training system can also generate reward values ​​computationally efficiently based on human feedback for further fine-tuning. For example, it can derive a reward model from the fine-tuned generative neural network.

[0045] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. Attached Figure Description

[0046] Figure 1 This is a diagram illustrating an example process for generating modified videos.

[0047] Figure 2 This is a diagram of an example system used to generate modified videos.

[0048] Figures 3A to 3C The example video shows the input video and sample frames of the modified video for the example text editing task.

[0049] Figure 4 This is a flowchart of an example process for generating modified videos.

[0050] Figure 5 This is a flowchart illustrating an example process for fine-tuning a generative neural network.

[0051] Figure 6 An example rendering of the user interface used to collect human feedback is shown.

[0052] Figure 7 This is a diagram illustrating an example process for fine-tuning a generative neural network.

[0053] In the various figures, the same reference numerals and names indicate the same elements. Detailed Implementation

[0054] Figure 1 This is a diagram illustrating an example process 100 for generating modified video. For convenience, process 100 will be described as being executed by a system of one or more computers located in one or more locations. For example, a system for generating modified video appropriately programmed according to this specification (e.g., Figure 2 The system 200) can execute process 100.

[0055] System 200 can receive input video 104. Input video 104 can display raw text and other underlying visual content, such as scenes or backgrounds. Figure 1 The example input video 104 is shown in the video frame. Figure 1 In the example, input video 104 displays the text "Foundation models are cool!". Input video 104 can display text using certain visual features (such as style, color, font, etc.). In some examples, input video 104 may not display text, display additional text, or display other text.

[0056] System 200 can receive input query 106. Input query 106 can specify the type of text editing task and / or the text to be modified from input video 104. Figure 1 In the example, input query 106 includes "Replace the text in the video: "Foundation models are cool!" with the text: "¡ Los modelos de baseson geniales!" (Replace the text in the video: "Foundation models are cool!" with the text: "¡ Losmodelos de base son geniales! (The basic models are great!)"). Input query 106 specifies the type of text editing task: text replacement. Input query 106 also specifies the text to be modified in input video 104: "Foundation modelsare cool!". Input query 106 also specifies the alternative text: "¡ Los modelos de base son geniales!". In some examples, input query 106 may specify the following references. Figures 3A to 3C Other types of text editing tasks described.

[0057] System 200 can be referenced as follows: Figure 2 The modified video 180 and status indicator 112 are described. Figure 1 The example modified video 180 is shown. Figure 1 In the example, system 200 can apply the modifications of input query 106 to the modified video 180. Therefore, Figure 1The modified video 180 displays "¡Los modelos de base son geniales!" at the location of "Foundation models are cool!". The modified video 180 can display the alternative text with the same visual characteristics as the text in the input video 104. In some examples, because the alternative text includes different numbers of characters, different numbers of words, or words with different numbers of characters, the modified video 180 can display the alternative text with a different layout or size compared to the original text in the input video 104.

[0058] Status indicator 112 indicates whether the modifications were successfully applied. Figure 2 In some examples, the status indicator 112 may include text such as "Text found". In some examples, the input video 104 may not include the text "Foundation models are cool!". The status indicator 112 may include text such as "Text not found".

[0059] Status indicator 112 can also indicate the position of text displayed in one or more video frames of input video 104. For example, the position of the text can be represented by a bounding box. Figure 2 In one example, the status indicator 112 may include text such as "Text at: (10, 20) to (100, 400)". In some examples, the position of the text may change between video frames of the input video 104. The position of the text may be represented by the bounding box of the first video frame of the input video 104 in which the text is displayed. As another example, the position of the text may be represented by data representing a fine-grained mask of pixels displaying the text.

[0060] See below for reference Figure 5 The description, such as reference Figure 1 The training system of the described training system 140 can use the state indicator 112 for each training example as part of the calculation of the loss value of that training example.

[0061] In some examples, status indicator 112 can also indicate the time when the text is displayed in one or more video frames of the input video. For example, the time when the text is displayed can be represented by a timestamp indicating the time elapsed from the beginning of the input video to the point where the text is displayed. Figure 1In some examples, the status indicator 112 may include text such as "Text at 9:59 seconds". In some examples, the text may be displayed within a time interval of the input video 104. The time when the text is displayed may be indicated by the timestamp of the first video frame in which the text is displayed in the input video 104.

[0062] Figure 2 This is a diagram of an example system 200 for generating a modified video 280 comprising multiple video frames. System 200 includes multiple components, such as a generative neural network 210. System 200 is an example of a system implemented as a computer program on one or more computers at one or more locations, in which the systems, components, and techniques described below are implemented.

[0063] To generate the modified video 280, system 200 receives input video 204 and input query 206. In some implementations, system 200 may receive input video 204 and / or input query 206 from a user.

[0064] The input video 204 consists of multiple video frames. Each video frame is an image comprising multiple pixels, each pixel having one or more intensity values. Each video frame can display a visual scene, such as a scene or background.

[0065] Some of the video frames can also display text. The text can be overlaid text, in-scene text, and / or animated text.

[0066] Overlay text can include text that is not part of the scene and is added to a video frame. For example, overlay text can be added to a video frame after the scene's video has been recorded or generated. As one example, the video frame can display the scene with overlay text displayed on top of the scene. As another example, the video frame can display the scene and overlay text in a split-screen manner. For example, the video frame can display text next to or around the scene. Examples of overlay text are described below. Figures 3A to 3B and Figure 4 As shown in the image.

[0067] In-scene text can include text that appears within the scene. For example, text can be displayed in the environment where video is being recorded. Examples of in-scene text are provided below. Figure 3A As shown in the image.

[0068] Animated text can include text that moves and / or transforms between video frames. For example, animated text can be added to video frames after the video has been recorded or generated. Examples of transformations in animated text include overlay text or in-scene text that changes position, size, appearance, color, transparency, and / or orientation across video frames.

[0069] Input query 206 specifies the text editing task to be performed on input video 204. For example, input query 206 can specify the type of text editing task. In some examples, input query 206 can specify the text to be modified in input video 204. Example text editing tasks (such as modifying the visual features of text, removing text, and / or replacing text with alternative text) are referenced below. Figures 3A to 3C Describe it.

[0070] System 200 uses a generative neural network 210 to generate a modified video 280 and a status indicator 212 for the modified video 280, as described in further detail below.

[0071] System 200 can then provide, for example, a modified video 280 and / or a status indicator 212 for display to the user via a user interface.

[0072] Status indicator 212 includes text indicating whether the modification specified by input query 206 has been successfully applied. As an example, system 200 may provide input video 204 and input query 206 as input to generative neural network 210 to generate modified video 280, but generative neural network 210 may output status indicator 212 indicating that the text specified by input query 206 was not detected in input video 204. For example, status indicator 212 may indicate that no text was found. System 200 may provide status indicator 212, for example, through a user interface for display to a user. In some implementations, system 200 may not output modified video 280. In some implementations, system 200 may output input video 204 as modified video 280.

[0073] As another example, system 200 can provide input video 204 and input query 206 as input to generative neural network 210 to generate modified video 280, and generative neural network 210 can output state indicator 212 indicating that text specified by input query 206 has been detected in input video 204. Generative neural network 210 can also output modified video 280 with modifications applied to the text.

[0074] For example, state indicator 212 may indicate that text has been found. In some examples, state indicator 212 may also indicate the position of the text displayed in one or more video frames of the input video. During the training of the generative neural network 210, the training system 240 described below can use the position of the text displayed for each training example to calculate the loss for that training example. In some examples, state indicator 212 may also indicate the time when the text is displayed in one or more video frames of the input video. For example, state indicator 212 may include a timestamp indicating the time elapsed from the beginning of the input video to the point when the text is first displayed.

[0075] Generative neural network 210 is a generative neural network configured to generate videos conditioned on at least video and text. Generative neural network 210 may have been pre-trained to generate videos conditioned on video and text, and fine-tuned on training datasets for multiple text editing tasks, each requiring modification of pixels depicting text in the input video.

[0076] The generative neural network 210 can have any suitable architecture that allows it to generate videos conditioned on at least video and text. For example, the generative neural network 210 can be a diffusion model, a Transformer-based model, or an autoregressive model.

[0077] The generative neural network 210 is pre-trained, for example, by training system 240 or by one or more other systems. For example, training system 240 may pre-train the generative neural network 210 on a conditional video generation task. As a particular example, the generative neural network 210 may be pre-trained to reverse the forward diffusion process on a dataset of videos and corresponding text.

[0078] The training system 240 within system 200 or another training system can fine-tune the generative neural network 210. That is, the training system 240 can obtain a pre-trained generative neural network 210, and can then fine-tune the generative neural network 210 starting from the pre-trained values ​​of the network parameters of the generative neural network 210.

[0079] The training system 240 can fine-tune the generative neural network 210 on the training dataset 244. The training dataset 244 may include multiple training examples corresponding to different text editing tasks.

[0080] Alternatively or additionally, the training system 240 may fine-tune the generative neural network 210 to minimize a loss function based on aggregated reward values ​​generated by a set of reward models 242. Each reward model in the reward models 242 may be configured to generate a reward value representing an evaluation of the modified video 280 generated by the generative neural network 210 compared to the input video 204.

[0081] By using a pre-trained generative neural network 210, system 200 can generate a modified video 280 with natural and coherent text modifications. For example, the pre-trained generative neural network 210 may have learned to associate objects depicted on top of video frames. For instance, input video 204 may show a bag of chips moving across multiple video frames, with reflection distortion and other visual changes across these frames. The pre-trained generative neural network 210 can be trained to understand that the same bag of chips is moving across video frames, rather than processing the pixels of each video frame independently. By fine-tuning the pre-trained generative neural network 210, system 200 can use it to generate the modified video 280, which has contextual modifications to the text that appear natural and fit the scene.

[0082] As another example, the pre-trained generative neural network 210 may have learned to recognize faces or underlying visual content such as signs. The system 200 can fine-tune the pre-trained generative neural network 210 to follow visual constraints, such as not positioning text above a person's face, positioning text relative to the underlying visual content, or ensuring that all text content is displayed within a video frame.

[0083] Training is referenced below. Figures 5 to 7 To describe in more detail.

[0084] Figures 3A to 3C Example frames 310-360 of the input video and the modified video are shown for the example text editing task. For example, refer to... Figure 3A Frames 310a and 310b depict a text editing task involving modifying the visual features of text, specifically the color of the text. Frame 310a is an example frame of an input video depicting a scene with six movie characters. Frame 310a displays the text 311 “This is from a movie.” Text 311 is displayed in black. Text 311 is an example of overlaid text that is not part of the scene.

[0085] The system can receive input video including frame 310a and input query, and provide the input video and input query to a generative neural network (as referenced above). Figures 1 to 2(As described) to generate a modified video including frame 310b. Input queries can specify different colors for the text in the input video. In the examples of frames 310a and 310b, the input query could include "Change the color of the text in the video to gray".

[0086] Frame 310b is an example frame of the modified video generated by the system. Frame 310b depicts the same scene as frame 310a, where the modifications specified by the input query above are applied to text 311. Frame 310b displays text 312, which has the same content as text 311, but is displayed in a different color than text 311. For example, Figure 3A The example text 312 is displayed in gray.

[0087] Frames 320a and 320b depict a text editing task involving modifying the visual features of text, specifically changing the color of specified text. Frame 320a is an example frame of an input video depicting a scene with two movie characters. Frame 320a displays the text 321 “Scene from a film”. Text 321 is displayed in black. Text 321 is an example of overlaid text that is not part of the scene. Frame 320a also displays the text 325 “STOP” on a street sign in the background. Text 325 is an example of text within the scene.

[0088] The system can generate a modified video, including frame 320b, by receiving the input video and input query, including frame 320a, and by providing the input video and input query to a generative neural network (as described above). The input query can specify the text to be displayed in a different color and that different color. In the examples of frames 320a and 320b, the input query could include "Change the color of the text in the video: 'a film' to gray". In some examples, the case of the text specified in the input query can be different from that of text 322. For example, the input query could include "Change the color of the text in the video: 'A Film' to gray".

[0089] Frame 320b is an example frame of the modified video generated by the system. Frame 320b depicts the same scene as frame 320a, where the modifications specified by the input query above are applied to text 321. Frame 320b displays text 322, which has the same content as text 321, but a portion of text 322 is displayed in a different color than text 311. For example, the word "afilm" in text 322 is displayed in gray. The other words in text 322, "Scene from (the scene in...)," are the same color as in text 321. Text 325 remains the same in frame 320b as it is in frame 320a.

[0090] Another example of modifying the visual characteristics of text includes changing the text size, as described in reference frames 340a and 340b below. Other examples of modifying the visual characteristics of text include changing the text weight, font, or stroke thickness.

[0091] refer to Figure 3B Frames 330a and 330b depict a text editing task involving erasing specified text. Frame 330a is an example frame of the input video depicting a scene with two movie characters. Frame 330a displays text 331, “Something to delete.” Text 331 is an example of overlaid text that is not part of the scene. Frame 330a also displays text 335, “Other text to ignore.” Text 335 is also an example of overlaid text.

[0092] The system can generate a modified video, including frame 330b, by receiving the input video and input query, including frame 330a, and by providing the input video and input query to a generative neural network (as described above). The input query can specify text to be erased or removed. In the examples of frames 330a and 330b, the input query could include “Erase the text in the video: “Something to delete””. In some examples, the case of the text specified in the input query can differ from that of text 331. For example, the input query could include “Erase the text in the video: “something to delete””.

[0093] Frame 330b is an example frame of the modified video generated by the system. Frame 330b depicts the same scene as frame 330a, with the modifications specified by the input query above applied to text 331. Frame 330b does not display any text corresponding to text 331, nor does it display any alternative text in its position. The pixels in frame 330b corresponding to the pixels of text 331 in frame 330a have been padded or repaired according to the underlying scene. Text 335 remains the same in frame 330b as it is in frame 330a.

[0094] Frames 340a and 324b depict a text editing task involving modifying the visual features of text, specifically modifying the size of specified text. Frame 340a is an example frame of an input video depicting a scene with a movie character. Frame 340a displays the text 341 “A scene from a movie.” Text 341 is an example of overlaid text that is not part of the scene. Frame 340a also displays text 341 as adhering to certain visual constraints, such as not obscuring the face of the movie character and fitting within frame 340a.

[0095] The system can generate a modified video, including frame 340b, by receiving the input video and input query, including frame 340a, and by providing the input video and input query to a generative neural network (as described above). The input query can specify text to be displayed at different sizes and a definition for the size change. For example, the size change can be defined by the type of change (such as size increase or size decrease) and the magnitude of the change (such as a percentage or proportion relative to text 341). In the examples of frames 340a and 340b, the input query could include “Increase the size of the text in the video: “A scene from a movie” by 10%”. In some examples, the capitalization of the text specified in the input query can differ from that of text 341. For example, the input query could include “Increase the size of the text in the video: “A SCENE from A movie” by 10%”.

[0096] Frame 340b is an example frame of the modified video generated by the system. Frame 340b depicts the same scene as frame 340a, where the modifications specified by the input query above are applied to text 341. Frame 340b displays text 342, which has the same content and style as text 341, but is displayed at a different size. For example, text 342 is displayed at a larger size than text 341. Furthermore, because text 342 is larger, it is also displayed with increased pixel detail compared to text 341. Frame 340b also displays text 342 in accordance with the same visual constraints as text 341. That is, frame 340b also displays text 342 in a way that does not obscure the faces of the characters in the film and is displayed as if it fits within frame 340b.

[0097] Furthermore, text 342 has a different layout. Text 341 is divided into two lines: “A scene” and “from a movie”. Text 342 is divided into three lines: “A scene”, “from”, and “a movie”. If text 342 has a larger size but the same layout as text 341 and follows the same visual constraint as text 341 of not covering the faces of movie characters, then text 342 may not fit in frame 340b. This could violate another visual constraint of fitting text within a frame.

[0098] Frames 350a and 350b depict a text editing task where text is replaced with alternative text. Frame 350a is an example frame of the input video depicting a scene with an observatory and shapes in the sky. Frame 350a displays the text 351 “EVERYWHEREYOU LOOK WEIRD THINGS IN THE SKY”. Text 351 is displayed in black with a gray outline. Text 351 is an example of overlaid text that is not part of the scene.

[0099] The system can generate a modified video, including frame 350b, by receiving the input video and input query, including frame 350a, and by providing the input video and input query to a generative neural network (as described above). The input query can specify the text to be replaced and the alternative text. In the examples of frames 350a and 350b, the input query could include “Replace the text “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY” with the text “ÜBERALLMERKWÜRDIGE DINGE AU HIMMEL”. In some examples, the capitalization of the text specified in the input query can differ from that of text 351. For example, an input query could include “Replace the text “everywhere you look weird things in the sky” with the text “ÜBERALLMERKWÜRDIGE DINGE AU HIMMEL”.

[0100] Frame 350b is an example frame of a modified video generated by the system. Frame 350b depicts the same scene as frame 350a, where the modifications specified by the input query above are applied to text 351. Frame 350b displays text 352, which has the same visual characteristics (such as style and placement) as text 351, but with different content. For example, text 352 is displayed in black with a gray outline. Text 352 includes “ÜBERALL MERKWÜRDIGE DINGE AU HIMMEL” instead of “EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY”. In this example, text 352 includes characters not present in text 351, such as “Ü”.

[0101] refer to Figure 3CFrames 360a and 360b are another example of a text editing task where text is replaced with alternative text. Frame 360a is an example frame of an input video depicting a scene with four characters and a raccoon. Frame 360a displays the text 361, “There was a noisy raccoon in our hotel!”. Text 361 is an example of overlaid text that is not part of the scene. Text 361 is displayed to follow certain visual constraints, such as fitting within text box 365.

[0102] The system can generate a modified video, including frame 360b, by receiving the input video and input query, including frame 360a, and by feeding the input video and input query to a generative neural network (as described above). The input query can specify the text to be replaced and the alternative text. In the examples of frames 360a and 360b, the input query could include “Replace the text “There was a noisy raccoon in our hotel!” with the text “¡ Habia unmapache ruidoso en nuestro hotel!”. In some examples, the capitalization of the text specified in the input query can differ from that of text 361. For example, the input query could include “Replace the text “there was a noisy raccoon in our hotel!” with the text “¡ Habia un mapache ruidoso en nuestro hotel!”.

[0103] Frame 360b is an example frame of a modified video generated by the system. Frame 360b depicts the same scene as frame 360a, where the modifications specified by the input query above are applied to text 361. Frame 360b displays text 362, which has the same visual characteristics (such as style) and follows visual constraints as text 361, but has different content and typographical elements. For example, text 362 is also displayed within text box 365. Text 362 includes “¡Habia un mapache ruidoso ennuestro hotel!” instead of “There was a noisy raccoon in our hotel!”. In this example, text 362 includes characters not present in the original text 361, such as “¡”.

[0104] Figure 4 This is a flowchart of an example process for generating a modified video. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system for generating a modified video appropriately programmed according to this specification (e.g., Figure 2 The system 200) can execute process 400.

[0105] The system receives input video (step 410). For example, the system may receive input video from a user. The input video may include multiple video frames. One or more video frames may display text.

[0106] The system receives an input query (step 420). The input query can specify the text editing task to be performed on the input video. In some implementations, the system can receive the input query from the user.

[0107] For example, an input query can include a natural language description of how the text displayed in the input video should be modified. In some examples, the input query may also specify the text to be modified and / or alternative text to replace the text to be modified.

[0108] In some examples, the input query can specify a text editing task that modifies all text displayed in the input video. For instance, the input query could include a description of the modifications to be made, followed by "text in the video," without identifying the specific text to be changed.

[0109] In some examples, the input query can specify a text editing task to modify specific text displayed in the input video. The input query can identify the text to be modified displayed in one or more video frames of the input video. For example, the input query can use specific characters or symbols to identify the text. For example, the text to be modified can be enclosed in quotation marks. As an example, the input query can include a description of the modification to be made, followed by "text: 'noisy raccoon'". In some examples, the input query can specify the text to be modified using a different capitalization than the text to be modified displayed in one or more video frames of the input video. For example, the input query can include a description of the modification to be made, followed by "text: 'Noisy raccoon'".

[0110] Alternatively or additionally, the input query can identify the text using the time when the text is displayed or the time interval during which the text is displayed in the input video. For example, the text to be modified may be displayed in the input video at one or more timestamps or during a time interval. As an example, the input query could include "From 10:03 to 10:17 seconds in the video," followed by a description of the modifications to be made to the text displayed in the input video. As another example, the input query could include "Starting at 10:03 seconds in the video," followed by a description of the modifications to be made to the text displayed in the video.

[0111] As another example, input queries can use quotation marks and identify text by the time interval in which the text is displayed in the input video. For example, an input query could include "From 10:03 to 10:17 seconds in the video", followed by a description of the modification to be made, and further followed by the text "noisy raccoon".

[0112] In some examples, the input query can specify a text editing task that will replace specific text displayed in the input video with alternative text. The input query can also use specific characters or symbols to identify the alternative text. For example, the alternative text can be enclosed in quotation marks. As an example, the input query could include the description "the text: "noisy raccoon" with the text: "un mapache ruidoso" (replacing the text "noisy raccoon" with the text "un mapache ruidoso")". In some examples, the input query can specify the text to be replaced using a different capitalization than the text to be replaced displayed in one or more video frames of the input video. For example, an alternative input query to the example input query above could include the description "replacing the text: "Noisy raccoon" with the text: "un mapache ruidoso"".

[0113] Alternatively or additionally, as described above, the input query can identify text using the time when the text is displayed or the time interval during which the text is displayed in the input video. For example, the text to be replaced may be displayed in the input video at one or more timestamps or during a time interval. As an example, the input query could include "From 10:03 to 10:17 seconds in the video," followed by a description that replaces the specified text with the alternative text. As another example, the input query could include "Starting at 10:03 seconds in the video," followed by a description that replaces the specified text with the alternative text.

[0114] As another example, input queries can use quotation marks and identify text by the time interval in which the text is displayed in the input video. For example, an input query could include "From 10:03 to 10:17 seconds in the video", followed by the description "replacing the text: "Noisy raccoon" with the text: "un mapacheruidoso".

[0115] In some implementations, the system can generate input queries based on user input. For example, the system can provide a user interface that allows users to play and pause the input video. The user interface can also allow users to input a natural language description of how the text displayed in the input video should be modified. For example, a user can input a natural language description such as "Replace the text with 'Foundation models are cool!'". The user can pause the input video at the 10:05 timestamp. The system can determine that at 10:05 in the input video, the text displayed is "¡Los modelos de baseson geniales!".

[0116] The system can identify the time interval in which the text “¡Los modelos de base son geniales!” is displayed in the input video and includes a timestamp of 10:05. The system can use, for example, a machine learning model configured to identify the time interval in which a given video displays a given text, where the time interval includes the given timestamp. For example, the system can determine that the text “Foundation models are cool!” displayed in the input video at 10:05 is displayed from 10:02 to 10:19. The system can include the time interval in the input query. For example, the system can generate an input query to include “Replace the text in the video from 10:02 to 10:19 with “¡Los modelos debase son geniales!””.

[0117] Text editing tasks may require modifying text displayed in one or more frames of the input video. The text can be, for example, overlaid text, in-scene text, or animated text. Example text editing tasks are described below.

[0118] For example, a text editing task could include replacing text with alternative text that shares visual features with the original text. The input query can specify the text to be replaced and the alternative text. An example of replacing text with alternative text is provided in the reference above. Figure 3B Frames 350a and 350b and Figure 3C Frames 360a and 360b are described.

[0119] As another example, a text editing task may include modifying one or more visual features of the text. For example, visual features may include color, size, font, and font weight. A text editing task may also include modifying the visual features of all text displayed in the input video. For example, the input query may specify different colors for the text. An example of modifying the color of all text displayed in the input video is referenced above. Figure 3AFrames 310a and 310b are described.

[0120] As another example, a text editing task could include changing the size of all text displayed in the input video. The input query can specify a definition representing the size change of the text displayed in the input video.

[0121] Text editing tasks can also include selectively modifying the visual characteristics of text displayed in the input video. For example, the input query could specify text to be displayed in a different color and that different color. An example of selectively modifying the color of text displayed in the input video is referenced above. Figure 3A Frames 320a and 320b are described.

[0122] As another example, a text editing task could include selectively changing the size of text displayed in an input video. For instance, an input query could specify text to be displayed at different sizes and a definition representing the size change of the specified text. Examples of selectively changing the size of text displayed in an input video are referenced above. Figure 3B Frames 340a and 340b are described.

[0123] As another example, text editing tasks can include removing or erasing text. For instance, an input query can specify the text to be removed. An example of selectively removing text displayed in an input video is referenced above. Figure 3B Frames 330a and 330b are described.

[0124] As another example, a text editing task could include removing or erasing all text displayed in the input video. The input query can specify whether the text in the video should be removed or erased.

[0125] The system uses a generative neural network to process the input video and input query to generate an output that includes a modified video (step 430). The modified video may include modifications applied to text displayed in one or more video frames.

[0126] For example, a generative neural network can use, for instance, a text encoder neural network to process an input query to generate an encoded representation of the input query. For example, the encoded representation may include the terms of the input query. As used herein, "encoded representation" or "embedding" is a vector of numerical values ​​(e.g., floating-point values ​​or other values) having a predetermined dimension. The space of possible vectors having a predetermined dimension is called the "embedding space." A generative neural network can use, for instance, an encoder neural network to process an input video to generate an encoded representation of the input video. For example, an encoder neural network can process the intensity values ​​of pixels in the input video. For example, the encoded representation may include the terms of the input video.

[0127] Generative neural networks can generate one or more video frames of an output video conditioned on an encoded representation of an input query and an encoded representation of an input video. For example, a generative neural network can generate an input word sequence from the encoded representation of the input query and the encoded representation of the input video. A generative neural network can process the input word sequence to generate an output including a modified video. For example, a generative neural network can autoregressively generate one or more video frames of a modified video. That is, a generative neural network can autoregressively generate an output word sequence by generating each word in the output sequence conditioned on the current input sequence, which includes (i) the input word sequence followed by (ii) any word preceding a specific word in the output sequence. A generative neural network can, for example, use a decoder neural network to generate one or more video frames from the output word sequence. As a specific example, a generative neural network can include an autoregressive Transformer-based neural network comprising multiple layers, each applying a self-attention operation. The neural network can have any of a variety of Transformer-based neural network architectures.

[0128] As another example, generative neural networks can generate one or more video frames of a modified video by reversing the forward diffusion process conditioned on the representation of the input query and the representation of the input video.

[0129] In some implementations, the output may also include a status indicator for the modification. The status indicator can indicate whether the modification was successfully applied. For example, a generative neural network might not detect any text in the input video or the text specified by the input query. The generative neural network can generate an output that includes a status indicator indicating that the modification was not successfully applied. For example, the output could include "Text not found". The status indicator can be used during the fine-tuning of the generative neural network, as referenced below. Figure 5 As described.

[0130] The status indicator can also indicate the position of text displayed in one or more video frames of the input video. For example, the position of the text can be represented by a bounding box. The bounding box can include a rectangle defining the text displayed in one or more video frames of the input video, and can be defined by pixel coordinates. For example, the bounding box can be defined by the coordinates of the top-left corner and the bottom-right corner of the rectangle.

[0131] In some examples, the position of the text varies across video frames. For instance, the text could be animated text displayed in different video frames with different positions, sizes, curls, opacities, colors, etc. The position of the text can be represented by the bounding box of the first video frame in which the text appears.

[0132] In some examples, the position of the text can be represented by data that is a fine-grained mask of pixels. For example, the data representing a fine-grained mask of pixels could include data indicating the position of the pixels displaying the text.

[0133] In some implementations, the system may further provide modified video for display to the user. For example, the system may provide data representing the modified video to the user interface for display.

[0134] Generative neural networks can be pre-trained to generate videos conditionally, at least from both video and text. In some implementations, the generative neural network has been fine-tuned on a training dataset used for multiple text editing tasks. Each text editing task may require modification of pixels depicting text in the input video. In some examples, the generative neural network has been sequentially fine-tuned on each of the multiple text editing tasks in the training dataset. See references below. Figure 5 A more detailed description of fine-tuning on the training dataset.

[0135] In some implementations, the generative neural network has been further fine-tuned to minimize the loss function based on the aggregate reward value. Fine-tuning based on the aggregate reward value is described in the following section. Figures 5 to 7 To describe in more detail.

[0136] Figure 5 This is a flowchart of an example process 500 for fine-tuning a generative neural network. For convenience, process 500 will be described as being executed by a system of one or more computers located in one or more locations. For example, a training system appropriately programmed according to this specification (e.g., Figure 2 The training system 240 can execute process 500 to train Figure 2 Generative neural network 210.

[0137] The training system can obtain data representing a pre-trained generative neural network (step 510). For example, the generative neural network may have been pre-trained to generate videos conditionally with at least video and text. For example, the training system can clone a version of a generative neural network that has been pre-trained to generate videos conditionally with video and text. Examples of generative neural networks pre-trained to generate videos conditionally with video and text include VideoPoet, Gen-1, and Gen-2.

[0138] The training system can obtain a training dataset (step 520). For example, the training dataset could be a reference dataset. Figure 2The described training dataset is 244. The training dataset may include multiple training examples. Each training example may include a training video and a training query specifying a text editing task from multiple text editing tasks. Each training example may also include a corresponding training output, which at least includes an output video. In some implementations, the corresponding training output may also include a status indicator. For example, the status indicator may indicate whether any text is displayed in the training video, or whether the training video displays text specified by the training query. The status indicator may also indicate the position of the text displayed in the training video.

[0139] In some implementations, the training system can generate a training dataset. For example, the training system can receive or generate seed videos, each consisting of one or more video frames. The training system can then generate synthetic training videos and / or training outputs by adding text to the seed videos.

[0140] The training system can generate training videos and corresponding output videos that adhere to certain visual constraints. For example, visual constraints may include not positioning text above a person's face, positioning text relative to underlying visual content, or ensuring that all text content is displayed within video frames.

[0141] The training system can also generate different training videos and corresponding output videos that display text with different visual features (such as color, size, style, font, orientation, geometric distortion, lighting, pattern, position, layout, and animation).

[0142] The training system can also generate different training and output videos with different types of text characters. For example, the training system can generate different training and output videos displaying different words and characters (e.g., using different alphabets for different languages).

[0143] The training system can also generate training videos that do not display text or the text specified in the training query. The training system can also generate an expected status indicator indicating that no text was found.

[0144] As an example, a text editing task for a specific training example could be modifying the visual features of all text in an input video. For instance, the training query could specify a color, such as yellow. The training video could include a video displaying text in a first color, such as pink. The training output could include an output video displaying the same text content and other visual features as the text displayed in the training video, but in a second color (yellow).

[0145] The training system can generate specific training examples by obtaining a seed video that does not display text. It can generate a training video by creating a copy of the seed video and adding text that is pink and has certain other visual features. It can also generate an output video by creating a copy of the seed video and adding the same text content that is yellow and has the same other visual features. The training system can generate training queries that include, for example, "Change the color of the text in the video to yellow". The training system can also generate additional training examples. For example, it can use the training video of a specific training example as the output video of an additional training example, and vice versa. The training system can generate training queries for additional training examples that include, for example, "Change the color of the text in the video to pink".

[0146] The training system can generate different training examples for the same text. For example, it can generate different training videos displaying text in different colors. Specifically, it can generate training videos that display text using colors present in the scene of the training video or colors similar to those present in the scene of the training video. Therefore, the training system can fine-tune the generative neural network to determine which pixels correspond to the text, rather than learning to replace all pixels with a certain color.

[0147] As another example, a text editing task for a specific training example could be modifying the visual features of specified text in an input video. For instance, the training query could specify text to be displayed in different colors, such as yellow. The training video could include a video displaying text in a first color, such as pink. The training output could include an output video displaying the same text content and other visual features as the text displayed in the training video, but with the specified text displayed in a second color (yellow).

[0148] The training system can generate specific training examples by obtaining a seed video that may or may not display text. The training system can generate a training video by creating a copy of the seed video and adding text that is pink and has certain other visual features. The training system can generate an output video by creating a copy of the seed video and adding yellow and pink portions of the text content with the same other visual features. The training system can generate training queries to include, for example, "Change the color of the text 'a film' in the video to yellow". The training system can also generate additional training examples. For example, the training system can use the training video of a specific training example as the output video of an additional training example, and use the output video of a specific training example as the training video of an additional training example. The training system can generate training queries for additional training examples to include, for example, "Change the color of the text 'a film' in the video topink".

[0149] As another example, a text editing task for a specific training example could be erasing text from an input video. For instance, the training query could specify the text to be erased. The training video could include a video displaying the specified text. The training output could include an output video that does not display the specified text.

[0150] The training system can generate specific training examples by obtaining a seed video that may or may not display text. The training system can generate training videos by creating copies of the seed videos and adding text such as "Something to delete". The training system can also generate output videos by creating copies of the seed videos. The training system can generate training queries to include, for example, "Erase the text in the video: "Something to delete".

[0151] As another example, the text editing task for a specific training example could be changing the size of text in an input video. For instance, the training query could specify text to be displayed at different sizes, along with definitions for the size changes. The training video could include a video displaying the text. The training output could include an output video displaying the same text content and other visual features as the text displayed in the training video, but at a different size.

[0152] The training system can generate specific training examples by obtaining a seed video that may or may not display text. The training system can generate training videos by creating copies of the seed videos and adding text with certain visual features (such as following certain visual constraints). The training system can generate output videos by creating copies of the seed videos and adding the same text content at different sizes with other visual features that are the same as or similar to the text in the training videos. The training system can generate training queries to include, for example, "Increase the size of the text in the video: 'A scene from a movie' by 10%". The training system can also generate additional training examples. For example, the training system can use the training video of a specific training example as the output video of an additional training example, and use the output video of a specific training example as the training video of an additional training example. The training system can generate training queries for additional training examples to include "Increase the size of the text in the video: 'A scene from a movie' by 10%".

[0153] In some examples, the output video may display text with a different amount of pixel detail than the training video. For example, when displayed at a smaller size, the text may have less pixel detail. When displayed at a larger size, the text may have more pixel detail.

[0154] In some examples, the training system can generate corresponding output videos that display text content in different positions or layouts. The corresponding output videos should display the text content in roughly the same position as the training videos, but possibly in a different layout, and should adhere to visual constraints. For example, when displayed at a smaller size, some words of the text may be grouped into a single line. When displayed at a larger size, some words of the text may be separated into different lines.

[0155] As another example, a text editing task for a specific training example could be replacing specified text with alternative text. For instance, the training query could specify both the text to be replaced and the alternative text. Training videos could include videos displaying the specified text. Training output could include an output video that does not display the specified text and instead displays the alternative text with visual features similar to the specified text in the training video. Alternative text could include one or more different words or characters. For example, alternative text could be text in a different language than the specified text.

[0156] The training system can generate specific training examples by obtaining a seed video that may or may not display text. The training system can generate training videos by creating a copy of the seed video and adding text with certain visual features. The training system can generate output videos by creating a copy of the seed video and adding alternative text with the same or similar visual features as the text in the training video. The training system can generate training queries to include, for example, "Replace the text 'EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY' with the text 'ÜBERALLMERKWÜRDIGE DINGE AU HIMMEL'". The training system can generate additional training examples. For example, the training system can use the training video of a specific training example as the output video of an additional training example, and use the output video of a specific training example as the training video of an additional training example. The training system can generate training queries with additional training examples to include "Replace the text "ÜBERALL MERKWÜRDIGE DINGE AU HIMMEL" with the text "EVERYWHERE YOU LOOK WEIRD THINGS IN THE SKY".

[0157] In some examples, the alternative text may be longer or shorter than the text in the training video. To comply with visual constraints, the corresponding output video may display the text at a different size than the training video. The output video may display the text with a different amount of pixel detail than the training video, or in a different position or layout than described above.

[0158] The training system can fine-tune the generative neural network on the training dataset (step 530). For example, the training system can use a pre-trained generative neural network to process each training video to determine, for example, using backpropagation, updates to the parameters of the generative neural network. For example, the training system can fine-tune the generative neural network to minimize the loss based on the difference between the video generated by the generative neural network for the training video and the output video (also known as the expected video) of the corresponding training output. The loss can be measured, for example, per pixel. For example, the loss can be mean squared error (MSE) loss, mean absolute error (MAE) loss, or structural similarity index measure (SSIM).

[0159] In some implementations, the loss can also be measured between bounding boxes represented by the text of the state indicator. That is, the loss can be a combination of the difference between the generated video and the expected video, and the difference between the bounding boxes of the state indicator generated by the generative neural network (also known as the generated bounding boxes) and the bounding boxes of the state indicator output during training (also known as the expected bounding boxes). For example, the loss can be based on the difference in position and / or size between the generated and expected bounding boxes. For instance, if the generated and expected bounding boxes are not located in the same region of the video frame, the loss value can be higher than if they are in the same region. As another example, the loss can also be based on the overlap between the generated and expected bounding boxes. For example, the loss based on the difference between the generated and expected bounding boxes could be an MSE loss or a MAE loss.

[0160] In some implementations, the loss can also be measured between masks represented by the text of the state indicator. That is, the loss can be a combination of the difference between the generated video and the expected video, and the difference between the fine-grained masks of the pixels of the state indicator generated by the generative neural network and the fine-grained masks of the pixels of the state indicator from the training output. For example, the loss based on the difference between the pixel-wise masks could be a pixel-wise cross-entropy loss or a boundary loss.

[0161] In some implementations, the loss can also be based on the differences between the text of the state indicators. That is, the loss can be a combination of the differences between the generated video and the expected video, and the differences between the text of the state indicators. For example, for some training examples, the expected state indicator might indicate that no text is displayed in the training video, or that the specified text is not displayed in the training video. If the generated state indicator indicates that text was found, the combined loss value can be higher due to the differences between the text of the state indicators compared to a loss based solely on the differences between the generated video and the expected video.

[0162] As another example, the expected state indicator could indicate that the specified text is first displayed at 9:00. If the generated state indicator indicates that the specified text is first displayed at 11:00, the combined loss value can be higher due to the differences between the text in the state indicator, compared to the loss based solely on the differences between the generated video and the expected video.

[0163] In some implementations, the training system can use course learning to train the generative neural network on the training dataset. For example, the training system can sequentially train the generative neural network on training examples from different text editing tasks. Alternatively, the training system can sequentially fine-tune the generative neural network on each of the multiple text editing tasks in the training dataset.

[0164] Each text editing task in a text editing task sequence can have an associated complexity level. Text editing tasks can be ordered by increasing the complexity level within the sequence. For example, a text editing task sequence can begin with tasks of lower complexity. Lower complexity tasks may include one or more simpler tasks that require fewer steps or subtasks to complete compared to higher complexity tasks. The complexity of subsequent tasks in the sequence can gradually increase. For example, higher complexity tasks may require subtasks that require more steps or subtasks to complete compared to lower complexity tasks. In some examples, subtasks may include one or more preceding tasks in the sequence.

[0165] The training system can train the generative neural network on training examples of the same text editing task for multiple consecutive training stages before training it on training examples of the next text editing task in the text editing task sequence. For example, the training system can train the generative neural network on training examples of a specific text editing task until it has trained the network on all training examples of that specific text editing task in the training dataset. As another example, the training system can train the generative neural network on training examples of a specific text editing task for a predetermined number of training stages. As another example, the training system can train the generative neural network on training examples of a specific text editing task until a predetermined performance (e.g., training or validation accuracy) of the generative neural network is achieved. As another example, the training system can train the generative neural network on training examples of a specific text editing task until the marginal improvement in the performance (e.g., training or validation accuracy) of the generative neural network between corresponding training stages decreases below a predetermined threshold.

[0166] As an example, a text editing task sequence could include i) modifying the visual features of all text in an input video, ii) modifying the visual features of a specified text in the input video, iii) erasing text in the input video, iv) changing the size of text in the input video, and v) replacing the specified text with alternative text. In this example, modifying the visual features of all text in the input video is the least complex task. Replacing the specified text with alternative text is the most complex task. For example, replacing the specified text with alternative text might require subtasks such as erasing text in the input video.

[0167] In some implementations, the training system can further refine the generative neural network to minimize the loss function based on the aggregated reward value (step 540). The training system can derive the aggregated reward value from one or more reward values. Each reward value can be generated by a corresponding reward model. The reward model can have been trained to generate reward values ​​for corresponding comparison cues.

[0168] For example, to determine the aggregate reward value, the training system can receive a final dataset. The final dataset may include examples, each consisting of an original video displaying text and an input query specifying the text editing task to be performed. For each example, the training system can feed the original video to the generative neural network to generate a modified video. The training system can then feed the generated video to a reward model to generate one or more reward values. The training system can derive the aggregate reward value from the one or more reward values. Fine-tuning the generative neural network based on the aggregate reward value is described below. Figure 7 Further detailed description.

[0169] In some implementations, the training system can further refine the generative neural network at each of several fine-tuning iterations to minimize the loss function based on the aggregated reward value. For example, for each generated video in each iteration, the training system can use a reward model to generate one or more reward values ​​for the generated video. The training system can then derive the aggregated reward value for the generated video from the one or more reward values.

[0170] In some implementations, the training system can derive a reward model from a finely tuned generative neural network. For example, the training system can obtain a copy of the data representing the generative neural network. The training system can modify the copy of the data representing the generative neural network to obtain the reward model. For example, the training system can replace one or more of the last layers of the generative neural network with one or more layers that produce reward values.

[0171] The training system can train each reward model to generate reward values ​​for corresponding comparison cues. The training system can train each reward model on reward training examples, each including reward training input and reward training output. Reward training input may include the input video of the evaluation example, the generated output video for the input query of the evaluation example, and the comparison cues. Reward training output may include rating data for the generated output video evaluated by the comparison cues. The training system can train the reward models to learn rater preferences for the comparison cues regarding different aspects of the evaluation video.

[0172] The training system can obtain reward training examples for the reward model from the evaluation dataset and rating data. The evaluation dataset can include evaluation examples, each of which includes an input video displaying text, an input query specifying the text editing task to be performed, and a comparison video. In some examples, the comparison video is professionally generated. In other examples, the comparison video is the output video generated by different versions of the generative neural network. For example, different versions of the generative neural network could be earlier checkpoints of the generative neural network.

[0173] The training system can feed the input video for each evaluation example as input to the generative neural network. The training system can then obtain each generated output video from the generative neural network. The training system can generate rating data for each generated output video relative to a comparison video, tailored to specific aspects. For example, the training system can derive rating data from preference data received from one or more human raters.

[0174] The rating data of the generated output videos can represent the preference between the generated output videos and comparison videos for a specific aspect. The training system can receive preference data from the raters in response to presenting the generated output videos and comparison videos to the raters.

[0175] After the reward model has been trained, the training system can use it to generate a reward value for each generated video. The training system can combine these reward values ​​to obtain an aggregate reward value for the generated videos, as shown in the reference below. Figure 7 Further detailed description.

[0176] Example of a user interface for collecting human feedback: 600 Figure 6 The training system can obtain preference data through presentation 600. Presentation 600 can be displayed, for example, on the user device of the rater. The user device can be any type of computer or computing device with a display and configured to receive input from the rater. For example, the user device can be a computer, laptop, tablet, or mobile phone with a display and capable of receiving user input through an interface such as a keyboard, mouse, touchpad, or touchscreen.

[0177] Presentation 600 may include an input video 610 for evaluating the sample, a generated output video 620, and a comparison video 630. Presentation 600 may also include a description of the editing tasks between the input video 610 and the generated output video 620 and comparison video 630. Figure 6 In the example, the editing task is "Replaced "There was a noisy raccoon in our hotel!" with "¡ Habia un mapache ruidoso en nuestro hotel!".

[0178] Presentation 600 may also include prompts 640 for raters to respond to. Each prompt may include a query or question about evaluating the generated output video 620 and comparison video 630 based on specific aspects. For example, prompt 640a includes “Which video matches the look of the original text better?” and prompt 640b includes “Which video matches the layout of the original text better?”

[0179] exist Figure 6 In the example, the rater can indicate, based on user input to presentation 600, that the generated output video 620 better matches the appearance of the original text (the text in input video 610), and that video 630 better matches the layout of the original text. In some examples, the rater may indicate that the videos are indistinguishable or that they are “uncertain.”

[0180] The training system can receive preference data indicating the raters' preferences for each of the 640 prompts. In some examples, the training system can combine the preference data for the raters with preference data for other raters.

[0181] The training system can assign values ​​to the rating data of the generated output video 620 based on preference data. For example, a higher value for the generated output video 620 can indicate a higher preference for a specific aspect of the generated output video 620. A lower value for the generated output video 620 can indicate a lower preference for a specific aspect of the generated output video 620.

[0182] In some examples where comparison video 630 is an output video generated by a previous version of the generative neural network, the training system can assign a value to the generated output video 620 based on preference data indicating a preference for comparison videos generated by previous versions of the generative neural network. For example, the training system can provide a series of presentations, such as presentation 600, to the rater. Each presentation can include a comparison between an output video generated by a specific version of the generative neural network and output videos 620 generated by different versions of the generative neural network. The training system can receive preference data for each prompt for each comparison.

[0183] The training system can determine the order in which each version of the output video is preferred relative to other versions for each specific aspect. For example, the training system can determine that, for a specific aspect, the output video of version three is more preferred than the output video of version seven, and the output video of version seven is more preferred than the output video of version two. In some examples, the training system can determine this order by comparing the output video generated by each version with the output video of every other version. In some examples, the training system can assume transitivity to determine this order in O(log n) comparisons.

[0184] Based on this order, the training system can assign values ​​to the rating data of the output videos generated by the current version. For example, the output videos can be sorted in descending order of preference. Each output video in the sequence can correspond to a value in a predetermined interval between 1 and -1. For example, the first output video can correspond to the value 1, the second output video can correspond to the value 0.75, and the third output video can correspond to the value 0.5.

[0185] If the current output video is the one with the lowest preference, or the last output video, the training system can assign a value of -1 to the rating data. If the current output video is the one with the highest preference, or the first output video, the training system can assign a value of 1.0 to the rating data. If the current output video is an intermediate video, the training system can assign the corresponding value.

[0186] In some examples, preference data can indicate that two or more output videos are indistinguishable for a particular aspect. For instance, raters may have indicated they are "unsure" which video is better based on a prompt. As another example, there may be disagreements in preference among multiple raters regarding the output videos. Two or more output videos may correspond to the same value.

[0187] Figure 7 This is a diagram illustrating an example process 700 for fine-tuning a generative neural network based on an aggregated reward value 710. For example, the generative neural network could be... Figure 2 Generative neural network 210. Process 700 can be used as... Figure 5 Part of step 540 is made by Figure 2 The training system 240 executes. The training system can execute process 700 at each of the multiple fine-tuning iterations.

[0188] The training system can use the original videos from the samples in the final dataset. i Provided to the generative neural network 210. Figure 7 In the example, the current version of the generative neural network 210 is version m.

[0189] The training system can use the input of this example to query "replacement hints" i "Provided to the generative neural network 210. The generative neural network 210 can generate the output video G." i,m .

[0190] The training system can output video G i,m and the original video U i One or more reward models 242a-k (such as reward model 242a) are provided to generate reward values ​​750a-k. Each reward model in reward model 242 can be configured to generate a reward value 750a for a specific comparison cue k, 740 given a video. i,m,k A reward value of 750 can represent the video G generated based on the corresponding comparison prompt of 740. i,m and the original video U i The assessment.

[0191] For example, reward model 242a can receive videos, which include the original videos U from the final dataset. i, And for the original video U i The generated video G i,m This video can be U i + G i,m Cascade.

[0192] Comparison prompt 740a can correspond to reward model 242a. Comparison prompt 740a can include queries or questions about the same aspects of how reward model 242a is trained to predict rater preferences. Comparison prompt 740a can also indicate in U i and G i,m The text displayed in each video will be evaluated focusing on a specific aspect. For example, the original video U i It can display "text1". Input queries can include "Replace "text1" in the video with "text2". The generated video G i,m "text2" can be displayed. Therefore, comparison prompt 740a can include "evaluate "text1" and "text2" for visual similarity".

[0193] Figure 7 Other comparison tips for the 242b-k reward model may include information about the generated video G. i,m The text displayed in one or more video frames compared to the original video U i Questions or queries about aspects of the text displayed in one or more video frames (such as appearance, feel, layout, style, font, weight, or color).

[0194] The training system can receive reward value 750 from reward model 242. The training system can generate an aggregate reward value 710 from the reward value 750. In the example in Figure 8, the aggregate reward value 710 is the average of the reward values ​​750. In some implementations, each reward value in the reward value 750 can be multiplied by a corresponding weight. For example, the training system can give more weight to layout than font by assigning higher corresponding weights to the reward values ​​for layout.

[0195] The training system can then fine-tune the generative neural network 110 to minimize the loss function 720. The loss function 720 is based on the aggregated reward value 710. Figure 7 In the example, the loss function 720 is the sum of the aggregated reward value 710 and a constant 1.0 divided by a constant 2.0. Therefore, the training system can fine-tune the generative neural network 210 based on the reward value predicted by a reward model configured to predict human preferences.

[0196] This specification uses the term "configured" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform a specific operation or action, it means that the system has software, firmware, hardware, or a combination thereof installed thereon that causes the system to perform those operations or actions in operation. For one or more computer programs configured to perform a specific operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause that device to perform that operation or action.

[0197] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their equivalents), or in a combination of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of the data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals—e.g., machine-generated electrical, optical, or electromagnetic signals—to generate artificially generated propagation signals to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0198] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The device may also be or further include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.

[0199] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages); and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not need to, correspond to a file in a file system. A program may be stored as a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites and interconnected via a data communication network.

[0200] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same one or more computers.

[0201] The processes and logic flows described in this specification can be executed by one or more programmable computers, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., an FPGA or ASIC) or by a combination of a dedicated logic circuit system and one or more programmable computers.

[0202] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or incorporated into a special-purpose logic circuit system. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to or from them, or both. However, a computer does not necessarily have to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0203] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0204] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's device in response to a request received from a web browser. Similarly, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user in response.

[0205] Data processing devices used to implement machine learning models may also include, for example, dedicated hardware accelerator units for handling the general and computationally intensive portions of machine learning training or production (i.e., inference, workloads). Machine learning models can be implemented and deployed using machine learning frameworks such as TensorFlow or Jax.

[0206] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface, web browser, or app that a user can interact with through an implementation of the subject matter described in this specification), or any combination of one or more such back-end components, middleware components, or front-end components. Components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0207] A computing system may include clients and servers. Clients and servers are generally remote to each other and typically interact via a communication network. The client-server relationship is established by computer programs executed on respective computers that establish a client-server relationship between them. In some embodiments, the server transmits data (e.g., HTML pages) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from that user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.

[0208] While this specification contains numerous details of specific implementations, these details should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from that combination may be removed from the claimed combination, and the claimed combination may involve sub-combinations or variations thereof.

[0209] Similarly, although operations are depicted in the accompanying drawings and described in a specific order in the claims, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or requiring all shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0210] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require a specific order or sequence shown to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.

Claims

1. A computer-implemented method, comprising: Receive input video; Receive an input query that specifies a text editing task to be performed on the input video, the text editing task requiring modification of text displayed in one or more video frames of the input video; as well as A generative neural network is used to process the input video and the input query to generate an output that includes a modified video having modifications applied to the text displayed in the one or more video frames.

2. The method of claim 1, wherein receiving the input video includes receiving the input video from a user.

3. The method of any of the preceding claims, wherein receiving an input query includes receiving the input query from a user.

4. The method of any of the preceding claims, wherein the output further includes a status indicator for the modification.

5. The method of claim 4, wherein the status indicator indicates whether the modification was successfully applied, or the position of the text displayed in one or more video frames of the input video.

6. The method of any of the preceding claims, further comprising providing the modified video for display to a user.

7. The method of any preceding claim, wherein the text displayed in one or more video frames of the input video includes any one or more of the following: overlaid text, in-scene text, or animated text.

8. The method of any of the preceding claims, wherein the text editing task includes any one or more of the following: modifying one or more visual features of the text, removing the text, or replacing the text with alternative text that shares visual features with the text.

9. The method of any of the preceding claims, wherein the input query specifies i) the text to be replaced and ii) the alternative text.

10. The method of any of the preceding claims, wherein the input query specifies a different color for the text.

11. The method of any of the preceding claims, wherein the input query specifies i) text to be displayed in a different color and ii) the different color.

12. The method of any of the preceding claims, wherein the input query specifies the text to be removed.

13. The method of any of the preceding claims, wherein the input query specifies i) text to be displayed at a different size and ii) a definition indicating a change in the size of the specified text.

14. The method of any of the preceding claims, wherein the generative neural network has been pre-trained to generate videos conditioned on at least video and text.

15. The method of any of the preceding claims, wherein the generative neural network has been fine-tuned on a training dataset used for multiple text editing tasks, each text editing task requiring modification of pixels depicting text in an input video.

16. The method of claim 15, wherein the generative neural network has been sequentially fine-tuned on each of the plurality of text editing tasks in the training dataset.

17. The method of claim 15, wherein the training dataset includes a plurality of training examples, each training example including a training input and a corresponding training output, the training input including a training video and a training query specifying a text editing task among the plurality of text editing tasks, and the corresponding training output including at least an output video.

18. The method of any one of claims 14 to 17, wherein the generative neural network has been further fine-tuned to minimize a loss function based on aggregate reward values.

19. The method of claim 18, wherein the aggregated reward value is derived from one or more reward values, wherein each reward value is generated by a corresponding reward model.

20. The method of claim 19, wherein each corresponding reward model is configured to generate a reward value representing the evaluation of a given video based on the corresponding comparison cue and the original video.

21. The method of claim 20, wherein the corresponding comparison prompt includes questions regarding the visual similarity, layout, style, font, weight, or color of text displayed in one or more video frames of the given video compared to text displayed in one or more video frames of the original video.

22. The method of claim 21, wherein each corresponding reward model is trained to generate the reward value in response to the corresponding comparison cue.

23. A system comprising one or more computers and one or more storage devices, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations, the operations including: Receive input video; Receive an input query that specifies a text editing task to be performed on the input video, the text editing task requiring modification of text displayed in one or more video frames of the input video; as well as A generative neural network is used to process the input video and the input query to generate an output that includes a modified video having modifications applied to the text displayed in the one or more video frames.

24. One or more non-transitory computer-readable storage media, said one or more non-transitory computer-readable storage media storing instructions that, when executed by said one or more computers, cause said one or more computers to perform operations, said operations including: Receive input video; Receive an input query that specifies a text editing task to be performed on the input video, the text editing task requiring modification of text displayed in one or more video frames of the input video; as well as A generative neural network is used to process the input video and the input query to generate an output that includes a modified video having modifications applied to the text displayed in the one or more video frames.