Modifying objects in films

JP2024520059A5Active Publication Date: 2025-09-04FLAWLESS HLDG LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023573083
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-23
Filing Date
2022-05-26
Publication Date
2025-09-04
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

The production of foreign language versions of movies often results in a loss of nuance and quality due to the use of subtitles or audio dubbing, as the original film's visual and auditory elements cannot be accurately translated.

Method used

A computer-implemented method using machine learning models, particularly deep neural networks, to isolate and reconstruct object instances within video frames, allowing for the deep editing of photorealistic depictions, such as facial performances, to seamlessly integrate synthetic elements into the original footage.

Benefits of technology

Enables high-quality visual dubbing and performance transfer by maintaining the original film's visual and auditory fidelity, providing a seamless transition between original and translated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method for processing video data including a first sequence of image frames including a first instance of an object, the method comprising: isolating the first instance of the object in the first sequence of image frames, determining a first parameter value of a synthesis model of the object using the isolated first instance of the object, modifying the first parameter value of the synthesis model of the object, rendering the object using the trained machine learning model and the modified first parameter value of the synthesis model of the object, and replacing at least a portion of the first instance of the object in the first sequence of image frames with a corresponding at least a portion of the modified first instance of the object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to modifying objects or parts of objects in films, and has particular, but not exclusive, relevance to the visual dubbing of foreign language feature films. [Background technology]

[0002] The production of a live-action feature film (filmmaking) is a time-consuming and expensive process that typically requires the involvement of many skilled professionals performing many interdependent tasks under strict time and resource constraints. A typical filmmaking process involves a production phase spanning multiple shoots, where raw video footage (with accompanying audio) is shot for multiple takes of each scene in the film. In the post-production phase, the raw footage is copied and compressed before selected portions are assembled by editors and directors for offline editing. Visual effects (VFX) are then applied as required while the sections of raw video footage corresponding to the offline edit are taken and audio is mixed, edited, and re-recorded as necessary. Before the master copy of the film is delivered, the resulting footage and audio undergo a finishing phase where additional processes such as color grading may be performed.

[0003] The high cost and interdependence of the tasks involved in the film production process, as well as the typical time constraints and variability of factors such as weather and actor availability, mean that it is rarely feasible to reshoot scenes in a film. Thus, a film must be constructed from footage created in an initial production stage to which VFX is applied as necessary. The production stage typically generates hundreds of hours of high-resolution raw video footage, only a small portion of which is ultimately used in the film. The raw footage may not capture the desired combination of actor performance(s) and conditions such as weather, background, lighting, etc., which can only be corrected to a limited extent during the VFX and finishing stages. Summary of the Invention [Problem to be solved by the invention]

[0004] When the film production process is complete, a master copy of the film is distributed for screening in cinemas, streaming services, television, etc. For some films, foreign language versions are produced in parallel with the original film, and the foreign language versions are distributed at the same time as the original film. Foreign language versions of films typically use text subtitling or audio dubbing to recreate the dialogue in the desired language. In either case, it is generally accepted that the foreign language version loses much of the nuance and quality of the original film. [Means for solving the problem]

[0005] According to a first aspect, a computer-implemented method for processing video data including a sequence of a plurality of image frames is provided. The method includes identifying each instance of an object in at least some of the sequence of image frames. For at least some of the identified instances of the object, the method includes isolating the instance of the object in image frames including the instance of the object, and using the isolated instances of the object to determine associated parameter values ​​of a synthesis model of the object. The method includes training a machine learning model to reconstruct the isolated instances of the object based at least in part on the associated parameter values ​​of the synthesis model of the object. The method includes modifying a first parameter value for a first instance of the object appearing in a first sequence of image frames and an associated first parameter value for the synthesis model of the object, rendering a modified first instance of the object using the trained machine learning model and the modified first parameter value for the synthesis model of the object, and replacing at least a portion of the first instance of the object in the first sequence of image frames with a corresponding at least a portion of the modified first instance of the object.

[0006] By training a machine learning model to reconstruct isolated object instances from within the video data, the methodology enables "deep editing" of photorealistic depictions of the video data beyond the capabilities of traditional VFX. The multiple sequences of image frames may correspond to footage of different takes of a scene within a feature film, for example, and can provide a wealth of learning data for the machine learning model under relatively consistent lighting / environmental conditions. The first sequence of image frames may or may not be one of the multiple sequences of image frames. The methodology is suitable for integration into a film production pipeline, where training of the machine learning model occurs in parallel with an offline editing process, potentially using the same video data as the offline editing process.

[0007] The object may be a particular human face, in which case the method may be used in applications such as visual dubbing of foreign language versions of films or performance transfer, where an actor's performance in a particular take of a particular scene is transferred to another take of the same scene, to another scene or to another film. At least a portion of the object may be part of a human face, including the mouth but excluding the eyes. The inventors have found that by replacing only this region of the face, a realistic visual dubbing or performance transfer can be achieved with minimal impact on the actor's performance.

[0008] Modifying the first parameter values ​​may include determining target parameter values ​​for a synthetic model of the object and progressively interpolating between the first parameter values ​​and the target parameter values ​​over a subsequence of the first sequence of image frames. The interpolation may include linear and / or non-linear interpolation. In this manner, the original first instance may be progressively transitioned to the modified first instance in a smooth and seamless manner. Furthermore, the deviation of the modified first instance from the original first instance may be increased and decreased to allow continuous deep editing of the object instance. For example, if the purpose of modifying the first instance is to match an audio track, the deviation may be greatest when the mismatch between the original first instance and the audio track is most noticeable. This may minimize the perceived effects in the original footage while achieving the desired result.

[0009] In the example where the first parameter value is progressively interpolated as described above, the computer-implemented method may further comprise detecting an event in the sequence of image frames and / or an audio track associated with the first sequence of image frames, determining one or more image frames of the first sequence of image frames in which the detected event occurs, and determining a sub-sequence of the first sequence of image frames in response to the determined one or more image frames in which the detected event occurs. For example, the sub-sequence of the first sequence of image frames may be determined such that the sub-sequence ends before the event occurs. Thus, the first instance of the object may undergo the greatest modification upon the occurrence of the event. In the context of visual dubbing, the event may be, for example, an event in which a plosive or bilabial nasal sound is uttered in either the first language or the second language, since this is when the visual inconsistency between the first language and the second language becomes most noticeable.

[0010] The machine learning model may include a deep neural network configured to process one or more input images to generate an output image. For the at least some of the identified instances of an object, isolating the instance of the object may include generating a registered portion of each of the image frames containing the instance of the object, and training the machine learning model may include rendering, for each of the image frames containing the instance of the object, a synthetic image of a portion of the instance of the object using the synthesis model and associated parameter values ​​of the synthesis model, superimposing the synthetic image of the portion of the instance of the object on the registered portion of each of the image frames containing the instance of the object to generate a respective synthetic image for each of the image frames containing the instance of the object, and adversarially training a deep neural network to process the generated synthetic images to reconstruct at least one frame of the isolated instance of the object. By providing the synthetic images as inputs to the deep neural network, the network can learn how to take into account lighting, color, and other characteristics derivable from an outer region of at least a portion of the object instance being modified, while also learning to perform realistic inpainting to seamlessly integrate the modified portion of the object instance back into the original image frame. In another example, in addition to or as an alternative to a synthetic image, a synthetic image of the entire instance of the object may be provided as input to the deep neural network.

[0011] The deep neural network may be configured to process the attention mask together with each of the one or more input images to generate an output image. Training the machine learning model for the at least some of the identified instances of the object may include generating, for each of the image frames including the instances of the object, a respective attention mask that highlights one or more features of the instances of the object, and training the deep neural network to process each attention mask together with the generated synthetic image to reconstruct at least one frame of an isolated instance of the object. By providing the attention mask as an independent input to the deep neural network, the network may learn to focus attention on specific regions of the synthetic image as guided by the attention mask. The attention mask may have one or more layers that highlight different features of the object. Each attention mask may include, for example, a segmentation mask that separates the instances of the object from background regions, and / or a mask that indicates other features, such as facial features if the object is a face. The attention mask may be generated from a synthetic model of the object together with the synthetic image. Adversarial training of deep neural networks may use adversarial losses restricted to object regions defined by attention masks, focusing the efforts of the deep neural network to faithfully reconstruct the object regions.

[0012] The adversarial training of the deep neural network may use an adversarial loss and one or more other loss functions, such as a perceptual or photometric loss function that indicates the photometric difference between at least one frame of an isolated instance of an object and at least one reconstructed frame of the isolated instance of the object. The (one or more) other loss functions may be restricted to object regions defined by respective attention masks. By using the photometric and / or perceptual loss in combination with the adversarial loss, the network is taught to generate a photorealistic reconstruction of the original instance of the object. The photometric loss may be an L2 loss modified to reduce the contribution of small photometric differences, which the inventors have found to reduce artifacts in the renderings generated by the deep neural network.

[0013] The deep neural network may be configured to process the projected ST map together with each of the one or more input images to generate an output image. For the at least some of the identified instances of the object, training the machine learning model may include generating a respective projected ST map for each of the image frames including the instances of the object, each projected ST map having pixel values ​​corresponding to texture coordinates of a synthetic model of the object, and training the deep neural network to process each projected ST map together with the generated synthetic image to reconstruct at least one frame of the isolated instance of the object. The projected ST map may provide input that the deep neural network can use to associate a surface area of ​​the object with a location in the synthetic image, enhancing the ability of the deep neural network to accurately reconstruct the instance of the object.

[0014] The deep neural network may be configured to process the projection noise map together with each of the one or more input images to generate an output image. For the at least some of the identified instances of the object, training the machine learning model may include generating a respective projection noise map for each of the image frames including the instances of the object, each projection noise map having pixel values ​​corresponding to values ​​of a noise texture applied to a synthetic model of the object, and training the deep neural network to process each projection noise map together with the generated synthetic images to reconstruct at least one frame of the isolated instance of the object. The projection noise map provides an additional input upon which the deep learning model can learn to build a spatially dependent texture in its rendered output.

[0015] The computer-implemented method may include, for said at least some of the identified instances of the object, color normalizing the isolated instances of the object, and training the machine learning model may use the color-normalized isolated instances of the object. Color normalizing the isolated instances simulates similar lighting conditions throughout the training data, simplifying the task of the machine learning model.

[0016] In an example, identifying each instance of an object may include discarding image frames in which the instance of the object is rotated by an angle outside a predetermined range with respect to an axis coplanar with the image frame. In some cases, it may be difficult to train a machine learning model to reconstruct instances of an object in all possible orientations. To address this issue, a method may treat views of an object from different perspectives as entirely different objects and train separate models for them accordingly.

[0017] The associated parameter values ​​of the synthesis model for the at least some of the identified instances of the object may comprise a base parameter value encoding a base geometry of the object and a deformation parameter value encoding a deformation of the base geometry of the object for each of the image frames containing an instance of the object. And the first parameter values ​​of the synthesis model may comprise a first deformation parameter value encoding each deformation of the base geometry of the object for each of the image frames of the first sequence of image frames. Modifying the first parameter values ​​may comprise modifying the first deformation parameter value. In some use cases, the desired modification of the object is a deformation of a non-rigid object, in which case only the deformation parameter value may need to be modified.

[0018] Modifying the first deformation parameter values ​​may include obtaining a second sequence of image frames including an instance of a second object (the second object may be the same object as the first object or may be a different object from the first object), isolating the instance of the second object in the second sequence of image frames to generate isolated second instance data, determining second parameter values ​​of the synthetic model using the isolated second instance data, the second parameter values ​​including second deformation parameter values ​​encoding deformations of a base geometry of the second object for each of the image frames of the second sequence of image frames, and updating the first deformation parameter values ​​using the second deformation parameter values. In this way, the second sequence of image frames is used as driving data for modifying the first deformation parameter values. In the case of visual dubbing, the second object typically corresponds to the face of the actor performing the dubbing. In the case of performance transfer, the second object typically corresponds to the face of the original actor.

[0019] The first sequence of image frames may be at a higher resolution than the multiple sequences of image frames. In this case, rendering the modified first instance of the object may include rendering an intermediate first instance at a resolution consistent with the multiple image frames and applying a super-resolution neural network to the intermediate first instance to render the modified first instance. This allows the machine learning model to be trained using low-resolution image data, significantly reducing the computational demands of training, while at the same time generating high-resolution renderings suitable for incorporation into high-resolution video data.

[0020] The first instance of the object may be in a different form than any of the at least some of the identified instances of the object, in which case the method may further comprise obtaining a first sequence of image frames, isolating the first instance of the object within the first sequence of image frames, and determining first parameter values ​​for a synthetic model of the object using the isolated first instance of the object. Alternatively, the first instance of the object may be one of the at least some of the identified instances of the object.

[0021] According to a second aspect, there is provided a computer-implemented method for processing source video data comprising a sequence of a plurality of image frames, the method comprising: detecting respective instances of an object within at least some of the sequence of image frames, for a first instance of the object detected in the first sequence of image frames, determining a frame-by-frame position and size of the first instance of the object in the first sequence of image frames, using a neural renderer to obtain replacement video data comprising a modified instance of the object, and using the determined frame-by-frame position and size to replace at least a portion of the first instance of the object in the first sequence of image frames with at least a portion of the modified instance of the object.

[0022] Obtaining the replacement video data may include, for each of one or more instances of an object detected in the respective sequence of image frames of the source video data, processing at least a portion of each of the respective image frames of the sequence of image frames to generate a three-dimensional synthetic model of each instance of the object, generating a respective sequence of synthetic images from the three-dimensional synthetic model of each instance of the object, and training a neural renderer to reconstruct each instance of the object using the respective sequence of synthetic images. For the three-dimensional synthetic model of a first instance of the object, the method may include modifying the three-dimensional synthetic model, generating a first sequence of synthetic images from the modified three-dimensional synthetic model, and generating the replacement video data using the trained neural renderer and the generated first sequence of synthetic images.

[0023] The method may comprise determining, for each of one or more instances of the object, a frame-wise location of a box containing each instance of the object within each sequence of image frames, and determining at least a portion of each of the image frames of each sequence of image frames as a portion contained within the box. The method may further comprise determining, for all image frames of each sequence of image frames, a size of the box such that each instance is contained within the box.

[0024] Generating the three-dimensional synthetic model for each of the one or more instances of the object may include tracking landmarks of each instance of the object and generating the three-dimensional synthetic model responsive to positions of the tracked landmarks.

[0025] The method may include, for each of one or more instances of the object, determining a pose of each instance of the object for each image frame of the respective sequence of image frames using the generated three-dimensional synthetic model, and normalizing each instance of the object to be a substantially constant size between each image frame of the sequence of image frames using the determined pose for each image frame of the respective sequence of image frames.

[0026] The object may be a first object, and the method may include detecting each instance of the object within at least some of the sequence of image frames, performing object recognition to determine an identifier of the detected instance of the object, identifying each of the multiple instances of the first object as having a common identifier, and training a neural renderer to reconstruct each of the multiple instances of the first object identified as having a common identifier.

[0027] The object may be a human face, and the at least a portion of the first instance of the object may be a portion of a face including a mouth and excluding eyes. In an example where the object is a face, modifying the three-dimensional synthetic model may include obtaining driving data including an audio and / or video recording including speech, processing the driving data to determine modification parameter values ​​for the three-dimensional synthetic model corresponding to the speech, and using the modification parameter values ​​to modify the three-dimensional synthetic model. For example, the first instance of the object may be an instance of a face speaking in a first language, and the audio and / or video recording may be speech in a second language different from the first language.

[0028] Modifying the three-dimensional synthetic model may include gradually transitioning between unmodified parameter values ​​for the three-dimensional synthetic model and modified parameter values ​​for the three-dimensional model depending on when speech is occurring in the driving data. In this way, abrupt changes in expression can be avoided. For example, it may be preferable to use modified parameter values ​​only when speech is detected in the driving data to avoid irrelevant facial expressions in the driving data appearing in the processed video data. For example, it may be preferable to gradually transition between unmodified parameter values ​​at times preceding or following speech in the driving data, or to gradually transition between parameter values ​​corresponding to neutral facial expressions at times preceding or following speech in the driving data.

[0029] Modifying the three-dimensional synthetic model may include determining when the face is speaking in the first sequence of image frames, and reducing the amplitude of mouth movements of the three-dimensional synthetic model when it is determined that the first instance of the object is speaking in the first sequence of image frames, for example, suppressing the mouth movements of the first language actor when the first language actor is speaking but the second language actor is not.

[0030] Modifying the three-dimensional synthetic model may include modifying the mouth shapes of the three-dimensional synthetic model to match plosives or bilabial nasals detected in the driving data. Incorrect mouth shapes during plosives or bilabial nasals are particularly easy for a listener to detect, and therefore precise control of the synthetic model at these moments may be particularly appropriate.

[0031] Any of the above methods may include obtaining mask data indicative of a frame-by-frame shape of at least a portion of the modified instance of the object, and replacing at least a first instance of the given object with at least a portion of the replaced instance of the given object may use the mask data. For example, the mask data may be first mask data, and the method may include obtaining second mask data indicative of a frame-by-frame shape of at least a portion of the first instance of the object. Replacing at least a portion of the first instance of the object with at least a portion of the modified instance of the object may include determining, based on a comparison between the first mask data and the second mask data, that a boundary of at least a portion of the first instance of the object exceeds a boundary of at least a portion of the modified instance of the object, and performing clean background generation in a region of the sequence of image frames between the boundary of at least a portion of the first instance of the object and at least a portion of the modified instance of the object. The clean background generation may enable plausible insertion of the modified instance of the object in cases where unintended artifacts of the first instance remain after replacement with the modified instance.

[0032] The method may include adjusting a color palette of a first instance of an object to be consistent throughout a sequence of image frames, which may simplify the task of a machine learning model or neural renderer by eliminating the need to model variations in the color palette due to changes in lighting, etc.

[0033] Replacing at least a portion of the first instance of the object with at least a portion of the modified instance of the object may include softening edges of at least a portion of the modified instance of the object to achieve a gradual blend between the replaced portion and the underlying image frame.

[0034] For any of the above methods, replacing at least a portion of the first instance of the object comprises determining optical flow data estimating a warping relating the first instance of the object and a modified first instance of the object for a subset of the first sequence of image frames falling within the time window, progressively applying the estimated warping to the first instance of the object across the subset of the first sequence of image frames to determine a progressively warped first instance of the object, progressively applying an inverse of the estimated warping to the modified first instance of the object across the subset of the first sequence of image frames to determine a progressively warped modified first instance of the object, and progressively dissolving the progressively warped first instance of the object into the progressively warped modified first instance of the object across the subset of the first sequence of image frames. Progressively warping and dissolving the images allows the modified first instance to be seamlessly incorporated into the first sequence of image frames in situations where a gradual change is visible.

[0035] The gradual dissolving may be performed at a predetermined dissolve rate, and the gradual application of the estimated warping and the inverse of the estimated warping may be performed at a predetermined warping rate. The ratio of the dissolve rate to the warping rate may increase to a maximum value within a subsequence of the sequence of image frames and then decrease. In this manner, the gradual dissolving may be concentrated, for example, within a central set of image frames of the subset. The inventors have found that by concentrating the dissolve in this manner, a more seamless transition between the first instance of the object and the modified first instance of the object can be achieved and image sharpness can be maintained during warping.

[0036] For any of the methods described above, replacing at least a portion of the first instance of the object includes determining optical flow data indicative of an estimated warping associating the first instance of the object and a modified first instance of the object; applying the estimated warping to the first instance of the object to determine a warped first instance of the object; blurring the warped first instance of the object; adjusting a color of the modified first instance of the object based on a pixel-by-pixel ratio of the blurred warped first instance of the object and the blurred modified first instance of the object to generate a color-graded modified first instance of the object; and replacing at least a portion of the first instance of the object with a corresponding at least a portion of the color-graded modified first instance of the object. A pixel-by-pixel ratio of the blurred warped first instance and the blurred modified first instance represents a color grading map for matching a color of the modified first instance to the original first instance of the object, allowing short-scale local variations in lighting and color to be reproduced in the modified first instance of the object. Blurring the warped instance of the object and blurring the modified instance of the object may be performed using a blur filter having a characteristic length scale of 3 to 20 pixels.

[0037] For any of the methods described above, replacing at least a portion of the first instance of the object includes determining optical flow data indicative of an estimated warping associating the first instance of the object and a modified first instance of the object; applying the estimated warping to the first instance of the object to determine a warped first instance of the object; blurring the warped first instance of the object; blurring the modified first instance of the object; adjusting a color of the modified instance of the object based on a pixel-by-pixel ratio of the blurred warped instance of the object and the blurred modified instance of the object to generate a color-graded modified instance of the object; and replacing at least a portion of the first instance of the object with a corresponding at least a portion of the color-graded modified first instance of the object. A pixel-by-pixel ratio of the blurred warped first instance and the blurred modified first instance represents a color grading map for matching a color of the modified first instance to the original first instance of the object, allowing short-scale local variations in lighting and color to be reproduced in the modified first instance of the object. Blurring the warped instance of the object and blurring the modified instance of the object may be performed using a blur filter having a characteristic length scale of 3 to 20 pixels.

[0038] According to a third aspect, there is provided a computer-implemented method for processing video data comprising a plurality of sequences of image frames. The method comprises identifying respective instances of an object within at least some of the sequences of image frames. For at least some of the identified instances of the object, the method comprises isolating the instances of the object within image frames containing the instances of the object, and determining associated parameter values ​​of a synthetic model of the object using the isolated instances of the object. The method comprises training a machine learning model to reconstruct the isolated instances of the object based at least in part on the associated parameter values ​​of the synthetic model of the object.

[0039] According to a fourth aspect, there is provided a computer-implemented method for processing video data including a first sequence of image frames including a first instance of an object, the method comprising: isolating the first instance of the object in the first sequence of image frames, determining first parameter values ​​for a synthetic model of the object using the isolated first instance of the object, modifying the first parameter values, rendering a modified first instance of the object using the trained machine learning model and the modified first parameter values, and replacing at least a portion of the first instance of the object in the first sequence of image frames with a corresponding at least a portion of the modified first instance of the object.

[0040] According to a fifth aspect, there is provided a computer-implemented method for processing video data comprising a sequence of image frames, the method comprising: isolating an instance of an object within the sequence of image frames, generating a modified instance of the object using a machine learning model, and modifying the video data to progressively transition between at least a portion of the isolated instance of the object and a corresponding at least a portion of the modified instance of the object across a subsequence of the sequence of image frames.

[0041] The sub-sequence of the sequence of image frames may be a first sub-sequence of the sequence of image frames, and modifying the video data may be to gradually transition from at least a part of the isolated instance of the object to a corresponding at least a part of the modified instance of the object. The method may further comprise modifying the video data to gradually transition from at least a part of the modified instance of the object back to the corresponding at least a part of the isolated instance of the object over a second sub-sequence of the sequence of image frames. In this manner, the method may provide a smooth or gradual transition from the isolated instance of the object to the modified instance of the object and back again, for example in response to a particular event in the video data and / or associated audio data.

[0042] According to a sixth aspect, there is provided a non-transitory storage medium for storing video data, the video data including a first sequence of image frames including a photographic representation of an object, a second sequence of image frames in which at least a portion of the photographic representation of the object is replaced with a corresponding at least a portion of a composite representation of the object, and a third sequence of image frames between the first sequence of image frames and the second sequence of image frames, where in the third sequence of image frames at least a portion of the photographic representation of the object is modified to progressively transition between at least a portion of the photographic representation of the object at an end of the first sequence of image frames and a corresponding at least a portion of the composite representation of the object at an beginning of the second sequence of image frames.

[0043] Modifying at least a portion of the photographic representation of the object may comprise simultaneously warping and dissolving at least a portion of the photographic representation of the object into at least a portion of the synthetic representation of the object. The warping may be performed in stages at a predetermined warping rate, and the dissolving may be performed in stages at a predetermined dissolve rate, and a ratio of the dissolve rate to the warping rate may increase to a maximum value in the third sequence of image frames and then decrease. This allows the dissolve to be concentrated within a central set of image frames of the subsequence, and allows a seamless transition between the photographic representation of the object and the synthetic representation of the object to be achieved while maintaining image sharpness during the warping.

[0044] The composite representation of the object may be a first composite representation of the object, and modifying at least a portion of the photographic representation of the object may comprise incrementally interpolating between a second composite representation of the object and the first composite representation of the object, the second composite representation of the object corresponding geometrically to the photographic representation of the object. Thus, the photographic representation of the object may be replaced with the geometrically corresponding composite representation before modifying or transforming the composite representation. The composite representation may be transformed or modified in a way that is not feasible with the photographic representation. The effect of modifying the photographic representation may be achieved by spatially or geometrically aligning the photographic representation and the composite representation.

[0045] According to a seventh aspect, there is provided a data processing system comprising means for performing any of the methods described above. The data processing system may comprise one or more processors and a memory storing machine readable instructions which, when executed by the one or more processors, cause the one or more processors to perform any of the methods described above.

[0046] According to an eighth aspect, there is provided a computer program product (e.g. a computer program stored on a non-transitory storage medium) comprising instructions which, when the program is executed by a computer, cause the computer to perform any of the methods described above.

[0047] According to a ninth aspect, there is provided an audiovisual product manufactured using any of the methods described above.

[0048] Further features and advantages of the present invention will become apparent from the following description of preferred embodiments of the invention, given by way of example only with reference to the accompanying drawings, in which: [Brief description of the drawings]

[0049] [Figure 1] FIG. 1 illustrates a schematic diagram of a collocated data processing system according to an embodiment.

[0050] [Diagram 2] FIG. 2 illustrates a schematic of a method for training a machine learning model according to an embodiment.

[0051] [Diagram 3] FIG. 3 shows an example of an object instance being isolated from a sequence of image frames.

[0052] [Figure 4] FIG. 4 illustrates a schematic of a method for training a deep neural network model according to an embodiment.

[0053] [Figure 5A] FIG. 5A shows an example of an input to a deep neural network. [Figure 5B] FIG. 5B shows an example of an input to a deep neural network. [Figure 5C] FIG. 5C shows an example of an input to a deep neural network.

[0054] [Figure 6] FIG. 6 illustrates generally a method for modifying an instance of an object in a sequence of image frames.

[0055] [Figure 7] FIG. 7 illustrates a schematic example of modifying an object instance based on video driving data.

[0056] [Figure 8] FIG. 8 illustrates generally how the deep neural network of FIG. 4 can be used to modify an instance of an object in a sequence of image frames.

[0057] [Figure 9] FIG. 9 illustrates an example of modifying an instance of an object by varying degrees according to an embodiment.

[0058] [Figure 10] FIG. 10 illustrates generally a method for transitioning from an instance of an object in a sequence of image frames to a modified instance of the object according to an embodiment.

[0059] [Figure 11] FIG. 11 shows an example of processing video data according to the method of FIG.

[0060] [Figure 12] FIG. 12 illustrates generally a method for performing automatic color grading when transitioning from an instance of an object to a modified instance of the object in a sequence of image frames according to an embodiment.

[0061] [Figure 13] FIG. 13 illustrates a schematic of a movie production pipeline for a foreign language version of a movie including visual dubbing according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0062] Details of the system and method according to the embodiment will become apparent from the following description with reference to the drawings. For purposes of explanation, numerous specific details of particular embodiments are described herein. References herein to "one embodiment" or similar language mean that a feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment, but not necessarily in other embodiments. Furthermore, it should be noted that particular embodiments are generally described with the omission and / or necessarily simplification of certain features to facilitate description and understanding of the concepts underlying the embodiment.

[0063] FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate to modifying objects in movies. In this disclosure, a movie may refer to any form of digital video data or audiovisual product. In particular, the embodiments described herein address the challenges associated with modifying objects in feature films in a seamless manner, both in terms of the quality of the output and in terms of integrating the associated processes into the filmmaking workflow. The techniques disclosed herein provide methods related to tasks such as visual dubbing of foreign language films, performance transitions between movie scenes, and modifying background objects in movies.

[0064] FIG. 1 illustrates a schematic diagram of a data processing system 100 according to an embodiment. The data processing system 100 includes a network interface 102 for communicating with remote devices over a network 104. The data processing system 100 may be a single device, such as a server computer, or may include multiple devices, such as multiple server computers connected over a network. The data processing system 100 includes a memory 106, which in this disclosure refers to both non-volatile storage and volatile and non-volatile working memory. The memory 106 is communicatively coupled to a processing circuit 108, which may include any number of processing units, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU) or neural network accelerator (NNA), one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), etc.

[0065] The memory 106 is arranged to store various types of data for implementing the methods described below. In particular, the memory 106 may store video data 110 including a sequence of image frames, which may correspond to raw and / or processed video footage captured by one or more cameras. The video data 110 may comprise, for example, picture rushes captured during the production of a movie, and / or may comprise compressed or otherwise processed footage. The video data 110 may also comprise video footage that has been modified as a result of applying the methods described herein.

[0066] The memory 106 may further store isolated instance data 112 indicative of isolated instances of one or more objects appearing in the video data 110. In this disclosure, an instance of an object broadly refers to an uninterrupted appearance of an object in a sequence of image frames. For example, in a given scene of a movie, an object may appear in a first sequence of image frames, then be occluded or move out of the camera's field of view in a second sequence of image frames, and then appear again in a third sequence of image frames, in which case two instances of the object are recorded. The isolated instance data 112 may include a sequence of image frames extracted from the video data 110, and / or may include metadata such as a timestamp indicating in which portion of the video data 110 a given instance appears along with the position, scale, and / or orientation of the object in each of the video frames in which the instance appears. The isolated instance data 112 may further include a registered portion of each of the image frames in which the instance appears, such as a bounding box, which may be resized, rotated, and / or stabilized as described in more detail below.

[0067] The memory 106 may further store synthetic model data 114 that encodes a synthetic model of one or more objects appearing in the video data 110. The synthetic model of an object may approximate the color, texture, and other visual characteristics of the object as well as the geometric characteristics of the object. The synthetic model may be a three-dimensional model that allows for rendering a two-dimensional synthetic image corresponding to a view of the synthetic model from a given camera position and orientation. The synthetic model may have adjustable parameters to control aspects of the model. For example, a synthetic model may correspond to a particular class or type of object and may have adjustable parameters with different values ​​that correspond to different objects within the class and / or different instances of a given object within the class. For example, a synthetic model for a class of "human faces" may be able to represent a range of human faces and orientations, expressions, etc. by specifying values ​​for the adjustable parameters of the synthetic model. Alternatively, the synthetic model may correspond to a particular object. For example, the synthetic model may be a deformable model of a non-rigid object such that different deformations correspond to different values ​​of the adjustable parameters of the synthetic model.

[0068] The memory 106 may further store machine learning model data 116 corresponding to a machine learning model. A machine learning model is a class of algorithms that generate output data based at least in part on parameter values ​​that are learned from data instead of being manually programmed by a human. Of particular relevance to the present disclosure are deep learning models in which machine learning is used to learn parameter values ​​of one or more deep neural networks, as described in more detail below. The data processing system 100 may use the machine learning model to render, among other tasks, an instance of a photorealistic depiction of an object for incorporation into a video, based at least in part on parameter values ​​of a synthetic model of the object. The machine learning data 116 may include parameter values ​​learned in response to the video data 110 and other data, as described in more detail below.

[0069] The memory 106 may further store program code 118 including routines for performing the computer-implemented methods described herein. The routines may allow fully automated execution of the methods described herein and / or may allow user input to control various aspects of the process. The program code 118 may, for example, define software tools to allow a user to perform deep editing of objects in the video data.

[0070] FIG. 2 illustrates a method of processing video data 202 to train a machine learning model according to an embodiment. The method may be performed by any suitable data processing system, such as the data processing system 100 of FIG. 1. When the machine learning model is trained, it may be used to generate instances of photorealistic depictions of objects for incorporation into a video. The video data 202 includes a plurality of sequences of image frames, of which a sequence of image frames A and a sequence of image frames B are illustrated. Each sequence of image frames may correspond to footage of each take of a scene or part of a scene in a feature film. The video data 210 may include all takes (i.e., all picture rushes) of all scenes in the film, or a subset thereof. Before performing the method of FIG. 2, the footage may be optionally downsized or compressed and / or converted to a common format. For example, the footage may be converted to a 2K format (i.e., a format with a horizontal dimension of approximately 2000 pixels). In this way, the format of the video data 210 (including, for example, resolution, pixel depth, color format) may be made consistent for processing as described below. Furthermore, by downsizing the footage, the computational cost of training machine learning models can be significantly reduced. During the traditional film production process, this process is commonly performed to generate a smaller volume of data to work with in the offline editing process. The video data 202 may be the same data used for offline editing.

[0071] The method of FIG. 2 proceeds to perform object detection and separation 204 for each of the sequence of image frames. In this regard, the object detection may generate metadata that identifies image frames that contain instances of a predefined class of object and indicates the location of each such object in each of the image frames. The metadata may include, for example, the location and dimensions of the bounding box of the object in each of the image frames that contain the object. Depending on the object detection algorithm, the bounding box may have predefined dimensions (e.g., one or more fixed-size, fixed-aspect-ratio squares or rectangles) or may have variable dimensions. The bounding box of a given instance may have fixed dimensions, determined, for example, such that a given instance is completely contained within the bounding box of all image frames in which the instance appears. It is common for the apparent size of an object to vary from instance to instance and / or image frame to image frame, and therefore the object detection algorithm is preferably capable of detecting objects at multiple scales. The object detection algorithm may be a machine learning algorithm, such as a deep learning algorithm. Examples of suitable object detection algorithms are Region-based Convolutional Neural Networks (R-CNN), Fast R-CNN, Faster R-CNN, Region-based Fully Convolutional Networks (R-FCN), Single Shot Detector (SSD) and Only Look Once (YOLO - multiple versions up to v5 are available at the time of writing). Various techniques including data selection, data augmentation and bootstrapping may be used during the training of these algorithms to achieve a desired level of performance for a particular task.

[0072] Multiple instances of a given object may be detected and isolated by object detection and isolation 204. In this example, instance A is detected in sequence A of image frames, and instance B and instance C are detected in sequence B of image frames (indicating that the object disappears from view and reappears in sequence B of image frames).

[0073] In addition to detecting instances of objects of a given class, object detection and separation 204 may include recognizing distinct objects of the same class. In an example where the objects are human faces, each time an instance of a face is detected, the method may perform face recognition to determine whether the face is a new face or a previously detected face. In this manner, an instance of a first object can be distinguished from an instance of a second object, and so on. Thus, metadata stored with a detected instance of an object may include an identifier of the object.

[0074] In addition to detecting instances of objects, object detection and isolation 204 may include determining the locations of a sparse set of 2D landmarks for isolated instances of objects. 2D landmarks are two-dimensional feature points that roughly represent an object. These landmarks may be used to aid in fitting a synthetic model, as described below. If the object is a human face, the landmarks may include, for example, points around the eyes and mouth and following the nose ridge. 2D landmarks may be detected frame-by-frame using sparse keypoint detection methods. Additionally, optical flow may be used across a sequence of image frames to determine a time-consistent trajectory of detected landmarks to improve the accuracy of the estimation of the landmarks' locations.

[0075] Object detection and isolation 204 may further comprise stabilizing and / or registering the isolated instance of the object. Stabilization and / or registration may be performed, for example, to ensure that for each frame of a given isolated instance, the object appears at a constant rotation angle relative to an axis perpendicular to the plane of the image frame. A normalization step may be applied so that the object appears at a substantially constant size between frames. Thus, object detection and isolation 204 may comprise determining a stabilization point in each of said image frames containing an instance of the object, the stabilization point may be determined, for example, as a function of the position of one or more two-dimensional landmarks. The method may then comprise stabilizing the instance of the object around the determined stabilization point such that the stabilization point remains at a fixed position and the object does not rotate significantly around this point. This stabilization may be performed using any suitable image registration technique and may take advantage of two-dimensional landmarks if they have been determined. In some cases, registration may be performed without the need to define a stabilization point. The inventors have found that it is beneficial to stabilize the object instance to reduce the difficulty of downstream tasks including synthetic model fitting and / or machine learning. It has been found to be particularly useful to determine a stabilization point that is within or near the part of the object instance that is to be replaced. In the case of visual dubbing or human facial performance transitions, the stabilization point may be the center of the mouth.

[0076] Each isolated instance may be stored as a video clip along with metadata including, for example, data indicating which image frame contains the instance, as well as the location, size and orientation of the instance within each image frame that contains it. The location, size and orientation may be stored, for example, as the coordinates of the upper left and lower right corners of a bounding box within an image frame. Other metadata may include information identifying the object, the resolution and frame rate of the image frames. The isolated instances may optionally be stored along with associated guide audio.

[0077] The metadata includes information necessary for the portion of the sequence of image frames to be reconstructed from the separated instances. Figure 3 shows an example of a sequence of image frames 302 corresponding to footage of a particular take of a scene in a movie for which a foreign language version is to be generated using the methods described herein. In this example, a first actor instance 304, a second actor instance 306, and a third actor instance 308 are detected and recognized, and the different actors are treated as different "objects" within the meaning of this disclosure. In this example, the first actor instance 304 and the second actor instance 306 appear in a generally frontal view, while the third actor instance 308 appears in profile. In some examples, a profile view of a particular actor (or, more generally, a view whose Euler angles with respect to an axis coplanar with the image frames are outside a predetermined range) may be treated as a different object than a frontal view of the same actor. The instances 304, 306, along with their respective metadata 314, 316, are separated to generate separated instances 310, 312. In this example, it is determined that the third actor does not speak in the scene, and therefore the third actor instance 308 is not isolated. The metadata 314, 316 enable generation of a sequence of image frames 308 from the isolated instances 310, 312, which includes a generated sequence of overlaid frames 318 that include reconstructions 320, 322 of the instances 304, 306 in their original positions in the sequence of image frames 302.

[0078] The method of FIG. 2 continues with synthetic model fitting 206, which uses the isolated instance of the object to determine parameter values ​​of a synthetic model of the object. The synthetic model may be a synthetic dense three-dimensional model, such as a three-dimensional morphable model (3DMM) of the object, and may be comprised of a mesh model formed of polygons, such as triangles and / or quadrangles, each having sides and vertices. The synthetic model may be parameterized by a set of fixed parameters and a set of variable parameters. The fixed parameters encode properties of the object that are not expected to change (or can be reasonably modeled as not changing) between image frames, whereas the variable parameters encode properties that may change between image frames. The fixed parameters may comprise base parameter values ​​for encoding a base geometry of the object (e.g., a neurally represented facial geometry) that is treated as a starting point from which deformations are applied. The base geometry may comprise, for example, the positions of a set of vertices of a mesh model. The variable parameters may comprise deformation parameters for encoding changes to the base geometry of the object. These deformation parameters may, for example, control the deformation of each vertex of the mesh. Alternatively, the deformation parameters may control the weighting of a linear combination of a predetermined set of blendshapes, each of which corresponds to a particular global deformation of the base geometry. Alternatively, the deformation parameters may control the weighting of a linear combination of a predetermined set of delta blendshapes, each of which corresponds to a deformation over a particular subset of vertices. By specifying particular weightings, a blendshape or linear combination of delta blendshapes can represent a wide range of deformations to the base geometry.

[0079] The fixed parameters of the synthetic model may include, in addition to the basic parameters, parameters encoding a reflectance model of the object's surface (and / or other surface properties of the object) along with intrinsic camera parameter values ​​for projecting the synthetic model onto the image plane (although in some cases the intrinsic camera parameter values ​​may be known and do not need to be determined). The reflectance model may treat the object's surface as a perfectly diffuse surface that scatters incident illumination evenly in all directions. Such a model is sometimes called a Lambertian reflectance model. This model has been shown to achieve a reasonable trade-off between complexity and realistic results.

[0080] The variable parameters may further include parameters encoding the position and / or orientation of the object relative to the camera as viewed within the isolated instance of the object together with an illumination model that characterizes the irradiance of the object at a given point. The illumination model may model the illumination at a given point on the surface of the object using a predetermined number of spherical harmonic basis functions (e.g., the first three bands L0, L1, L2 of the spherical harmonic basis functions). The combination of the reflectance model and the illumination model may model the irradiance at a given point on the surface of the object according to a set of parameter values ​​determined during fitting of the model.

[0081] As described above, parameter values ​​of a synthesis model of an object are determined for each instance of the object, with at least some of the parameter values ​​being determined on a frame-by-frame basis. In the example of FIG. 2, a respective set of parameter values ​​is determined for each of the object instances A, B, and C detected at 204. The parameter values ​​of the synthesis model may be determined independently for each instance of the object. Alternatively, some of the fixed parameter values ​​(such as the fixed parameter values ​​encoding the base geometry and reflectance model) may be fitted across multiple instances of the object, which may improve accuracy, particularly for instances of an object that include a relatively small number of image frames or where the object is not clearly visible.

[0082] The synthetic model of the object, together with the parameter values ​​determined for a particular isolated instance of the object, may be used to generate a synthetic image corresponding to a projection of the object onto the image plane. These synthetic images may be compared to corresponding frames of the isolated instances to determine parameter values ​​that minimize a metric difference or loss function that characterizes the deviation between the synthetic image and the corresponding frames of the isolated instances. In this manner, parameter values ​​that fit the synthetic model to the isolated instances of the object may be determined. Additional techniques may be used to improve the accuracy of the model fitting, including, for example, a loss term that compares the positions of 2D landmarks detected in the isolated instances of the object with corresponding feature vertices of the synthetic model or a loss term that compares the contours of the isolated instances of the object with corresponding contours of the synthetic model.

[0083] The method of Figure 2 continues with machine learning 208, where the isolated instances of the object and associated parameter values ​​of the synthesis model are used to train a machine learning model to reconstruct isolated instances of the object. By performing this training on multiple instances of the object (e.g., instances from multiple takes of a movie scene), the machine learning model may learn to generate photorealistic instances of the object based on a set of parameter values ​​of the synthesis model. The machine learning 208 process generates trained parameter values ​​210 for the machine learning model.

[0084] The machine learning model may include one or more neural networks. For example, the machine learning model may include a conditional generative adversarial network (GAN) having a generative network configured to generate images depending on parameter values ​​of a synthetic model and a discriminative network configured to predict whether a given image is real or generated by the generative network. The generative network and the discriminative network may be trained in parallel with each other using an adversarial loss function that rewards the discriminative network for making accurate predictions and rewards the generative network for causing the discriminative network to make incorrect predictions. This type of training may be referred to as adversarial training. The adversarial loss function may be supplemented with one or more other loss functions, such as a photometric loss function that penalizes differences between pixel values ​​of isolated instances of an object and pixel values ​​of images output by the generative network and / or a perceptual loss function that compares images output by the generative network with isolated instances in the feature space of an image encoder (such as a VGG net trained on ImageNet). By combining the adversarial loss function with the photometric loss function and / or the perceptual loss function, the generative network may learn to generate sequences of images that are photometrically similar to isolated instances of an object and stylistically indistinguishable from isolated instances of the object. In this way, a generative network may learn to generate photorealistic reconstructions of isolated instances of objects.

[0085] In one example, the machine learning model may have a generative network that receives as input a set of parameter values ​​derived from a sequence of one or more frames of isolated instances of an object and generates an output image. During training, the output image may be compared to a given frame of the sequence (e.g., an intermediate frame or a final frame), where the generative network may learn to reconstruct that frame. By using parameter values ​​from multiple frames, the generative network may take into account information before and / or after the frame being reconstructed, which may enable the generative network to take into account dynamic properties of the object.

[0086] As an alternative to direct processing of the parameter values ​​of the synthesis model, the machine learning model is arranged to receive input derived from the synthesis model itself. For example, the machine learning model may be arranged to process input data based at least in part on a synthetic image rendered from the synthesis model. FIG. 4 illustrates an example of a method using a sequence of synthetic images 402 rendered from a synthetic model of an object and corresponding to isolated instances 404 of the object to generate input data for a neural network. Each of the synthetic images 402 may include the entire object or a portion of the object, e.g., a portion of the object to be replaced or modified. The synthetic images may include an alpha channel that encodes an alpha matte or binary mask that designates background regions surrounding the object or portion of the object as transparent. The isolated instances 404 are optionally subjected to color normalization 406, which may be performed on a frame-by-frame basis or for all frames of the isolated instances 404. Color normalization simulates similar coarse lighting conditions across all isolated instances of the object, which may aid the learning process described below by reducing the spatial extent of the images that the machine learning model needs to learn to generate.

[0087] For each frame containing an isolated instance 404 of an object, a portion of the corresponding composite image 402 is overlaid on the (possibly color normalized) frame of the isolated instance 404 resulting in a composite image 408. As described above, each of the frames of the isolated instance 404 may be a registered portion of the image frames containing the instance of the object. The portion of the composite image 402 to be overlaid may be defined using a segmentation mask, which may be generated using a composite model of the object. To generate the mask, an ST map is obtained with linearly increasing U and V values ​​encoded in the red and green channels, respectively. UV mapping may then be used to map the ST map onto the composite model. Regions suitable for the mask may be defined in the ST map manually or automatically, for example by referencing certain feature vertices of the composite model (as described above). Projections of the mapped regions may then be rendered for each of the composite images 402, and the rendered projections may be used to define the geometry of the mask for the overlay process. With this approach, a mask is obtained that conforms to the geometry of the composite model, and this approach only needs to be defined once for a particular object or a particular instance of an object. The mask used for registration may be a regular binary segmentation mask or a soft mask, the latter of which provides a gradual blend between the separated instances 404 and the overlaid parts of the composite image 402 .

[0088] 5A shows an example of the synthetic image described above in which a portion 502 of a synthetic image of a face rendered from a synthetic model of a face is superimposed on a frame 504 containing an isolated instance of a face. In this example, portion 502 includes the mouth but excludes the eyes, and is defined using a binary mask generated using the ST map as described above.

[0089] Returning to FIG. 4 , a synthetic image 410 is provided as an input to a generative network 412. The generative network 412 is configured to process the synthetic image 410 to generate candidate reconstructions 414 of an instance of an object. By processing the synthetic image 410 instead of the full synthetic image 402 generated by the synthetic model, more information is available to the generative network 412 regarding the lighting and color characteristics of the reconstructed instance, particularly in the area surrounding the portion to be replaced. This has been found to enhance the ability of the generative network 412 to reconstruct instances of the object as it learns how to perform inpainting such that the reconstructed portion of the object blends seamlessly into the area surrounding the object. It should be noted that while in this example the synthetic image is provided as an input to the generative network 412, in other examples a full rendering of the synthetic model may alternatively or additionally be provided as input.

[0090] In a single forward pass, the generative network 412 may be configured to process a spatiotemporal volume containing a predetermined number of composite images 410 (e.g., 1, 2, 5, 10, or any other suitable number of composite images 410) to generate one or more frames of reconstruction candidates 414 corresponding to a predetermined one or more of the composite images 410. A space-time volume in this context refers to a set of images that appear consecutively within a time window. The generative network 412 may output a reconstruction candidate for a single frame, for example, corresponding to the last composite image 410 of the space-time volume. By processing multiple composite images 410 simultaneously, the generative network 412 may learn to use information about how objects move over time to achieve a more realistic output. By performing this process in a sliding window manner in time, the generative network 412 may generate a reconstruction candidate of an object for each frame that contains an isolated instance of the object. A space-time volume may not be defined for the first few frames or the last few frames. Alternatively, the space-time volume may be expanded by replicating the first and / or last frame X times, where X is the size of the time window that can effectively affect the Dirichlet boundary conditions. In this way, the space-time volume remains defined, but is biased by the first and last few image frames. Other boundary conditions may instead be used to expand the space-time volume.

[0091] The generative network may have an encoder-decoder architecture with an encoder portion configured to map the spatio-temporal volume to latent variables in a low-dimensional latent space and a decoder portion configured to map the latent variables to one or more frames containing candidate reconstructions of the object. The encoder portion may be composed of several downsampling components that may reduce the resolution of the input. The given downsampling components may include convolution filters and non-linear activation functions (such as rectified linear unit ReLU, activation functions). The decoder portion may be composed of several upsampling components that may increase the resolution of the input. The given upsampling components may include convolution filters and non-linear activation functions, optionally with other layers or filters. At least some components of the encoder portion and / or decoder portion may utilize batch normalization and / or dropout during training. In a particular example, the generative network 412 has eight downsampling components to reduce the resolution from 256×256 to 32×32 and eight upsampling components to restore the resolution to 256×256. Each of the downsampling components uses a 4x4 convolutional layer with stride 2, followed by batch normalization, dropout and a leaky ReLU activation function. Each of the upsampling components utilizes a cascade refinement strategy, using a 4x4 deconvolutional filter with stride 2, followed by batch normalization, dropout and a ReLU activation function, followed by two 3x3 convolutional filters with stride 1 each, followed by another ReLU activation function. The output of the final upsampling components is passed through a TanH activation function to generate a single frame of candidate reconstructed instances of the object.Batch normalization may be omitted from the first downsampling component and the last upsampling component, and as an improvement, the architecture may employ skip connections from the input layer to one or more decoder components to allow the network to transfer detailed structure. It is understood that other architectures for the generative network 142 are possible and this architecture is provided by way of example only.

[0092] A generative network 412 is adversarially trained to reconstruct the separated instance 404 of the object. In this example, a discriminator network 416 is used that receives as input the same spatiotemporal volume of the synthetic image 410 used by the generative network 412 to generate the one or more frames of the reconstructed instance 414, along with one or more frames of the separated instance 402 (which may be considered the "ground truth" in this context) of the reconstructed instance 414 generated by the generative network 412. The discriminator network attempts to predict whether it has received the reconstructed instance 414 or the ground truth separated instance 412. An adversarial loss 418 is determined that rewards the discriminator network 416 for making accurate predictions and rewards the generative network 412 that caused the discriminator network 416 to make incorrect predictions. Backpropagation (indicated by dashed arrows in FIG. 4 ) is then used to determine the gradient of the adversarial loss 418 with respect to the parameters of the generative network 412 and the discriminative network 416, and the parameter values ​​of the generative network 412 and the discriminative network 416 are updated according to the determined gradient of the adversarial loss, for example, using stochastic gradient descent or a variant thereof. The adversarial loss 418 may be supplemented with one or more other losses (not shown), such as a photometric loss that penalizes the difference between the pixel values ​​of the separated instance 402 and the pixel values ​​of the reconstructed instance 414 output by the generative network 412, or a perceptual loss that penalizes the difference between the image features of the separated instance 402 and the image features of the reconstructed instance 414 output by the generative network 412. The photometric loss may be, for example, an L1 loss, an L2 loss, or any other suitable loss based on the ratio between the pixel values ​​of the separated instance 402 and the pixel values ​​of the reconstructed instance 414. In a particular example, the photometric loss may be a modified L2 loss that is modified to reduce the contribution of small photometric differences.In this way, contributions from training samples on which the generative network 412 performs well (i.e., easy samples) are reduced relative to contributions from training samples on which the generative network 412 struggles (i.e., difficult samples). For example, the photometric loss may be a modified L2 loss that multiplies the squared photometric difference by a sigmoid function or a soft-step function that reduces the contribution of photometric differences below a predetermined value. The inventors have found that using this type of loss function during training causes the generative network 412 to produce more accurate renderings with fewer artifacts than other loss functions.

[0093] By combining an adversarial loss function with a photometric loss function, the generative network 412 can learn to generate reconstructed instances of an object that are photometrically similar to and stylistically indistinguishable from ground truth instances of the object, meaning that the reconstructed instances preserve the distinctiveness of the isolated instances.

[0094] The generative network 412 may further be configured to process an attention mask 420 together with each of the synthetic images 410, which may be applied to the input of the discriminatory network 416 during masking operations 422, 424. This has the effect of restricting the loss function to the region defined by the attention mask 420. The photometric loss (if any) may be similarly restricted to the region defined by the attention mask 420. The attention mask 420 may be a regular binary mask or a soft mask and may define the boundaries of a region that includes the whole object or a part of the object. The attention mask 420 may be output from a synthetic model of the object, or may be generated from an isolated instance of the object, for example, using semantic segmentation. By providing the attention mask 420 as an additional input to the generative network 412 and restricting the loss function to the region defined by the attention mask 420, the generative network 412 may learn to focus attention on the object rather than the background. This may be particularly important for dynamic backgrounds such as those expected in movies. The attention mask 420 may define a larger area than the portion of the object being modified and replaced, so that the generative network 412 focuses its attention on the area surrounding the portion being modified, thereby learning to integrate the replaced portion with the area surrounding the object. Alternatively or additionally, the attention mask 420 may be applied to the synthetic image before the synthetic image is input to the generative network 412, to provide the attention mask 420 as an input to the generative network 412. In either of these cases, the generative network 412 may generate "hallucinated" outputs for areas outside the attention mask 420, because there is no training signal associated with these areas of the output. FIG. 5B shows an example of an attention mask corresponding to the synthetic image of FIG. 5A. In this case, it is observed that the attention mask defines a larger area of ​​the face than the portion of the face 502 being replaced.

[0095] The generative network 412 may further be configured to process a projected ST map (not shown in FIG. 4) with each of the composite image frames 410. As described above, the projected ST map may be generated using a synthesis model from which the composite images 404 are generated. In particular, a general ST map may be obtained with linearly increasing U and V values ​​encoded in the red and green channels, respectively. The ST map may be applied to the synthesis model using UV mapping to render a projection of the ST map for each of the composite images 404. FIG. 5C shows an example of a projected ST map corresponding to the synthetic image of FIG. 5A and the attention mask of FIG. 5B. It is observed that the color of the projected ST map varies across the surface of the face such that red (R) increases (when viewed from the left side of the face to the right side) and green (G) increases (when viewed from the bottom part of the face to the top part of the face). Because the ST map conforms to the surface of the synthetic model, pixels corresponding to the same position of the face in two different synthetic images have a common pixel value (color). The projected ST map may enable the generative network 412 to relate surface regions of the synthetic model to locations in the synthetic image, helping the generative network 412 generate surface details at consistent locations of the object. In another example, a projected normal coordinate code (PNCC) image may be used instead of an ST map as a spatially dependent input to the generative network 412. However, a projected ST map may be preferred because it more directly maps surface regions of the object to locations in the synthetic image and uses only two channels compared to the three channels of a PNCC image. In another example, in addition to or as an alternative to the projected ST map or PNCC image, other types of projected maps may be input to the generative network 412, which may further improve the quality of the output of the generative network 412. For example, one or more projected maps may be generated from the synthetic model to highlight particular features or aspects of the object. For example, a projected topology map may be generated that shows the topology of the object surface. This may help the generative network 412 generate details that are consistent with the topology of the object surface.In an example where the object is a human face, the topology map may indicate the topology of facial features such as the nose and mouth.

[0096] The generative network 412 may further be configured to process a projected noise map (not shown) along with each of the composite image frames 410 (and optionally one or more other maps). Similar to the projected ST map, the projected noise map may be generated using a synthesis model from which the composite images 404 are generated. In particular, a noise map may be obtained whose pixel values ​​do not depend on identically distributed random variables (such as Gaussian variables) or on which the noise pixel values ​​do depend. In a particular example, the noise map may be a Perlin noise map. The noise map may be applied to the synthesis model using UV mapping, and a projection of the noise map may be rendered for each of the composite images 404. The noise map provides additional resources that the generative network 412 can use to generate rich textures that conform to the surface of the object. Perlin noise is particularly suitable for representing complex natural textures. For example, the noise map may be stored in the blue channel of the ST map (since the ST map uses only the red and green channels by default), in which case only one UV mapping run is required. Additional maps may further be provided as input to the generator (e.g., as additional channels of ST and / or noise maps) to improve the quality of the output rendered by the generative network 412. For example, one or more maps derived from a synthetic model of the object, such as a generic map emulating grain detail or a normal map and / or a displacement map, may be provided to the generative network 412.

[0097] A machine learning model trained using the above method may be used to generate modified instances of photorealistic representations of objects, as described below. FIG. 6 illustrates a method of using the machine learning model of FIG. 2 with trained parameter values ​​210 to modify instances of objects in a first sequence of image frames 602. The first sequence 602 may correspond to one of the sequences of image frames used to train the machine learning model, although in some examples, the first sequence 602 may not have undergone the downsizing or compression described in connection with the training process. This is because the objective of this stage is to generate the highest quality and most photorealistic rendered instances of the object, and computational costs are less of an issue than during training. In the context of a movie production pipeline, the first sequence 602 of image frames may be the highest resolution at which the movie needs to be delivered. The first sequence 602 may be determined manually (e.g., a user may determine that instances of objects in this sequence need to be replaced) or automatically upon detection of a talking character, for example, in the context of visual dubbing. The first sequence 602 may be one of the sequences of image frames used to train a machine learning model, but this is not required.

[0098] The method of FIG. 6 proceeds with object detection and isolation 604, in this example, obtaining an isolated first instance of the object and metadata for replacing the first instance of the object, as described above. Then, synthetic model fitting 608 is performed to generate first parameter values ​​610 of a synthetic model of the object. A more accurate synthetic model fitting 608 may be obtained by increasing the resolution of the first instance 606 compared to the resolution used for training. In other examples where the first sequence 602 is one of the sequences of image frames used to train a machine learning model, it may not be necessary to perform the object detection and isolation 604 and / or synthetic model fitting 608 steps again, as these steps may have already been performed on the first instance 606 during training.

[0099] The synthetic model's first parameter value 610 is modified 612 to obtain a modified first parameter value 614. The modification 612 of the first parameter value 610 modifies the appearance of the synthetic model, and ultimately allows for rendering of a modified instance of the object. The modification 612 of the first parameter value 610 may be performed manually, for example by receiving user input via a user interface from which the modified first parameter value can be derived, allowing for deep editing of the object instance beyond what can be achieved using typical VFX techniques. Alternatively, the modification 612 of the first parameter value 610 may be performed at least partially automatically, for example in response to driving data, such as video driving data and / or audio driving data.

[0100] FIG. 7 illustrates an example of modifying parameter values ​​of a particular human face synthesis model in response to video driving data 704. As a result, a face instance 702 may be modified. The instance 702 may, for example, correspond to an actor speaking lines from a movie in a first language, and the video driving data 704 may correspond to a dubbing actor speaking a translation of the same lines in a second language. In this example, the video driving data 704 and / or the instance 702 are clipped such that the video driving data 704 and the instance 702 span the same number of frames. In this example, primary synthesis model parameters 706 are derived for the instance 702 using the methods described above. The primary synthesis model parameters include fixed parameters that encode the intrinsic camera parameters, base geometry and reflection model, and variable parameters for each of the frames of the instance 702 that encode the respective poses and transformations relative to the base geometry. Secondary model parameters 708 are then derived for the video driving data 704 using the same methods described above. The secondary parameter values ​​708 include secondary deformation parameter values ​​710 for each of the frames of the video driving data 704, which indicate deformations of the base geometry determined for the dubbing actor (in the context of a human face, the deformations may represent facial expressions). A style transfer 712 is optionally performed, in which the secondary deformation parameter values ​​710 are adjusted for stylistic consistency with the deformation parameter values ​​derived for the primary object (in this case, the first language actor). The style transfer 712 may be performed manually, for example, by a VFX artist, or may be performed automatically or semi-automatically. The style transfer 712 may be performed using a style transfer neural network trained to modify the deformation parameter values ​​derived from a secondary source (e.g., a video source) for stylistic consistency with the deformation parameters derived from the primary source.Training may be performed using two style transfer neural networks, each configured to transform primary transformation parameter values ​​to secondary transformation parameter values ​​and back again. The generative network may be trained adversarially with circular consistency.

[0101] Style Transformation 712 allows for the transformation derived from a given secondary source to be "transformed" into a stylistically consistent transformation of the primary object. Style Transformation 712 may be unnecessary in some cases, for example, when the secondary source is stylistically similar to the primary source or when the primary and secondary sources depict the same object. The latter occurs, for example, when an actor's performance is transferred from one take of a scene to another take of a scene.

[0102] The primary parameter values ​​706 of the synthesis model, excluding the primary transformation parameter values, may be combined with the (possibly stylized) secondary transformation parameter values ​​710 to generate modified parameter values ​​714 of the synthesis model.

[0103] It should be noted that in the example of Figure 7, video driving data is used to determine the modified parameter values ​​of the synthetic model, whereas in other examples audio driving data or a combination of video and audio driving data may be used to determine the modified parameter values ​​of the synthetic model, in which case a separate audio driving neural network or a mixed mode neural network may be trained to determine the parameter values ​​of the synthetic model.

[0104] Returning to Figure 6, rendering 616 is performed where the machine learning model uses the trained parameter values ​​210 of the machine learning mode to render a modified first instance 618 of the object in response to modified first parameter values ​​614 for the synthetic model of the object. The rendering 616 may include processing the modified first parameter values ​​614 using, for example, a conditional GAN. Alternatively, the rendering may include generating a synthetic image from the synthetic model of the object and using the synthetic image to generate the modified first instance 618 of the object.

[0105] Figure 8 shows an example of a method for rendering modified instances of an object using the trained generative network 212 of Figure 2. The method is equivalent to that of Figure 2, but without the associated functionality of training the discriminative network 416 and generative network 212. The synthetic image 804 of Figure 8 is rendered from a synthetic model using modified parameter values. The synthetic image 810 is thus a hybrid image in which the portions of the synthetic image that are overlaid correspond to different parameter values ​​than the remainder of the isolated instance. The trained generative network 212 nevertheless transforms these hybrid images into modified instances of photorealistic depictions of the object.

[0106] It should be noted that the isolated instances 802, and therefore the synthetic image 810, may be of a higher resolution than the images used to train the generative network 212. In some examples, the generative network 212 may be a fully convolutional network (i.e., does not include fully connected layers). In this case, the generative network 212 may be able to process high-resolution input images to generate high-resolution output images despite being trained on low-resolution images. Alternatively, the isolated instances 802 (or synthetic image 810) may be downsized or compressed before being input to the generative network 212. In this case, a super-resolution neural network may be applied to the output of the generative network 212 to generate a photorealistic output at the appropriate resolution. The inventors have found that the latter approach produces highly plausible rendering outputs.

[0107] In some examples, such as the example of FIG. 7 above, parameter values ​​derived from an isolated instance of an object are replaced with parameter values ​​derived from a driving data source. Thus, an instance of an object may be modified for all image frames that contain an instance of the object. In other examples, it may be sufficient, and indeed may be preferred, to modify an instance of an object only for a subset of frames that contain an instance of the object, for example to smoothly transition between an unmodified instance of the object and a modified instance of the object. In the case of visual dubbing, modifying the shape of the first language actor's mouth for all image frames that contain the first language actor or all image frames that contain the first language actor in which the first language actor speaks may produce unrealistic results that adversely affect the viewing experience. The inventors have found that the viewing experience may be less adversely affected by, for example, modifying the shape of the first language actor's mouth only at specific times when either the first language actor or the second language actor has their mouth closed, which is easily detectable as being incompatible with most generated sounds. Furthermore, it may be preferred to smoothly transition between the performance of the first language actor in non-dialogue moments and the performance of the second language actor in dialogue moments or at least when the second language actor is speaking.

[0108] FIG. 9 illustrates an example of deriving two sets of parameter values ​​for a synthetic model of an object. In particular, a primary set of parameter values ​​is derived from a primary video source containing an instance of the object, and a secondary set of parameter values ​​is derived from a secondary driving data source (which may be, for example, a video source or an audio source). In this example, the primary set of parameter values ​​and the secondary set of parameter values ​​differ only in the transformation parameter values. The synthesis model may be controlled to change between the corresponding transformations by interpolating the transformation parameter values ​​of the synthesis model between the primary and secondary values. In this way, the extent to which the object is modified may be controlled, e.g., maximized only when a predefined event occurs, or relaxed to limit the extent to which the object is modified. For example, a synthetic model of a first language actor's face may be changed between the first language actor's performance and the second language actor's performance. The resulting blended performance may be better than maximally modifying the first language actor's performance throughout the entire period the actor is speaking. In the example of FIG. 9, the transformation parameter values ​​are interpolated between the first language actor's performance P and the second language actor's performance S. In particular, the second language actor's performance is phased in and out around event 902, when the second language actor's mouth is closed, and around event 904, when the first language actor's mouth is closed. The second language actor's mouth is closed for a short period of time when the second language actor makes a plosive sound (such as the letter "p") at event 902, and the first language actor's mouth is closed for a longer period of time when the first language actor makes a bilabial nasal sound (such as the letter "m"). Alternatively, event 902 may indicate that the first language actor is speaking, and event 904 may indicate that the second language actor is speaking.

[0109] Events 902 and 904 may be determined manually, for example, by an editor reviewing footage of the first language actor and the second language actor and marking the time when a certain event occurs, such as a mouth closed event. Alternatively, such events may be detected automatically from the audio or video data. For example, appropriate audio filters or machine learning models (e.g., recurrent neural networks or temporal convolutional neural networks) may be used to identify specific auditory events, such as plosives or bilabial nasals or the speech of a particular person in the audio data. Alternatively, appropriate machine learning models may be trained to visually identify such events. In the example of FIG. 9, events 904 and 902 may be automatically detected in the primary audio track 906 and the secondary audio track 908, respectively, and the interpolation of transformation parameter values ​​may be automated such that the second language actor's performance gradually or gradually phases in a predetermined time before the events 902, 904 and gradually phases out a predetermined time after the events 902, 904.

[0110] Other effects of using interpolation of deformation parameter values ​​include amplifying or suppressing the amplitude of expressions controllable by the deformation parameter values. For example, when a first language actor is speaking but a second language actor is not, instead of substituting deformation parameter values ​​derived from the second language actor, the deformation parameter values ​​may be adjusted to reduce the amplitude of the first language actor's mouth movements to correspond to a roughly neutral mouth shape. As another example, when a second language actor produces a plosive or bilabial consonant, the deformation parameter values ​​may be automatically or manually adjusted to ensure that the generated mouth shape is closed at such times.

[0111] After rendering the modified first instance 618 of the object, the method of FIG. 6 proceeds to object replacement 620, where at least a portion of the first instance 606 of the object is replaced with at least a corresponding portion of the modified first instance 618 of the object. This replacement may be achieved by compositing the modified first instance 618 into the first sequence of image frames 602, where the compositing process comprises superimposing the modified first instance 618 of the object (or a portion thereof) onto the first sequence of image frames 602 using metadata stored in association with the first instance 606 of the object. Any stabilization, registration or color normalization applied during isolation of the first instance 606 is inversely applied (i.e., reversed) to the modified first instance 618 prior to superimposition. To achieve a gradual blend between the replaced portion and the underlying image frame, a soft mask (alpha matte) may be applied to a portion of the modified first instance of the object to be superimposed (e.g., the lower region of the face including the mouth but excluding the eyes). The mask may be generated using a synthesis model of the object. In particular, appropriate regions may be defined in the ST map described above, applied to the synthesis model using UV mapping, and a projection may be rendered for each image frame of the modified instance of the object. The rendered projection may be used to define the geometry of the mask for the synthesis process. With this approach, a mask is obtained that conforms to the geometry of the object, and the approach only needs to be defined once for a particular object or a particular instance of an object. The regions of the ST map may be defined manually or automatically, for example by referencing certain feature vertices of the synthesis model (as used for the synthesis model fitting described above). The ST map may be identical to the map used to generate the synthesis image, as described with reference to FIG. 7.

[0112] To facilitate accurate replacement of the first instance 606 of the object with the modified first instance 618 of the object, a synthetic model of the object may be used to generate first mask data indicative of the frame-wise shape of the first instance 606 (or a portion thereof) and second mask data indicative of the frame-wise shape of the modified first instance 618 (or a portion thereof). The object replacement 620 may then include comparing the first mask data and the second mask data to determine whether the boundaries of the first instance 606 exceed the boundaries of the first instance 618 for any of the image frames in which the first instance 606 is replaced. This may occur, for example, when the first instance 606 represents a face with an open mouth whereas the modified first instance 618 represents a face with a closed mouth. In this case, a portion of the first instance 606 may still be visible after overlay of the replaced first instance 618. In such cases, a clean background generation may be performed, for example, using a visual effects tool such as Mocha Pro in Boris FX (RTM) or by application deep inpainting techniques to replace the trace of the first instance 606 with a suitable background.

[0113] In some examples, noise may be applied to the replaced portion of the object to match digital noise or grain appearing in the first sequence of image frames (which may not appear in the rendered portion of the object. For example, Perlin noise may be applied at a scale and intensity that matches the digital noise appearing in the image frames.

[0114] The compositing process produces a sequence of modified image frames that replace the instance of the object. In some cases, the sequence of modified image frames can simply replace the original image frames of the video data. This may be possible if an instance of the object is replaced or modified in every image frame in which it is visible. In other cases, transitioning directly from the original image frames to the modified image frames may produce undesirable effects and artifacts. In the example of visual dubbing, transitioning from footage of an actor speaking in a first language to a synthetic rendering of the actor speaking in a second language may, for example, cause the shape of the actor's mouth to instantly change from an open position to a closed position or vice versa. To mitigate these problems, the inventors have developed a technique that allows for a more seamless transition from the original instance of an object to the modified instance of the object and vice versa.

[0115] FIG. 10 illustrates an example of a method for processing video data including a sequence of original image frames 1002 and a sequence of modified image frames 1004. The modified image frame 1004 is identical to the original image frame 1002, except that in the modified image frame 1004, instances of objects appearing in the image frames have been modified and replaced using techniques described herein. To generate the modified image frame 1004, at least a portion of the modified instances of the objects are composited with the original image frame 1004, thereby replacing the original instances of the objects. In this example, the sequence of original image frames 1002 precedes another sequence of original image frames (not shown) that is replaced with a corresponding modified image frame. The other sequence of image frames may, for example, include footage of an actor speaking in a first language that is to be dubbed into a second language. The sequence of original image frames 1002 may include an image frame in which the actor first begins speaking in the first language or an image frame immediately before the actor first begins speaking in the first language.

[0116] The method proceeds to optical flow determination 1006. For the original image frame 1002 and the corresponding modified image frame 1004, optical flow data 1008 is generated that determines how to displace pixels of the original image frame 1002 so that the displaced pixels approximately match the images of the modified image frame 1004. The optical flow data 1008 may indicate or encode the displacement or velocity for each of the pixels of the original image frame 1002 or for a sub-region of the original image frame 1002 in which the replaced object appears. Optical flow is typically used to estimate how an object moves in a sequence of image frames that contain footage of the object. In this case, optical flow is used instead to determine a mapping of pixel locations from the original footage of the object to pixel locations of a synthetic rendering of the object. This is made possible by the photorealistic rendering generated by the machine learning model described herein. The optical flow determination 1008 may be performed using any suitable method, for example phase correlation, block-based methods, differential methods, general variational methods or discrete optimization methods.

[0117] The method of Figure 10 continues with warping 1010, where pixels of the original frame 1002 are displaced in a direction indicated by the optical flow data 1008 to generate a warped original image frame 1012, and the optical flow data 1008 is used to displace pixels of the modified image frame 1004 in a direction opposite to that indicated by the optical flow data 1008 into the generated warped modified image frame 1014. This process warps objects appearing in the original image frame 1002 towards objects appearing in the modified image frame 1004, and vice versa. To warp the original image frame 1002 to the modified image frame 1004 in stages, the distances the pixels of the original image frame 1002 and the pixels of the modified image frame 1004 are displaced vary from image frame to image frame. At the start of the sequence, the original image frame 1002 is unmodified and the modified image frame 1004 is maximally warped (corresponding to pixels being moved by 100% of the distance indicated by the optical flow data 1008). At the next time step of the sequence, the pixels of the original image frame 1002 are displaced by a fraction F1 of the distance indicated by the optical flow data 1008 (e.g., F1=5%, 10%, 20% or other suitable value) and the pixels of the modified image frame 1004 are displaced by a fraction 100%-F1 of the maximum distance. At the next step of the sequence, the pixels of the original image frame 1002 are displaced by a fraction F2 of the distance indicated by the optical flow data 1008 and the pixels of the modified image frame 1004 are displaced by a fraction 100%-F2 of the maximum distance indicated by the optical flow data 1008, where F2>F1. This process continues stepwise with increasing percentages F1, F2, F3... at each time step until the last time step in the sequence where the original image frame 1002 is maximally warped and the modified image frame 1004 is unchanged. In this way, at each time step the pixels of the warped original image frame 1012 and the warped modified image frame 1014 approximately match each other.The percentages F1, F2, F3... may increase linearly with frame number or may increase according to another increasing function of the frame number.

[0118] The method proceeds to a dissolve 1016, where the warped original image frame 1012 is progressively dissolved into the warped modified image frame 1014 to generate a composite image frame 1018. This causes the composite image frame 1018 to transition from the original image frame 1002 at the start of the sequence to the modified image frame 1004 at the end of the sequence. For at least some time steps of the sequence, the dissolve 1016 may determine pixel values ​​for the composite image frame 1018 based on a weighted average of pixel values ​​of the warped original frame 1012 and the warped modified image frame 1014, where the weighting of the warped original image frame 1012 decreases for each time step and the weighting of the warped modified image frame 1014 increases for each time step. The weights of the warped original image frame 1012 may decrease from 1 to 0 according to a linear or non-linear function of the frame number, whereas the weights of the warped modified image frame 1014 may increase from 0 to 1 according to a linear or non-linear function of the frame number. Thus, an incremental dissolve is realized as an incremental interpolation between pixel values ​​of the warped original image frame 1012 and the warped modified image frame 1014.

[0119] The inventors have found that a more realistic transition that maintains image sharpness when warping from the original image frame 1002 to the modified image frame 1004 (or vice versa) can be achieved by concentrating the incremental dissolve 1016 in a central set of image frames where the incremental warping 1010 is performed. For example, the rate of the incremental dissolve 1016 may increase and then decrease relative to the rate of the incremental warping 1010. The incremental dissolve 1016 may be performed quickly relative to the incremental warping 1010, around the middle of the incremental warping 1010. The dissolve 1016 may start at a frame number that is later than the warping 1010 and end at a frame number that is earlier than the warping 1010, and / or the dissolve 1016 may be performed using a function that changes more rapidly than the warping 1010. In this way, the incremental dissolve is concentrated in the central few image frames where the incremental warping 1010 is performed. In one example, the incremental warping 1010 may be performed linearly, while the incremental dissolve 1016 may be performed with coefficients corresponding to a smooth step function or a sigmoid-like function that smoothly transitions from a substantially flat horizontal section of 0 to a substantially flat horizontal section of 1.

[0120] To illustrate the method, FIG. 11 shows an example of a sequence of original image frames O1-O5 and a sequence of modified image frames M1-M5. Optical flow data OF1-OF5 is determined for each time step, where the optical flow data for a given time step indicates an estimated warping relating the original and modified image frames. For example, optical flow data OF1 indicates an estimated warping for transforming the original image frame O1 into the modified image frame M1. The graphs show examples of how incremental warping and incremental dissolving are applied. For incremental warping, the coefficients on the vertical axis represent the distance the pixels are displaced as a percentage of the maximum distance indicated by the optical flow data. A coefficient of 0 means that the pixels stay in their original positions, and a coefficient of 1 means that the pixels are displaced by the maximum distance. For incremental dissolving, the coefficients on the vertical axis represent how much of the (warped) original image frame is replaced by the (warped) modified image frame. Coefficients of 0 correspond to the original (warped) image frame and coefficients of 1 correspond to the modified (warped) image frame.

[0121] In this example, warping is applied in linearly increasing increments, with the first warped frame being frame number 1. The dissolve is applied with a smooth step function. Before the most rapidly changing section of the smooth step function, the rate at which incremental dissolve occurs increases relative to the rate at which incremental warping occurs. After the most rapidly changing section of the smooth step function, the rate at which incremental dissolve occurs decreases relative to the rate at which incremental warping occurs. The incremental dissolve is centered within a central frame of the incremental warping. In this example, the rate of dissolve relative to the rate of warping increases and decreases smoothly, but in other examples, the rate of dissolve relative to the rate of warping may increase and decrease non-smoothly, for example, instantaneously.

[0122] Although the machine learning models described herein may be able to learn to recreate lighting and color characteristics that appear consistently in the training data, in some cases, the rendered instance of an object may fail to capture other lighting or color characteristics that vary locally or from one instance to another. This may occur, for example, when a shadow moves across an object in a movie scene. Such issues may be addressed using color grading, which changes visual attributes of the image, such as contrast, color, and saturation. Color grading may be performed manually, but is a time-consuming process that requires input from skilled VFX artists.

[0123] 12 shows an example of a method for performing automatic color grading that can be used instead of or in addition to the manual color grading described above. The method comprises processing video data including a sequence of original image frames 1202 and a sequence of modified image frames 1204, where in the modified image frames, instances of objects are modified using the techniques described herein. It is preferable that the color and lighting characteristics of the modified image frames 1204 closely resemble those of the corresponding original image frames 1202, although this may not be guaranteed for regions that include modified instances of objects. To achieve this, the method proceeds to optical flow determination 1206 to estimate a warping that relates the original instances of the object to the modified instances of the object. The estimated warping is indicated by optical flow data 1208, which may represent or encode the displacement or velocity for each of the pixels of the original image frame 1202 or sub-regions of the original image frame 1202 in which the replaced object appears. The method continues with warping 1210, where the optical flow data 1008 is used to displace pixels of the original frame 1202 in the direction indicated by the optical flow data 1208 to generate a warped original image frame 1212. Unlike the method of FIG. 10, which performs warping incrementally, the warping 1210 of FIG. 12 may be performed to the extent indicated by the optical flow data 1208 for the original image frame 1202. As a result, the pixels of the warped original image frame 1212 and the pixels of the modified image frame 1204 approximately match. In another example, the modified image frame 1204 may be warped to match the original image frame 1202. In yet another example, partial warping may be performed on the original image frame 1202 and the modified image frame 1204 (in either of these other examples, the pixels of the modified image frame 1204 need to be warped back to their original positions after the color grading process).

[0124] The method continues with blurring 1214, where a blur filter is applied to the warped original image frame 1212 to generate a blurred warped original image frame 1216, and a blur filter is applied to the modified image frame 1204 to generate a blurred modified image frame 1218. The blur filter may be a 2-D Gaussian filter, a box blur filter, or any other suitable form of low pass filter. The blur filter may have a finite size or a characteristic size in the range of a few pixels, for example, 3-20 pixels or 5-10 pixels. In the context of a 2-D Gaussian filter, the characteristic size may refer to the standard deviation of the Gaussian filter distribution. The effect of the blur 1214 is to remove high resolution details such that the pixels of the resulting image frame represent the surrounding colors within the region of those pixels. By choosing an appropriate size for the blur filter, local variations in surrounding color and lighting can be captured on a relatively short scale.

[0125] The method proceeds to color grading 1220, in this case using the blurred warped original image frame 1216 and the blurred modified image frame 1218 to modify color characteristics of the modified image frame 1204 to generate a color graded modified image frame 1220. Because the warped original image frame 1212 approximates the modified image frame 1204, the pixels of the blurred warped original image frame 1216 represent the desired ambient color of the corresponding pixels of the modified image frame 1204. Thus, the ratio of pixel values ​​of the blurred warped original image frame 1216 to the pixel values ​​of the blurred modified image frame 1218 represents the spatially-varying color correction map that is applied to the modified image frame 1204. Thus, color grading 1220 may be performed by dividing the blurred warped original image frame 1216 pixel-wise by the blurred modified image frame 1218 and multiplying the result pixel-wise by the modified image frame 1204 (or performing an equivalent mathematical operation). The resulting color graded modified image frame 1222 preserves the local color characteristics of the original image frame 1202 while preserving the fine scale details of the modified image frame 1222.

[0126] Figure 13 illustrates a movie production pipeline for a foreign language version of a movie in which visual dubbing is performed according to certain methods described herein. The solid arrows in Figure 13 represent the paths of video data, and the dashed arrows represent the paths of audio data. In this example, the production picture rush 1302 undergoes a face-off process 1304 in which instances of actors' faces are detected and isolated (possibly at reduced resolution). The resulting isolated instances of actors' faces are used for neural network training 1306. In this example, a separate neural network is trained for each actor who speaks in each scene (due to the fact that different scenes are likely to have different visual features).

[0127] While neural network training 1306 is taking place, the production picture rush 1302 and associated production audio rush 1308 are used in a first language (PL) edit workflow 1310, which includes an offline edit, where footage selected from the production picture rush for the final product is edited. The resulting offline edit (images and audio) is used to guide a second language (SL) recording 1312, which may include recording audio for multiple actors in a second language for multiple actors in a first language and / or multiple actors in a second language in multiple second languages. In this example, the SL recording 1312 includes video recordings and audio recordings. In other examples, the SL recordings may include only audio recordings. Additionally, the offline edit may be used to determine which instances of the first language actors' faces need to be translated.

[0128] The video and / or audio data obtained from the SL recordings 1312 are used as driving data for visual translation 1314, in this case using the neural network trained in 1306 to generate translated instances of photorealistic depictions of the actors' faces in the first language as needed for incorporation into the film. The resulting translated instances undergo a face-on process 1316 where the translated instances are combined with a full resolution master image. VFX 1318 are then applied as needed, followed by mastering 1320 of the full resolution master image and second language audio to create a final second language master image 1322 for distribution.

[0129] The above embodiments should be understood as illustrative examples of the present invention. Alternative embodiments of the present invention are envisioned. For example, in the context of visual dubbing, a machine learning model can be trained on footage of actors from various sources, such as various movies, and the machine learning model can be used later for visual dubbing of the actors in new movies. If a sufficiently expressive synthesis model (e.g., including a more sophisticated lighting model) is used, the methods described herein may be capable of generating renderings of photorealistic depictions of actors in scenes or movies with different visual characteristics, or renderings of actually different actors. In the latter case, a general machine learning model may, for example, be trained on many different instances of human faces, optionally in scenes with different visual characteristics, and may be capable of generating photorealistic renderings of a variety of different human faces, depending on parameter values ​​of the sufficiently expressive synthesis model. Furthermore, it should be noted that the specific techniques described herein for rendering modified instances of an object using synthesis models and machine learning models are merely exemplary, and that the methods described herein for replacing a first instance of an object with a modified first instance of an object may be used as well as other methods of generating a modified first instance. Some such methods utilize a combination of synthetic models and machine learning models such as the neural network models described above, while other methods may omit the use of an explicit synthetic model and instead model the object implicitly by a neural network, as is the case for neural radiance fields and certain GAN-based approaches such as StyleGAN and its variants. A neural network that is configured and trained to generate instances of photorealistic representations of objects based on explicit or implicit models of the objects may be referred to as a neural renderer.

[0130] The methods described herein may be used for detailed editing of objects other than human faces that appear in a movie. For example, the methods may be used to manipulate entire humans, animals, vehicles, etc. Additionally, deep inpainting may be used to composite the modified object back into the video, for example, if the object's contour moves as a result of the modification.

[0131] It is to be understood that any feature described in connection with any one embodiment may be used alone or in combination with the other features described, and any feature described in connection with any one embodiment may be used in combination with one or more features of any other embodiment or in any combination with other embodiments. Moreover, equivalents and modifications not described above may be used without departing from the scope of the invention as defined in the appended claims.

Claims

1. 1. A computer-implemented method for modifying video data comprising a sequence of image frames, comprising: isolating an original instance of an object within the sequence of image frames, wherein a geometry of the original instance of the object changes throughout the sequence of image frames; generating a modified instance of the object for the sequence of image frames using a machine learning model, wherein a geometry of the modified instance of the object changes throughout the sequence of image frames; generating a replacement instance of the object, wherein a geometry of the replacement instance of the object gradually transitions between a geometry of the original instance of the object and a geometry of the modified instance of the object over the sub-sequence of image frames; modifying the video data by replacing, for each image frame of the sequence of image frames, at least a portion of the original instance of the object with a corresponding at least a portion of the modified instance of the object; 1. A computer-implemented method comprising:

2. Creating the replacement instance of the object comprises: determining parameter values ​​for a composite model of the object, a first parameter value corresponding to the geometry of the original instance of the object; modifying the first parameter value of the composite model of the object to determine a second parameter value of the composite model of the object, the second parameter value corresponding to a geometry of the modified instance of the object; Stepwise interpolating between the first parameter value and the second parameter value across the subsequence of the sequence of image frames, thereby determining interpolated parameter values ​​for the synthetic model of the object; generating the replacement instance of the object based on the interpolated parameter values ​​of the synthetic model of the object using the machine learning model; The computer-implemented method of claim 1 , comprising:

3. The step of generating the replacement instance of the object, comprising: determining, for the subsequence of the sequence of image frames, optical flow data indicative of an estimated warping relating the instance of the object to the modified instance of the object; progressively applying the estimated warping to the original instance of the object across the subsequence of the sequence of image frames to determine progressively warped instances of the object; progressively applying the inverse of the estimated warping to the modified instance of the object across the sub-sequence of the sequence of image frames to determine the progressively warped modified instance of the object; progressively dissolving the progressively warped first instance of the object into the progressively warped modified first instance of the object across the subsequence of the sequence of image frames; The computer-implemented method of claim 1 , comprising:

4. The gradual dissolving is performed at a predetermined dissolve rate, the stepwise applying of the estimated warping and the inverse of the estimated warping is performed at a predetermined warping rate; 4. The computer-implemented method of claim 3, wherein the ratio of the dissolve rate to the warping rate increases to a maximum value within the subsequence of the sequence of image frames and then decreases.

5. determining optical flow data indicative of an estimated warping relating the original instance of the object to the replaced instance of the object across the sequence of image frames; applying the estimated warping to the instance of the object to determine a warped instance of the object; blurring the warped instance of the object; blurring the replaced instance of the object; adjusting a color of the replacement instance of the object based on a pixel-by-pixel ratio of the blurred warped instance of the object to the blurred replacement instance of the object to generate a color-graded replacement instance of the object; updating the replacement instance of the object to the color-graded replacement instance of the object before replacing at least a portion of the original instance of the object with a corresponding at least a portion of the modified instance of the object; The computer-implemented method of claim 1 further comprising:

6. 6. The computer-implemented method of claim 5, wherein blurring the warped instance of the object and blurring the modified instance of the object are each performed using a blur filter having a characteristic length scale of between 3 and 20 pixels.

7. The computer-implemented method of claim 1 , wherein the object is a human face.

8. The computer-implemented method of claim 7 , wherein at least some of the isolated instances of the object include a mouth but exclude eyes of a human face.

9. The sequence of image frames includes image frames of a human face speaking; 8. The computer-implemented method of claim 7, wherein a subsequence of the sequence of image frames precedes an image frame in which a human face is speaking.

10. detecting events in the sequence of image frames and / or an audio track associated with the sequence of image frames; determining one or more image frames of the sequence of image frames in which the detected event occurs; and determining the sub-sequence of the sequence of image frames in response to determining one or more image frames in which the detected event occurs; The computer-implemented method of claim 1 further comprising:

11. The computer-implemented method of claim 10 , wherein the object is a human face and the event relates to the human face speaking.

12. The computer-implemented method of claim 1, wherein the machine learning model includes a neural network.

13. the sub-sequence of the sequence of image frames is a first sub-sequence of the sequence of image frames; the gradual transitioning is from the geometry of the original instance of the object to the geometry of the modified instance of the object; 2. The computer-implemented method of claim 1, wherein the geometry of the replacement instance of the object gradually transitions from the geometry of the modified instance of the object back to the original instance of the object over a second subsequence of the sequence of image frames.

14. Creating the modified instance of the object comprises: determining parameter values ​​for a synthetic model of the object using the separated instance of the object; modifying parameter values ​​of the synthetic model of the object; Rendering the modified instance of the object using the trained machine learning model and the modified parameter values ​​of the synthetic model of the object; The computer-implemented method of claim 1 , comprising:

15. the sequence of image frames is a first sequence of image frames, the instance of the object is a first instance of the object, and the parameter value of the instance of the object is a second parameter value for a second instance of the object; identifying the second instances of each of the objects within a second plurality of sequences of image frames; For at least some of the identified second instances of the object, isolating the second instance of the object within the image frame containing the instance of the object; determining associated second parameter values ​​of the synthetic model of the object using the isolated second instance of the object; using the isolated second instance of the object and the associated second parameter values ​​of the composite model of the object, training the machine learning model to reconstruct the isolated second instance of the object based at least in part on the associated second parameter values ​​of the composite model of the object; The computer-implemented method of claim 14 further comprising:

16. A non-transitory storage medium comprising machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform a method of modifying video data comprising a sequence of image frames, the method comprising: isolating an original instance of an object within the sequence of image frames, wherein a geometry of the original instance of the object changes throughout the sequence of image frames; generating a modified instance of the object for the sequence of image frames using a machine learning model, wherein a geometry of the modified instance of the object changes throughout the sequence of image frames; generating a replacement instance of the object, wherein a geometry of the replacement instance of the object gradually transitions between a geometry of the original instance of the object and a geometry of the modified instance of the object over the sub-sequence of image frames; modifying the video data by replacing, for each image frame of the sequence of image frames, at least a portion of the original instance of the object with a corresponding at least a portion of the modified instance of the object; A non-transitory storage medium comprising:

17. A non-transitory storage medium for storing video data, the video data comprising: a first sequence of image frames including a photographic representation of an object; a second sequence of said image frames in which at least a portion of said photographic representation of said object is replaced with a corresponding at least a portion of a synthetic representation of said object; a third sequence of image frames between the first sequence of image frames and the second sequence of image frames, wherein at least a portion of the photographic representation of the object has been modified to gradually transition between at least a portion of the photographic representation of the object at the end of the first sequence of image frames and a corresponding at least a portion of the composite representation of the object at the beginning of the second sequence of image frames; and Non-transitory storage media, including

18. 20. The non-transitory storage medium of claim 17, wherein the synthetic representation of the object is a synthetic representation generated using a neural renderer.

19. modifying at least a portion of the photographic representation of the object comprises simultaneously warping and dissolving at least a portion of the photographic representation of the object into at least a portion of the composite representation of the object; the warping is performed in stages at a predetermined warping rate; The dissolving is performed in stages at a predetermined dissolving rate, 20. The non-transitory storage medium of claim 17, wherein the ratio of the dissolve rate to the warping rate increases to a maximum value within the third sequence of image frames and then decreases.

20. the composite representation of the object is a first composite representation of the object; modifying at least a portion of the photographic representation of the object comprises progressively interpolating between a second composite representation of the object and the first composite representation of the object; 20. The non-transitory storage medium of claim 17, wherein the second synthetic representation of the object corresponds geometrically to the photographic representation of the object.