Style-driven automated video editing

A learning-based framework using a transformer architecture predicts transitions and selects video clips to automate video editing, ensuring the content creator's intended style is maintained, thus addressing the challenge of automating video editing while preserving the creator's vision.

WO2025128665A1PCT designated stage expired Publication Date: 2025-06-19DOLBY LABORATORIES LICENSING CORP

Patent Information

Application Number
PCT/US2024/059513
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-24
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Automating the video editing process while maintaining the content creator's intended style is challenging due to the need to determine optimal camera angles, cuts, and transitions, which are highly dependent on the creator's style.

Method used

A learning-based framework that uses a transformer encoder-decoder architecture to predict transitions and select the next video clip based on a target style embedding, incorporating activation maximization to adapt the video to the desired style.

Benefits of technology

The framework effectively automates the video editing process by generating style-adapted transitions and selecting clips that align with the creator's intended style, reducing the time and effort required for manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059513_19062025_PF_FP_ABST
    Figure US2024059513_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems are described to create transitions between a set of video clips according to a style choice and to pick a next video clip in a sequence of ordered clips. In one embodiment, a device receives a plurality of video clips. In addition, the device may encode the plurality of video clips. The device further may compute a set of predicted transition embeddings using the target style and the encoded plurality of video clips. Furthermore, the device may determine a set of transitions for the plurality of video clips using the set of predicted transition embeddings in a recurrent manner.
Need to check novelty before this filing date? Find Prior Art

Description

STYLE-DRIVEN AUTOMATED VIDEO EDITINGCROSS REFERENCE TO RELATED APPLICATIONS[1] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 608,651, filed on 11 December 2023, and from European Patent Application No. 24 159 534.7, filed on 24 February 2024, which are both incorporated by reference herein in their entirety.TECHNOLOGY[2] The present invention relates generally to editing video. More particularly, an embodiment of the present invention relates to editing video using a style driven parameter.BACKGROUND[3] Digital cameras make it easier for content creators to record many versions (or takes) of a scene. Each new take can provide a unique camera framing or performance of the actor. At the end of this process, the content creator is left with a huge amount of camera shots and needs to perform the extremely time consuming and labor-intensive task of “video editing.” Video editing is a critical and creative stage in video production that involves assembling, arranging, and manipulating various video clips, audio elements, and visual effects to create a coherent and engaging final video. This process plays a pivotal role in shaping the overall quality, impact, and effectiveness of the video content.[4] The journey from recording different scene shots to making the final video involves dealing with the following three questions. (1) What camera angles to use: The choice of camera angles can convey emotions, perspectives, and storytelling elements. The selection of the shots with specific camera angles lays the foundation for how the content will be presented. (2) Which cuts to choose: Cuts are crucial for maintaining the video’s pacing, storytelling flow, and overall coherence. They determine which shot from the selected bag of camera angle shots to consider for shaping the narrative and visual rhythm of the video. (3) What gradual transitions to use: These include visual effects like fades, dissolves, and wipes to create a smooth and seamless connections between scenes or shots.They convey a change in time, location, or mood, and they contribute to the overall aesthetic and emotional impact of the video.[5] The answer to each of these questions is tied tightly with the content creator’s style of videomaking. All the three elements, e.g., camera angles, cuts and transitions can work together to ensure that the final video is a visual treat for the viewer. If the style intended changes, then answers to these questions change too. That’s where the challenge lies when you try build systems that can try to automate the process of video editing to some extent.BRIEF DESCRIPTION OF THE DRAWINGS[6] The present invention is illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements.[7] Figure 1 shows an example of a system that can be used in one or more embodiments the invention;[8] Figure 2 shows an example of creating a style adapted video from a bag of different camera shots;[9] Figure 3 shows an example of an illustration of a workflow to determine transitions for a sequence of ordered video clips;

[0010] Figure 4 show, in a flow diagram, an example of a process for determining that can be used with one or more embodiments of the invention;

[0011] Figure 5 shows, in a flow diagram, an example of a process for extracting a final set of clips that can be used with one or more embodiments of the invention.

[0012] Figure 6 shows an example of an ordered sequence of video clips that can be used with one or more embodiments of the invention;

[0013] Figure 7 shows an example of a workflow for determining pretrained embeddings that can be used with one or more embodiments of the invention;

[0014] Figure 8 shows an example of using contrastive learning to obtain style embeddings that can be used with one or more embodiments of the invention;

[0015] Figure 9 shows an example of a zero-shot way of obtaining style embeddings that can be used with one or more embodiments of the invention;

[0016] Figure 10 shows, in a flow diagram, an example of a process for choosing the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention;

[0017] Figure 11 shows, in a flow diagram, an example of a workflow to choose the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention;

[0018] Figure 12 shows an example of predicting transitions and the next clip that can be used with one or more embodiments of the invention;

[0019] Figure 13 shows, in a flow diagram, an example of rearranging video clips and choosing the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention;

[0020] Figure 14 shows an example of a workflow to rearrange video clips and choosing the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention; and

[0021] Figure 15 shows an example of a data processing system that can be used to perform or implement one or more embodiments of the invention.DETAILED DESCRIPTION

[0022] Various embodiments and aspects will be described with reference to details discussed below, and the accompanying drawings will illustrate the various embodiments. The following description and drawings are illustrative and are not to be construed as limiting. Numerous specific details are described to provide a thorough understanding of various embodiments. However, in certain instances, well- known or conventional details are not described in order to provide a concise discussion of embodiments.

[0023] Reference in the specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in conjunction with the embodiment can be included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in the specification do not necessarily all refer to the same embodiment. The processes depicted in the figures that follow are performed by processing logic that comprises hardware (e.g., circuitry, dedicated logic, etc.), software, or a combination of both. Although the processes are described below in terms of some sequential operations, it should be appreciated thatsome of the operations described may be performed in a different order. Moreover, some operations may be performed in parallel rather than sequentially.

[0024] The embodiments described herein can be used to create transitions between a sequence of ordered clips according to a style choice and to pick a next video clip in a set of video clips. In one embodiment, a method is described that can create transitions between pairs of video clips using a style input. In one embodiment, the method initially encodes each of video clips using an initial encoding and a positional encoding to create positionally encoded clip embeddings. The positionally encoded clip embeddings are input into a transformer encoder that outputs a latent embedding corresponding to the video clip. These latent embeddings can be style adapted using an activation maximization mechanism using a target style embedding. With the style adapted latent embeddings, transformer decoder decodes the latent embeddings to output embeddings at every time step t. As used throughout the specification, timestep is referring to the clips between which the output transition is being applied. The output embeddings are then compared to a pretrained set of embeddings to determine a corresponding transition. In addition, the method can determine, given an ordered set of video clips, a next video clip. Furthermore, the method can identify an order for a set of unknown ordered video clips.

[0025] As used herein the term ‘embedding’ denotes a learned deep feature or representation of an input video clip. An embedding can be a high-dimensional vector. In another context, the term embedding may denote a learned deep feature or representation of any kind of input data and can be a high-dimensional vector.

[0026] In one embodiment, digital cameras make it easier for content creators to record many versions (or takes) of a scene. Each new take can provide a unique camera framing or performance of the actor. At the end of this process, the content creator is left with a huge amount of camera shots and needs to perform the extremely time consuming and labor-intensive task of “video editing.” Video editing is a critical and creative stage in video production that involves assembling, arranging, and manipulating various video clips, audio elements, and visual effects to create a coherent and engaging final video. This process plays a pivotal role in shaping the overall quality, impact, and effectiveness of the video content.

[0027] The journey from recording different scene shots to making the final video involves dealing with the following three questions. (1) What camera angles touse: The choice of camera angles can convey emotions, perspectives, and storytelling elements. The selection of the shots with specific camera angles lays the foundation for how the content will be presented. (2) Which cuts to choose: Cuts are crucial for maintaining the video’s pacing, storytelling flow, and overall coherence. They determine which shot from the selected bag of camera angle shots to consider for shaping the narrative and visual rhythm of the video. (3) What gradual transitions to use: These include visual effects like fades, dissolves, and wipes to create a smooth and seamless connections between scenes or shots. They convey a change in time, location, or mood, and they contribute to the overall aesthetic and emotional impact of the video.

[0028] The embodiments described herein can be used in apparatuses which include one or more processors in a processing system, and which include memory and which are configured to perform any one of the methods described herein. Moreover, the embodiments described herein can be implemented using non-transitory machine- readable storage media storing executable computer program instructions which when executed by a machine cause the machine to perform any one of the methods described herein.

[0029] In one embodiment, an answer to each of these questions can be associated with the content creator’s style of videomaking. In this embodiment, the three elements, e.g., camera angles, cuts and transitions can work together to ensure that the final video is a visual treat for the viewer. If the style intended changes, then answers to these questions change too. That’s where the challenge lies when you try build systems that can try to automate the process of video editing to some extent.

[0030] In one embodiment, a learning-based framework to predict transitions that needs to be applied between all the clips of a video is presented. In this embodiment, this approach takes into consideration temporal dependency that exists between different transitions being used in the same video. The framework is a clip-level implementation to estimate transitions.

[0031] In a further embodiment, a mechanism to estimate transitions between clips of a video conditioned on the intended style is presented. In this embodiment, an activation maximization is leveraged to modify the choice of transitions being made based on style. In another embodiment, an approach to predict the next clip that should be used in the video with respect to the intended style. The task is analogous to a next frame prediction.However, the difference is in the fact that the approach operates on a clip-level, with style as the conditioning factor.

[0032] Figure 1 shows an example of a system 100 that can be used in one or more embodiments the invention. In one embodiment, the system 100 includes a video processing device 104 is coupled to a video clip datastore 102. In this embodiment, the video processing device 104 is a device that can analyze a sequence of ordered video clips and determine transitions according to a style parameter, or predict a next clip based on the intended style. In one embodiment, the video processing device 104 includes a video process 106 that can perform these actions. In this embodiment, the video process 106 can access a bag of video clips from the video clip store 108 on the video clip datastore 102. If the bag of video clips is ordered in time, the video process 106 can create transitions between these video clips using an intended style. In this embodiment, the video process 106 can store the generated transitions in the transition clips 110. In addition, the video process 106 can select a subset of the bag of video clips, where the subset is ordered, and select another video clip as the next video clip as per the intended style. In one embodiment, the video processing device 104 can be a server, personal computer, laptop, camera, smartphone, or another device that can process video clips. In one embodiment, the video clip 102 datastore can be a server, cluster, personal computer, laptop, camera, smartphone, or another device that can store video clips.

[0033] In one embodiment, video editing is the process of manipulating and rearranging video footage, audio, and visual elements to create a cohesive and engaging final product. It involves selecting and arranging shots, adding effects, music and sound effects and refining the content to convey a specific message, story, or concept. In this embodiment, video editing is a crucial step in the post-production phase of filmmaking, television production, online content creation, advertising, and various other forms of visual media. It shapes the way audiences perceive and interact with content, making it a key factor in the success of any video project.

[0034] In a further embodiment, a video cut refers to an editing technique in video production where one shot is instantly replaced by another shot. It involves an instantaneous switch from one visual element to another, resulting in a direct and immediate change. Cuts are used to maintain the pace, continuity, and flow of a video by movie from one scene to the next without any gradual effects or visual effectsbetween them. Examples of cuts include but are not limited to match cut, jump cut, dutch cut, cross-cutting, and cut-away.

[0035] Transitions are a widely used postproduction technique used to connect or move between shots in a more stylized or creative manner. They include effects that help smoothen the change from one shot to another, adding a visual and sometimes thematic link between the shots. There are several types of transitions, including fades, dissolves, wipes, crossfades and more, each with its own distinct effect on the viewer. Unlike cuts, transitions often involve a gradual blending, overlay or movement between shots rather than an immediate switch.

[0036] As is described below, a neural network can be used to determine either the transitions between video clips or the next video clip. In one embodiment, a SlowFast Network type of neural network can be used. Alternatively, a machine learning method can be used that takes a video input and outputs an embedding for temporal-related downstream tasks, such as action recognition, video prediction, or another temporal- related downstream tasks. For example, a Backbone type of machine learning method can be used (e.g., machine learning methods such as VideoMAE, VideoMAE v2, UniFormer, or another Backbone machine learning method can be used). SlowFast Networks are a class of learning-based designs utilized for tasks involving the analysis and comprehension of videos, such as recognizing actions within them. They are specifically crafted to proficiently grasp both swift and gradual movements present in videos, thus enabling adept scrutiny of actions and occurrences happening at varying time scales within a video sequence. The term "SlowFast" originates from the fusion of two distinct streams: one tailored to apprehend more gradual temporal nuances (slow pathway), and the other to capture quicker temporal details (fast pathway). In one embodiment, both the slow and fast pathways employ a 3D ResNet model, which involves processing multiple frames concurrently and applying 3D convolutional operations to them.

[0037] In addition, an activation maximization involves utilizing backpropagation through a neural network's weights to discover an arrangement of input values that would result in a desired alteration in the network's outputs.Given:- 0: Parameters of the neural network (comprising weights and biases)- hij (0,x): Activation of a specific unit z from a given layer j in the network for a given input xIn this context, if p is a defined bound value, then activation maximization can be defined as:

[0038] In one embodiment, the architecture of the transformer network uses an encoderdecoder framework. However, the transformer network is different from the conventional strategies that typically involve employing sequential arrangements of recurrent memory networks or computationally demanding convolutional networks. In place of these conventional approaches, the transformer introduces a multi-head self-attention mechanism. This mechanism that enables the network to capture intricate dependencies among elements positioned at various time steps within both the input and target sequences. This departure from sequential or convolutional processing contributes to the transformer's remarkable performance in a variety of tasks and its ability to handle long- range dependencies in data sequences more efficiently. In essence, the transformer leverages this attention mechanism to simultaneously consider relationships between all elements in a sequence, regardless of their temporal distance. This not only allows the model to effectively capture short-range interactions but also facilitates the understanding of long-range contextual information.

[0039] In one embodiment, consider a set of clips c1(c2, . . . , cm, where m < n that can be combined to form a video v in a way to adhere to the style s of the content creator using transitionsand cuts. Figure 2 shows an example of creating a style adapted video from a bag of different camera shots. In Figure 2, a bag of different camera shots 202 and a style tag 204 is used by a video process 210 to create a transition sequence(s) 212 and a set of cuts 206. In one embodiment, the transition sequence(s) 212 and the cuts 206 can be combined into a style adopted video 208. In this embodiment, the video process 210 generates the transition sequence(s) 212 using the style tag 204.

[0040] In one embodiment, the video process 210 can perform different action. In this embodiment, the actions are:1. Given a set of video clips (c1;c2, , cn) in the order the video clips are to appear in the video, find the best set of transitions (tr1 2, tr2 3, ... trn-l n) that can be applied to adapt the video for a content creator’s style s. This is further described in Fig. 3 below.2. Given the set of clips (c1;c2, ... , cm) in the order the video clips are to appear in the video, identify the order of clips in known bag of clips, (xltx2, ■ ■■ , xn-m) that can appear after cm. Additionally, find the best set of transitions (tr1 2, tr2 3, ... trn-l n) to be applied between the clips such that the video overall is adapted to the content creator’s style s. This is further described in Fig. 10 below.3. Given the bag of clips (c1;c2, ... , cn). identify the order of clips in unknown bag of clips, (%■£, x2, ... , xn-m) that can appear after cm. In one embodiment, first learn to choose the most optimal bag of clips and then how to arrange them. Additionally, find the best set of transitions (tr1 2, tr2 3, ... trn-l n) to be applied between the clips such that the video overall is adapted to the content creator’s style s. This is further described in Fig. 13 below.Typically, users can select one of the three actions based on their specific use cases or their preference for the level of automation desired in the editing process. While in one embodiment, audio is not used to determine the transitions, in alternate embodiments, audio can be used to determine the transitions (e.g., by using a different training set).

[0041] In one embodiment, generating the transitions is used for each of the three different actions. The subsequent embodiments can build onto this. In this embodiment, there can be some assumptions for determining transitions between the video clips: (1) The clips have been provided in the order they will appear in the final video; and (2) The clips will be used as it is and are not expected to be trimmed further or manipulated in any manner.

[0042] In one embodiment, there are two main sub-tasks that are performed: transition estimation and determining a style adherent transition. For transition estimation, a system estimates transition between any two clips appearing in the video. Given that these clips will together form a long video, it is important that the choice of transitions being made between any two clips considers the transitions that have occurred in the past.

[0043] In a further embodiment, for style adherent transitions, the choice of transitions needs to be modeled based on the required style. Creators have a diverse array of transitions to choose from, but their distinctive video-making style significantly shapes their selection. In this embodiment, the system can offer a fitting transition for the various clips. It’s worth noting that multiple styles might share applicabletransitions, meaning transitions within a style's set are not necessarily entirely distinct. Rather, the comprehensive choice and interaction of transitions throughout the entire video underscore the aspect of stylization. In simpler terms, a style is characterized by how and which transitions are utilized across the complete video, rather than just between two clips, unless a style's set includes one transition. Consequently, the interdependence among transition choices, as mentioned earlier, can play a role in conveying the intended style in the resulting video.

[0044] In one embodiment, to generate the set of transitions between successive pairs of video clips, there is a two-step process. The first step is training this framework without the style conditioning block. This helps understand and solve the first task of transition estimation. Second step would be to define the style conditioning block and solve for the style adherent transitions.

[0045] In one embodiment, a video is divided into small clips c1(c2, ... , cnby removing the parts of the video where the different transitions occur. These clips are in the order that the video clips are expected to appear in the video. Each clip can have different durations and different number of frames. In this embodiment, separating a video into multiple video clips is agnostic to these factors if the video encoder being used does not have any specific requirements such as minimum frame requirement or other criteria. If the video encode does, then obtaining the clips will change. However, once the clips are obtained, the generating the style adherent transitions is the same procedure.

[0046] Figure 3 shows an example of an illustration of a workflow to determine transitions for a sequence of ordered video clips according to a particular style which is referred to as target style. The choice of transitions is to be modeled based on the target style In Figure 3, system 300 includes a set of video clips 302, in the order the video clips are needed. System 300 represents a processing pipeline for the set of video clips 302. In one embodiment, the video encoder includes an initial encoder 304. The video clips of the plurality of video clips are fed into the initial encoder 304 as input. The initial encoder 304 outputs, for each of the plurality of video clips, a corresponding initial video clip embedding which is fed to the next element in the processing pipeline. In one embodiment, the next element in the processing pipeline is a positional encoder 306 receiving the respective initial video clip embedding as input. The positional encoder 306 outputs, for each initially encoded video clip, a corresponding video clip embedding. The video clip embeddings are concatenated to a singular video embedding for processing at the next element in the processingpipeline, the transformer encoder 308. In one embodiment, the purpose of an initial encoder 304 is to encode the different video clips together in a meaningful way to obtain an embedding that can be then passed to another encoder, e.g. the positional encoder 306 or directly the transformer encoder 308 for further processing. In one embodiment, an embedding is a learned deep feature or representation of an input video clip. In this embodiment, the embedding is a high-dimensional (1024, 2048, or another dimensional) vector. In another embodiment, the embedding generally means the learned deep feature / representation of any kind of input data, and usually is a high-dimensional vector. In one embodiment, any encoder helping achieve this goal can be used here. In one embodiment, Backbone can be used as the initial encoder 304 of choice. In this embodiment, the pretrained model available for Backbone has been obtained after training on a large dataset which makes the representation available from this encoder rich. It is equivalent to being referred to as the BERT for videos. In one embodiment, Backbone is trained on a large dataset, so using it as initial encoder 304 can extract useful feature of each video clip and make the learning of following models easier. Furthermore, irrespective of the video clip size, the embedding size obtained from the video encoder is constant. For example, and in one embodiment, the shape of the initial video clip embedding obtained from Backbone corresponding to a video clip is 1x768x1568. Therefore, as part of video encoding, the Backbone embeddings corresponding to each video clip , . . . , cnis predetermined. The singular video embedding that will be passed to the transformer encoder 308 will be obtained after concatenating each of the video clip embeddings obtained from the positional encoder 306, or in the absence of the positional encoder 306, by concatenating the initial video clip embeddings obtained from the initial encoder 304.

[0047] In one embodiment, the initial encoder 304 may require a minimum frame requirement. For example, the initial encoder may need the minimum number of frames in the video clip that needs to be encoded to be 16 (or some other threshold number of frames). This can be a problem because a number of the video clips in a dataset may have a number of frames less than 16 frames. Handling video clips with less than the requisite number of frames is described with reference to Figure 5 below.

[0048] In one embodiment, the embedding size obtained from initial encoder 304 for each video clip is 768 x 1568. The embedding size does not change based on video length, i.e. the number of frames in the video clip. In this embodiment, the video clips are passed as input to the framework, with the initial encoder 304 representing the first element of theframework. The initial encoder 304 of the video encoder will be computing the initial video clip embedding corresponding for each video clip and then combine them to get the overall singular video embedding to be passed on as input to the transformer encoder 308. Depending on the embodiment, the combining takes place after positionally encoding the initial video clip embeddings. Therefore, in the embodiment specifying an embedding size for each video clip of 768 x 1568, the overall size of the data, i.e. the singular video embedding, being processed in each batch of data would be a batch size of n x 768 x 1568, where n is the number of video clips in each video. Therefore, if n is high then the data size being worked with increase a lot and will be computationally expensive. Hence, a mechanism can be used to reduce the dimensions of the initial encoder embedding. In one embodiment, the video encoder can apply a pooling layer right after the initial encoder 304 and prior the next element in the processing pipeline, for example a mean pooling layer on the initial encoder embedding and reduce the number of feature dimensions, preferably to one. For example, the embedding size for each video clip is reduced to a size of 1x1568 instead of an embedding size of 768x1568. In another embodiment, other dimension reduction methods, like max pooling layer can be used. In other embodiments, the embedding size can be different sizes.

[0049] In the embodiment, where the video encoder 304 includes a positional encoder 306, positional encoding is employed. Positional encoding is a technique used in the field of natural language processing (NLP) and other areas involving sequence data to incorporate information about the positions of elements in a sequence. In this embodiment, instead of working with a sequence of words, the positional encoder 306 works with a sequence of video clips represented by corresponding initial video clip embeddings. The positional encoder 306 transforms the respective initial video clip embedding as input to obtain a corresponding video clip embedding as output that encodes information about the order of the respective video clip in the plurality of video clips. The positional encoder 306 is relevant for encoding the order of the video clips such that the processing pipeline is capable of converting a sequence of video clips to a sequence of transitions that needs to be applied between them. In one embodiment, this type of task can be similar to a machine translation task where the objective is usually to convert a sentence in one language to another. The difference is how the language is defined. Unlike other video-based works that exists in the literature today, the positional encoder 306 works on a “clip level” instead of “frame” level. The order in which these video clips appear is an important piece of information to determine theorder of transitions. Therefore, the video clips in a video can be viewed like the words in a sentence. In one embodiment, the embedding obtained from initial encoder for each clip is equivalent to having word level embeddings. In one embodiment, before proceeding further with the transformer encoder, the positional encoder 306 performs positional encoding on the clip-level embeddings to encode information regarding their order.

[0050] In one embodiment, with the initial and positional encoding completed, the result is a singular video embedding (e.g. 1 x 1568 or another dimension embedding) is passed to transformer encoder 308. In one embodiment, transformer encoder 308 takes in the singular video embedding comprising of positionally encoded video clip embeddings and outputs, for each video clip represented in the singular video embedding, a latent embedding corresponding to the respective video clip. While in one embodiment, Py Torch™ can be used for the transformer encoding, in alternate embodiments, another type of transformer encoder can be used. While the transformer encoder 308 receives a singular video embedding as input, the singular video embedding representing the entire plurality of video clips, the transformer encoder outputs a plurality of latent embeddings one after each other. Each latent embedding corresponds to a particular timestep t with an associated video clip. Each timestep represents an iteration of the processing pipeline after the transformer encoder 308.

[0051] In one embodiment, the output of the transformer encoder 308, the latent embedding corresponding to a video clip of the plurality of video clips, can be style adapted using an activation maximization mechanism as performed in style conditioning block 310 of the presented processing pipeline. This approach comprises (i) transformer decoding the respective latent embedding to obtain a corresponding estimated transition embedding, (ii) style conditioning the respective latent embedding using a target style embedding associated with the desired style to produce a corresponding style conditioned latent embedding, wherein the style conditioning employs activation maximization, and (iii) transformer decoding the style conditioned latent embedding to update the estimated transition embedding, thereby obtaining a predicted transition embedding that mimics the desired style. In one embodiment, this approach causes the output of the transformer encoder 308, i.e. the respective latent embedding, to be processed in two distinct branches. In a first branch, the latent embedding is fed into the transformer decoder 312 as input to output an estimated transition embedding which is described in greater detail below. In a second branch performed after having processed the first branch, the estimated transitionembedding corresponding to the currently processed latent embedding is used as input to the style conditioning block 310 as described in greater detail below. The style conditioning block 310 causes calculating a modified latent embedding which is referred to as style conditioned latent embedding. The style conditioning of the latent embedding in style conditioning block 310 causes a second run of the transformer decoder 312 for the same timestep but now based on inputting the style conditioned latent embedding as input, which causes an update of the estimated transition embedding which is now referred to as predicted transition embedding as it is the final estimate for the current timestep. When referring to estimated transition embeddings, only the updated estimated transition embedding is relevant, i.e. the predicted transition embedding. In this embodiment, moving closer towards the required style by making use of appropriate transitions moving to a stable region from an unstable region. In other words, the stable region can be thought of as the style positive region (e.g., if the video clip is in the stable region, then it means the video has characteristics of the requested style) and the unstable region to be style negative region (this means that the video does not showcase characteristics of the requested style). The way to make the video move towards the style positive region is by changing some or all the transitions currently being used in the video. Therefore, the system would know the sequence of transitions that can potentially be applied between the different clips of a video. Lastly, the system can introduce one or more modifications to this sequence to bring the overall video closer to having the required style’s characteristics. Each point in the trajectory in this problem would be the combination of transition and cuts that can be applied between clips t and t+1. The trajectory optimization can be done using the concept of activation maximization.

[0052] In one embodiment, a style conditioning block 310 conditions the latent embedding according to requested style. In this embodiment, given a target style embedding estyie, the aim is to produce a transition embedding configuration with a sequence of transition embeddings (e0, e1(e2, ... eT-1) that drive the overall video closer to exhibiting the required style computing a distance, for example the Euclidean distance, between the target style embedding estyieand an averaged estimated transition embedding e / (, e.g., by computing 11 e / (— estyie11 . The target style embedding estyiemay be determined from one or more style video clips corresponding to the desired style. The averaged estimated transition embeddings e / (represents the video clips from the beginning up to the currently processed timestep t, i.e. is representative of the video clips associated with timesteps 0...t.In one embodiment, the averaged estimated transition embedding e / (may be the average of a plurality of estimated transition embeddings as determined from previous runs of the transformer decoder 312, i.e. the estimated transition embeddings correspond to latent embeddings processed in earlier iterations of the transformer decoder 312. For the preceding latent embeddings in the processing pipeline, the estimated transition embeddings are the predicted transition embeddings 316, i.e. the updated transition embeddings obtained in the second run of the transformer decoder 312, while for the most recent latent embedding as currently processed the estimated transition embedding from the first (and up to now) only run is to be considered. In some embodiments, the average of the plurality of estimated transition embeddings is the average of the transition embeddings estimated for all latent embeddings preceding and up to the respective latent embedding which is currently style conditioned. This is described further below. The number of transitions is T assuming the number of clips in the video are T+l. The style conditioning block 310 iteratively employ activation maximization T times to achieve this. Applying activation maximization causes minimizing the computed distance between the target style embedding estyieand the averaged estimated transition embedding e / (. The style conditioning block 310 causes adjusting the respective latent embedding ztbased on a scaled gradient of the computed distance to generate a style conditioned latent embedding z^. By subtracting the scaled gradient of the computed distance in each of the iterations to process all timesteps, the computed distance will be iteratively minimized. In one embodiment, the equations used for the style conditioning block 310 are:Here ztis the latent embedding obtained from the transformer encoder for video clip t and a is a user-defined parameter, k denotes the timesteps 0...t processed so far in the processing pipeline. In one embodiment, a can be used by the user to decide the extent to which style needs to be emphasized while estimating transitions.

[0053] In order to implement these equations, estyieshould be determined. In one embodiment, contrastive learning can be used to obtain style embeddings. In this embodiment, Figure 8 gives an overview of the system that can potentially be used to learn embeddings corresponding to a style. The basic idea revolves around leveraging videos exhibiting the style required and a Large Language Model (LLM) (e.g., a Contrastive Language-Image Pre-training (CLIP)) embeddings obtained from the text input to obtain ajoint rich embedding corresponding to a style. The embedding will not only capture the style content visually present in the videos but also certain additional information that LLMs can provide.

[0054] Figure 8 shows an example of using contrastive learning 800 to obtain style embeddings that can be used with one or more embodiments of the invention. In Figure 8, the input (802) consists of videos that were collected corresponding to particular styles (SAVE videos). The styles videos are encoded (804) using the encoder that outputs a corresponding embedding. In one embodiment, Backbone can be used for the encoder when this encoder is used for the initial encoder above. In another embodiment, the same or different encoder can be used for (804) as in the initial encoder.

[0055] In one embodiment, a fully connected layer (806) is employed with the purpose of aligning the dimensions of video embeddings and text embeddings, which can bring the embeddings to a common dimensionality. In addition, the linear layer is employed to enrich the embeddings as per above. With the two sets of embeddings, a contrastive learning is performed to learn the shared embedding space for both video and text.

[0056] In another embodiment, a style tag 812 is inputted to a LLM module 814. In this embodiment, the style tag is a text / sentence describing the desired style, e.g., "documentary style" or "YouTube UGC style". This describes the style of entire video instead of the classes of transitions. The LLM module 814 is used to obtain embeddings corresponding to the text informing the system about the required style. An example from a trial is presented below. The input to the LLM model (e.g., via the LLM module 814) is a generic video summary describing a coffee commercial. The LLM model was asked it to then stylize the text as if it was being written by the director Christopher Nolan, in one embodiment, that the output text seems to decently capture Nolan-like style signatures. Similar results were observed in other trials as well. This shows that in the process of being trained on vast amounts of text data, these LLMs have manage to learn certain concepts around stylization with respect to different categories and people. The idea behind using LLM embeddings was to capture exactly this while developing a style embedding.

[0057] The results from the LLM module 814 is input to a projection layer (816) that is employed with the purpose of aligning the dimensions of video embeddings and text embeddings, which can bring the embeddings to a common dimensionality. In addition, the linear layer is employed to enrich the embeddings as per above. With the two sets of embeddings, a contrastive learning is performed that outputs style embeddings. In oneembodiment, contrastive learning here brings the video embedding and text embedding to a shared embedding space. Given training data of video-text pairs, contrastive learning minimizes the distance between video and text that are paired and maximizes the distance between video and text that are not paired.

[0058] In another embodiment, a zero-shot way can be used to obtain the target style embedding from one or more style video clips including transitions representing the same particular target style. Figure 9 shows an example of a zero-shot way of obtaining a target style embedding that can be used with one or more embodiments of the invention. In Figure 9, the inputs 902 to the framework was a set of style video clips and the output was transition embeddings 910 at every time step t. The zero- shot way of obtaining style embeddings would be to pass one or more style video clips, all corresponding to the same particular style, and obtaining the corresponding first embedding 910 from the decoder 908. Using this workflow, there would be an embedding corresponding to each of the one or more style video clips. The overall style embedding 912 in that case would be the mean of the embeddings obtained for each of the one or more style video clips.

[0059] As above, using a zero-shot way of obtaining the style embeddings is similar to the workflow of Figure 3 above, without the style adapting. Instead, selected input 902 of a known style is used to generate style embedding 912. This style embedding is further used above. The pipeline is processed for each style video clip of the one or more style video clips. In one embodiment, the input 902, the respective style video clip to be processed, is fed into the video encoder 904, which includes an initial encoder 916 and positional encoder 918. In one embodiment, the initial encoder 916 and 918 performs as described above with reference to Figure 3. The output of the video encoder 904 is fed to the transformer encoder 906. As described above, the transformer encoder 906 takes in the positionally encoded clip embeddings and outputs a latent embedding corresponding to the video as in described in above.

[0060] In addition, the latent embedding is passed to a decoder 908. In one embodiment, the decoder 908 is a transformer decoder as described below with reference to Figure 3. The output of the decoder 908 are the predicted transitions embeddings 910 for each timestep t. After having processed all style video clips of the one or more style video clips associated with a particular target style, by y taking the mean of all predicted transition embeddings, the final predicted transition embedding 912 is computed and may serve as target style embedding.

[0061] Returning to Figure 3, the respective style conditioned latent embedding is fed into the transformer decoder 312 to obtain the predicted transition embedding 316 corresponding to the currently processed latent embedding. This is the second run when the transformer decoder 312 is invoked when processing a latent embedding at timestep t. The first time, the transformer decoder 312 processes the latent embedding as it was output from the transformer encoder 308 to obtain an estimated transition embedding. The estimated transition embedding corresponding to the currently processed latent embedding is used as input in the style conditioning block 310 as described above to obtain the style conditioned latent embedding. The predicted transition embedding output from the transformer decoder 312 is an update of the estimated transition embedding of the first run and used at following iterations in the activation maximization approach, i.e. when processing the next latent embeddings for following timesteps. In one embodiment, the transformer decoder 312, which bears similarity to the transformer encoder 308 in terms of its foundational structure comprising multiple layers. However, the transformer decoder 312 employs an approach involving masked multi-head self-attention within a sequence. This tailored attention mechanism ensures that the attention for each element is focused exclusively on elements preceding it in the sequence. This mechanism guarantees causality, making it applicable during testing when the complete target sequence isn't known beforehand. The process begins with the masked multi-head self-attention (MMH) operation. Here, individual self-attention operations (SA) are carried out by each head, and their outputs are concatenated. This concatenated result is transformed through a weight matrix (Wc ncat). This step ensures the sequence-level attention mechanism considers only past elements, preserving the autoregressive property essential for generating sequences.The next step involves another multi-head self-attention mechanism (MH), where the output of the MMH operation functions as both the key and value, while the encoded representation (Fw) serves as the query. This additional self-attention layer lacks masking, allowing it to incorporate future information based on the key-value relationships established by the MMH operation. The outcomes of this MH operation are concatenated and subjected to further transformation via a weight matrix (Wconcat). Subsequently, the output of this multi-head self-attention is passed through two fully connected layers, finalizing this processing block. To encapsulate, a decoder block in the transformer structure follows the sequence of operations: self-attention (SA), masked multi-head self-attention (MMH), self-attention (SA), multi-head self-attention (MH), and fully connected (FC) layers.The entire transformer network consists of N such blocks, each block designed as an (SA- MMH-SA-MH-FC) sequence. Additionally, positional encoding is introduced to the target sequence upfront, contributing vital temporal information. Furthermore, residuals around each layer within these blocks are employed during backpropagation, aiding the network's learning process by mitigating vanishing gradient issues and enabling smoother training convergence.

[0062] The transformer decoder 312 outputs an estimated transition embedding e for every time step t by transforming the currently processed latent embedding. The transformer decoder 312 is invoked twice due to involving the style conditioning block and providing a style conditioned latent embedding for which the transformer decoder 312 outputs an updated estimated transition embedding which is referred to as predicted transition embedding 316. Timestep is basically referring to the video clips between which the transition corresponding to the estimated transition embedding is being applied. Therefore, at timestep / , the output transition embedding is referring to the transition between clip t and t+1. This output is passed back in step 314 into the transformer decoder 312 to estimate the next transition embedding when processing the next latent embedding in the following timestep. This helps the framework to know which transitions have been applied in the past and understand the temporal dependency that exists between the different transitions for the same video. At time t=0, e.g., when there is no past transition available, a vector of zeros is passed in having the same shape as the transition embedding vector.

[0063] At this stage, in order to obtain the class of the transition being predicted by the decoder, the system makes use of the pretrained transition embeddings received from training the transition classifier. The pretrained transition embeddings are expected to be ranked by their distances with et. e) is then assigned the corresponding transition category. To achieve this, the system utilizes the triplet margin loss to optimize the similarity between the transition embedding and e. For the embedding etcorresponding to every time step t for a video with ground truth label c, the training objective is defined with a triplet loss asWhere T calculates the triplet margin loss for each triplet (et, erans, e^ransy.T a, p, r) = max ( (a, p) — (p (a, ) + M, 0)M is the soft margin (used 0.3 in this work), a, p and n are anchor, positive sample, and negative sample, respectively. <p follows the following definition -In one embodiment, the system takes the predicted embedding etas the anchor, the pretrained transition embedding with category c as the positive sample, other pretrained embeddings as negative samples. This results in:

[0064] By optimizing the above equation, the model encourages the similarity between the embedding etand its ground truth transition embedding higher than the similarity between non-matching pairs with a margin of M. In one embodiment, transition pretrained embeddings being used in the masked triplet loss have been obtained after training a transition classifier. The transition classifier is further described in Figure 7 below. In one embodiment, the pretrained embeddings are the pretrained embeddings as described in Fig. 3, block 318 above.

[0065] Figure 4 shows, in a flow diagram, an example of a process 400 that determines transitions that can be used with one or more embodiments of the invention. In one embodiment, process 400 begins by receiving a set video clips at block 402. In one embodiment, the set of video clips is an ordered set of video clips as described above in Figure 3. At block 404, process 400 encodes the set of video clips. In one embodiment, process 400 performs an initial encoding and positional encoding as described in Figure 3, blocks 304 and 306 above. In this embodiment, process outputs an initial embedding. Process 400 transformer encodes the initial embedding into a positionally encoded clip embedding at block 406. In one embodiment, process 400 transformer encodes the positionally encoded clip embeddings and outputs a latent embedding, as described in Figure 3, block 308 above.

[0066] At block 408, process 400 style conditions that latent embedding. In one embodiment, process 400 style conditions that latent embedding as described in Figure 3, block 310 above. Process 400 transformer decodes the style conditioned latent embeddings, at block 410. In one embodiment, process 400 transformer decodes the style conditioned latent embeddings as described in Figure 3, block 312 to output the reduced transition embedding. Process 400 returns the predicted transition clips at block 412.

[0067] Figure 5 shows, in a flow diagram, an example of a process 500 for extracting a final set of clips that can be used with one or more embodiments of the invention. As per above, the initial video encode may require that each clip have a minimum number of frames in each clip. If any of the clip ckhappens to have less than the threshold number of frames, the system combines ckwith the subsequent clips until the system has a clip Ckof at least the threshold number of frames. Combining clips would also mean that the system retains the transition occurring between these clips used to form this Ck. Therefore, if originally, the system reported N transitions in the video, the system will now report N — p transitions where p is the number of clips used to form Ck.

[0068] In one embodiment, the following two other potential ways of approaching this issue of handling clips having less than the threshold number of frames were also discussed: (1) video frame interpolation and (2) and dropping clips.

[0069] In one embodiment, video frame interpolation is a technique used to generate the intermediate frames between existing frames in a video sequence. This process can be used to smooth out motion and create illusion of higher frame rates. Interpolated frames are calculated based on the motion and content of adjacent frames. This technique is particularly useful when converting videos from lower frame rates (e.g., 24 or 30 frames per second) to higher frame rates (e.g., 60 or 120 frames per second) for smoother motion. Frame interpolation algorithms typically use various approaches such as optical flow, neural network, and mathematical interpolation methods, to predict the appearance of frames between the original frames. This can enhance the visual quality of videos, especially in scenarios like slow-motion effects or when content is played back on displays with higher refresh rates.

[0070] In one embodiment, while this method could have helped the system to increase the number of frames and enable each invalid clip to have more than 16 frames, the system can be used to determine the quality of the frames being created and where to insert model frames. (A) Ensuring quality of the frames being created - The frames being generatedshould be meaningful with respect to the video they were being inserted in. This would require the system to build an additional module to just do quality check. Relatively poor quality of frames as well incorrect visual content can dramatically affect the data distribution associated with the different transitions and hence affect the training process. In one embodiment, the system is used to minimize manipulating the content for this reason. (B) Where to insert more frames - Given a clip, it was a challenge to decide between which frames should the interpolation be done. Location could matter as it could again affect the data distribution as well as the meaning of the content being showcased in the clip. Transition choices do depend on understanding the visual content of the clips between which it is being applied.

[0071] In a further embodiment, the system can be used to drop clips. In this embodiment, this system drops a video clip wherein the number of frames is less than 16 or another threshold for minimum number of frames for each clip. In other words, the system skips the clips not satisfying the required criteria. Doing so reduced the data size by 3000. Additionally, the location of the invalid clips can also create a problem. The invalid clips can prevent the system from using clips appearing before and / or after them. This might further reduce the dataset size. Taking all these concerns into consideration, the system can combine clip.

[0072] In Figure 5, process 500 begins by analyze transition metadata at block 502. In one embodiment, the autotransition dataset provides information about the start and end time stamps of the transition frames. This allows process 500 to know the duration of video clips right before and after the transition. Process 500 uses this information and the video frame rate to compute the number of frames in each of these clips. At block 504, process 500 finds the invalid clips. In one embodiment, the information regarding the number of frames in each potential clip of the video allows process 500 to decide which of these are invalid. While in one embodiment, a clip as invalid if it contains < 16 frames, in alternate embodiments, an invalid clip may have more or less number of frames.

[0073] Process 500 adjusts a timestamp of each invalid clip at block 506. In one embodiment, adjusting an invalid clip timestamp can explain the process opted for combining clips. The input to process 500 is the timestamps of the original set of clips and output is the timestamps of final set of clips to be considered. In one embodiment, given a video, process 500 traverses through the clips in the order they appear. Assume that process 500 is processing a clip cL. In this embodiment, consider the following main cases for each clip. Case 1: clip ctis valid, ci+1is valid. Since, the sequential clips are valid,process 500 can consider the pair of clips c(-, ci+1in the dataset and consider the transition between this pair of clips. Case 2: clip ctis valid, ci+1is invalid. This can be a problem because process 500 cannot consider the transition between a and ci+1unless process 500 makes the latter valid. If process 500 cannot then a will not be useful and will therefore become invalid. Therefore, process 500 checks if frames ci+1+ frames ci+2> 16 ? If yes, then a is valid. If not, then mark a as doubtful. Case 3: ct< 16. In this case, process 500 traverses through the clips ci+1onwards and continue till the number of total frames is > 16. For example, assume that clips : ci+6> 16 frames. Then, process 500 groups C . c(+6makc a new clip . Therefore, process 500 modifies the timestamps of ctaccordingly and jump to analyze ci+7directly. However, if process 500 does not find Cj then ctis considered invalid. Once process 500 has marked the clips accordingly, then process 500 can re-iterate through the doubtful clips and re-assign them as valid and invalid.

[0074] At block 508, process 500 extracts the final set of clips from video. In one embodiment, process 500 uses the final timestamps obtained from the previous step for the different timestamps.

[0075] Figure 6 shows an example of an ordered sequence of video clips 600 that can be used with one or more embodiments of the invention. In Figure 6, there are six video clips 602A-F in the ordered sequence of video clips 600.

[0076] Figure 7 shows, in a flow diagram, an example of a framework 700 for determining pretrained embeddings that can be used with one or more embodiments of the invention. In one embodiment, the framework 700 is used to obtain the transition embeddings corresponding to different classes. In Figure 7, a set of video transitions 702 are input into the backbone 704. In one embodiment, the backbone 704 is a neural network employed to transform the frames of the video into a singular embedding. The neural network can be a SlowFast network or another type of neural network. The singular embeddings outputted by the backbone 704 are input to the linear layer 706. In one embodiment, the linear layer is a fully connected layer that is used for projecting the features obtained from the backbone network to a dimension that is easier to work with.

[0077] The output of the linear layer 706 is processed with a normalization module 708 that normalizes the embedding. The embedding that is output form the normalization is the transition embedding 710. The size of this transition embedding corresponding to one sample (irrespective of its size) is 1x2048 (or another normalization size). The transitionembeddings are processed by a linear classifier 712. In one embodiment, the linear classifier 712 is a MLP consisting of a fully connected layer. The first layer has an input dimension of the transition embeddings, and the last layer output dimensions equal the number of classes or categories of transition. In one embodiment, framework 700 is used to train the transition classifier. The output of the transition classifier network is a transition class 714.

[0078] In one embodiment, once the transition classifier network is trained, framework 700 obtains the embedding weights corresponding to each of the transitions being considered in the training data from the linear classifier module. In this embodiment, framework 700 is trained on 30 classes. Keeping in mind the requirement of the future steps, framework 700 can be trained this framework on 31 classes. These classes include - direct_cut, pull_in, mix, pull_out, circle, open, windmill, cube, switch, left, pane, circle, right, black_fade, tum_page, clock_wipe, blinds, heart, squeeze, floodlight, down, kaleidoscope, memory, white_flash, memory, blur, gradient_wipe, superimpose, dissolve, star, and blanch. Framework 700 can also be trained on a greater number of classes, if training data is available.

[0079] As described above, the next clip can be computed using the parts of the analysis from Figure 3, because in Figure 3, the system has determined an understanding of how transitions relate to style. In Figure 10 below, the system gains an understanding of cuts, e.g., the arrangement of clips of the video and its relation to style. In one embodiment, the system can comprehend how the arrangement of clips is influenced by the style input when there is no explicit direction provided regarding their organization. In this embodiment, an iterative zero- shot framework can be used that leverages the process Figure 4 above. Assume there are k clips and wish to find the k+1 clip. In one embodiment, the system includes a process that determines the next clip choice using a style input. In Figure 3, one of the assumptions made is that the clip arrangement is according to style. Here, however, the choice of k+1 clip needs to take style into consideration. Therefore, in this embodiment, the process can do a grid search mechanism wherein the process picks a potential k+1 clip and obtain the e when clips 1 to k+1 is passed into the step 2 framework. In addition, the process compares the e with the style embedding. Every clip in the bag of clips is processed in this way. The clip that has ethaving least distance from estyieis taken as the next clip. In one embodiment, once the next video clip is determined, transitions for the updated ordered sequence of video clips can be determined using process 300 as described in Fig. 3.

[0080] Figure 10 shows, in a flow diagram, an example of a process 1000 for choosing the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention. In Figure 10, In one embodiment, process 1000 begins by receiving a set video clips at block 1002. In one embodiment, the set of video clips is an ordered set of video clips as described above in Figure 3. At block 1004, process 1000 encodes the set of video clips. In one embodiment, process 1000 performs an initial encoding and positional encoding as described in Figure 3, blocks 304 and 306 above. In this embodiment, process 1000 outputs an initial embedding. Process 1000 transformer encodes the initial embedding into a positionally encoded clip embedding at block 1006. In one embodiment, process 1000 transformer encodes the positionally encoded clip embeddings and outputs a latent embedding, as described in Figure 3, block 308 above.

[0081] Process 1000 transformer decodes the latent embeddings, at block 1008. In one embodiment, process 1000 transformer decodes the style conditioned latent embeddings as described in Figure 3, block 312 to output the reduced transition embedding. In one embodiment, unlike with computing the transition clips as described in Figure 4 above, process 1000 computes the distance between the embedded transition clips and the style embeddings at block 1010. In one embodiment, process 1000 computes a Euclidean distance between predicted transition embeddings using the equation:||e ~ estyie | |2In this embodiment, the computed distances can be used to choose a transition embedding the closest corresponds to a selected style and used to choose the next video clip in the sequence. At block 1012, process 1000 stores the computed distances for a later comparison. Process 1000 chooses the next potential video clip at block 1016. In one embodiment, process 1000 uses the stored distance to determine which of the video clips has a transition that is closest to the style embedding. In one embodiment, the video clip corresponding to the smallest distance between a predicted transition embedding and style embedding is chosen as the next video clip.

[0082] Figure 11 shows an example of a workflow 1100 to choose the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention. In Figure 11, the workflow begins by receiving input video clips (1102). In one embodiment, these video clips are ordered. The video encoder 1104includes an initial encoder and a positional encoder as described in Figure 3, blocks 304 and 306 above. The output of the video encoder is a positionally encoded clip embedding. A transformer encoder 1106 encodes the positionally encoded clip embedding into a latent embedding. A transitional decoder 1108 receives the latent embeddings and outputs the predicted transition embedding 1110. In one embodiment, the predicted transition embedding 1110 has a dimension of 1x2048 (or another dimension). The workflow 1100 continues, in one embodiment, by computing a distance between the predicted transition embedding and the style embedding (1112). In one embodiment, the computed distance is a Euclidean distance (or another type of distance). The computed distances 1112 are stored for later comparison (1114). The next potential video clip from a set of unordered video clips (1120) is chosen (1118), in one embodiment, using the stored computed distances. The next potential video clip chosen can be appended to the input video clips (1102) so that the process can be done iteratively to output a sequence of next video clips in a growing window manner.

[0083] Figure 12 shows an example of predicting (1200) transitions and the next clip that can be used with one or more embodiments of the invention. This approach is an evolution of the action 1 framework (as described above), where a notable departure is made: employing two decoders instead of the initial singular decoder. The first decoder is dedicated to approximating the transition embedding, whereas the second decoder focuses on approximating the embedding that corresponds to the subsequent video clip. The embedding hk+1obtained has shape 1x1568, same as the shape of the embedding obtained from backbone. This embedding is passed into the video encoding module along with the previous k clip embeddings of the video. The method involves a two-step training process in which the decoder is trained to predict the next clip and then predict train the decoder to predict the transitions.

[0084] In addition, and in one embodiment, the system that can predict a next video clip can be adapted to rearrange a set of clips using the style embeddings. In this embodiment, the system does not know the bag of xltx2, ... xkvideo clips. Therefore, using existing style features and other user defined parameters, the system will use a retrieval method to obtain this set first and then proceed to rearrange them in the order that aligns with the intended style and consider appropriate transitions. Figure 13 shows, in a flow diagram, an example of a process 1300 for rearranging video clips and choosing the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention. In Figure 13, in one embodiment, process 1300begins by receiving a set video clips at block 1302. In one embodiment, the set of video clips is an ordered set of video clips as described above in Figure 3. At process 1300, block 1304encodes the set of video clips. In one embodiment, process 1300 performs an initial encoding and positional encoding as described in Figure 3, blocks 304 and 306 above. In this embodiment, process 1300 outputs an initial embedding. Process 1300 transformer encodes the initial embedding into a positionally encoded clip embedding at block 1306. In one embodiment, process 1300 transformer encodes the positionally encoded clip embeddings and outputs a latent embedding, as described in Figure 3, block 308 above.

[0085] Process 1300 transformer decodes the latent embeddings, at block 1310. In one embodiment, process 1300 transformer decodes the latent embeddings as described in Figure 3, block 312 to output the predicted transition embedding. Process 1300 computes the distance between the embedded transition clips and the style embeddings at block 1312. In one embodiment, process 1300 computes a Euclidean distance between predicted transition embeddings using the equation:||e ~ estyie | |2

[0086] In this embodiment, the computed distances can be used to choose a transition embedding the closest corresponds to a selected style and used to choose the next video clip in the sequence. At block 1314, process 1300 stores the computed distances for a later comparison. Process 1300 rearranges the video clips at block 1316. In one embodiment, because process 1300 uses existing style features and other user defined parameters, process 1300 will use a retrieval method to obtain this set 1420 first and then proceed to rearrange them in the order that aligns with the intended style and consider appropriate transitions. To be specific, the retrieval method retrieves a set of clips 1420 from the possible video clips. The retrieval method can be user defined to enforce certain properties, such as chronological order or other distribution, of the video clips. For example, if a user has 27 video clips where 9 are from morning, 9 are from afternoon and 9 are from night, the user may apply a constraint such that the final video consists and in the order of 3 clips from each time of the day. In this case, the retrieval method will return different set of video clips depending on the current of iteration n in Fig. 14. For another example, the user may want to limit the number of video clips of certain scene in the final video. In this case, when the number of such clips is reached in the selected video, the retrieval method will not retrieve such clipfrom all possible video clips. In brief, the method allows any retrieval methods defined by user.

[0087] Process 1300 chooses the next potential video clip at block 1318. In one embodiment, process 1300 uses the stored distance to determine which of the video clips has a transition that is closest to the style embedding. In one embodiment, the video clip corresponding to the smallest distance between a predicted transition embedding and style embedding is chosen as the next video clip.

[0088] Figure 14 shows an example of a workflow to rearrange video clips and choosing the next video clip for a sequence of ordered video clips that can be used with one or more embodiments of the invention. In Figure 14, the workflow 1400 begins by receiving input video clips (1402). In one embodiment, these video clips are returned from a user-defined retrieval method. The video encoder 1404 includes an initial encoder and a positional encoder as described in Figure 3, blocks 304 and 306 above. The output of the video encoder is a positionally encoded clip embedding. A transformer encoder 1406 encodes the positionally encoded clip embedding into a latent embedding. A transitional decoder 1408 receives the latent embeddings and outputs the predicted transition embedding 1410. In one embodiment, the predicted transition embedding 1410 has a dimension of 1x2048 (or another dimension). The workflow 1400 continues, in one embodiment, by computing a distance between the predicted transition embedding and the style embedding (1412). In one embodiment, the computed distance is a Euclidean distance (or another type of distance). The computed distances 1412 are stored for later comparison (1414). The next potential video clip from a set of unordered video clips (1420) is chosen (1418), in one embodiment, using the stored computed distances. The next potential video clip chosen can be appended to the input video clips (1402) so that the process can be done iteratively to output a sequence of next video clips in a growing window manner.

[0089] Figure 15 shows an example of a data processing system 1500 that can be used by or in a camera or other device to provide one or more embodiments described herein. The systems and methods described herein can be implemented in a variety of different data processing systems and devices, including general-purpose computer systems, special purpose computer systems, or a hybrid of general purpose and special purpose computer systems. Data processing systems that can use any one of the methods described herein include a camera, a smartphone, a set top box, a computer,such as a laptop or tablet computer, embedded devices, game systems, and consumer electronic devices, etc., or other electronic devices.

[0090] Figure 15 is a block diagram of data processing system 1500 hardware according to an embodiment. Note that while Figure 15 illustrates the various components of a data processing system, it is not intended to represent any particular architecture or manner of interconnecting the components as such details are not germane to the present invention. It will also be appreciated that other types of data processing systems that have fewer components than shown or more components than shown in Figure 15 can also be used with one or more embodiments of the present invention.

[0091] As shown in Figure 15, the data processing system 1500 includes one or more buses 1509 that serve to interconnect the various components of the system. The system in Figure 15 can include a camera or be coupled to a camera. One or more processing devices 1503 are coupled to the one or more buses 1509 as is known in the art. Memory 1505 may be DRAM or non-volatile RAM or may be flash memory or other types of memory or a combination of such memory devices. This memory is coupled to the one or more buses 1509 using techniques known in the art. The data processing system can also include non-volatile memory 1507, which may be a hard disk drive or a flash memory or a magnetic optical drive or magnetic memory or an optical drive or other types of memory systems that maintain data even after power is removed from the system. The non-volatile memory 1507 and the memory 1505 are both coupled to the one or more buses 1509 using known interfaces and connection techniques. A display controller 1521 is coupled to the one or more buses 1509 in order to receive display data to be displayed on a display device which can be one of displays. The data processing system 1500 can also include one or more input / output (1 / 0) controllers 1515 which provide interfaces for one or more 1 / 0 devices, such as one or more cameras, touch screens, ambient light sensors, and other input devices including those known in the art and output devices (e.g., speakers). The input / output devices 1517 are coupled through one or more 1 / 0 controllers 1515 as is known in the art. The ambient light sensors can be integrated into the system in Figure 15.

[0092] While Figure 15 shows that the non-volatile memory 1507 and the memory 1505 are coupled to the one or more buses directly rather than through a network interface, it will be appreciated that the present invention can utilize non-volatile memory that is remote from the system, such as a network storage device which iscoupled to the data processing system through a network interface such as a modem or Ethernet interface. The buses 1509 can be connected to each other through various bridges, controllers and / or adapters as is well known in the art. In one embodiment the I / O controller 1515 includes one or more of a USB (Universal Serial Bus) adapter for controlling USB peripherals, an IEEE 1394 controller for IEEE 1394 compliant peripherals, or a Thunderbolt controller for controlling Thunderbolt peripherals. In one embodiment, one or more network device(s) 1525 can be coupled to the bus(es) 1509. The network device(s) 1525 can be wired network devices (e.g., Ethernet) or wireless network devices (e.g., Wi-Fi, Bluetooth) that receive images from a camera, etc.

[0093] Although separate embodiments are enumerated below, it will be appreciated that these embodiments can be combined or modified, in whole or in part, into various different combinations. The combinations of these embodiments can be any one of all possible combinations of the separate embodiments.

[0094] Embodiment 1 is a method comprising: receiving a plurality of video clips; encoding the plurality of video clips; computing a set of predicted transition embeddings using the target style embedding and the encoded plurality of video clips; and determining a plurality of transitions for the plurality of video clips using the set of predicted transition embeddings in a recurrent manner.

[0095] Embodiment 2 is a method of embodiment 1 wherein the plurality of video clips are ordered in time.

[0096] Embodiment 3 is a method of embodiment 1 or embodiment 2, wherein a transition embedding is a learned deep representation of an input video clip.

[0097] Embodiment 4 is a method of any of embodiment 1 to embodiment 3, wherein the computing comprises: decoding each of the plurality of video clips iteratively.

[0098] Embodiment 5 is a method of embodiment 4, wherein the decoding comprises: outputting a predicted transition embedding for sequential pair of the plurality of video clips.

[0099] Embodiment 6 is a method of any of embodiment 1 to embodiment 5, wherein the determining comprises: computing the set of transitions using a pretrained plurality of transition embeddings.

[0100] Embodiment 7 is a method of embodiment 6, wherein the pretrained plurality of transition embeddings includes a plurality of set of the pretrained transition embeddings, wherein each set of the pretrained transition embeddings is associated with a different class.

[0101] Embodiment 8 is a method of embodiment 7, wherein the class is a transition style class.

[0102] Embodiment 9 is a method of embodiment 8, wherein the class is one of direct_cut, pull_in, mix, pull_out, circle, open, windmill, cube, switch, left, pane, circle, right, black_fade, tum_page, clock_wipe, blinds, heart, squeeze, floodlight, down, kaleidoscope, memory, white_flash, memory, blur, gradient_wipe, superimpose, dissolve, star, and blanch.

[0103] Embodiment 10 is a method of any of embodiment 1 to embodiment 9, further comprising: encoding each of the plurality of video clips to output a plurality of a latent embedding representing each of the plurality of video clips; and style conditioning each of the latent embeddings to produce style conditioned latent embeddings.

[0104] Embodiment 11 is a method of embodiment 10, wherein the style conditioning is conditioning a latent embedding using a target style.

[0105] Embodiment 12 is a method of embodiment 10, wherein the style conditioning comprises: computing a distance between the target style embedding and an average of a plurality of estimated transition embeddings; minimizing the distance using activation maximization; and computing the style conditioned latent embeddings using a transformer encoder and decoder or using a large language model.

[0106] Embodiment 13 is a method of embodiment 12, wherein the decoding uses a transformer decoder to output a style adapted latent embedding.

[0107] Embodiment 14 is a method of embodiment 12, wherein the target style is computed using a large language model.

[0108] Embodiment 15 is a method of any of embodiment 1 to embodiment 14, further comprising: determining a next video clip by computing the distance between the target style embedding and predicted transition embedding; anddetermining a growing window of next video clips iteratively.

[0109] Embodiment 16 is a method of any of embodiment 1 to embodiment 15, wherein a target style embedding is a latent embedding that can be learned using one of a contrastive learning with a large language model or videos corresponding to a particular style in a zero- shot way.

[0110] Embodiment 17 is a method of any of embodiment 1 to embodiment 16, further comprising: determining an order of the plurality of video clips by computing the distance between the target style embedding and predicted transition embedding.

[0111] Embodiment 18 is an apparatus comprising a processing system and memory and configured to perform any one of the methods in embodiment 1 to embodiment 17.

[0112] Embodiment 19 is a non-transitory machine-readable storage storing executable program instructions which when executed by a machine cause the machine to perform any one of the methods of embodiment 1 to embodiment 17.

[0113] It will be apparent from this description that one or more embodiments of the present invention may be embodied, at least in part, in software. That is, the techniques may be carried out in a data processing system in response to its one or more processors executing a sequence of instructions contained in a storage medium, such as a non-transitory machine- readable storage medium (e.g., DRAM or flash memory). In various embodiments, hardwired circuitry may be used in combination with software instructions to implement the present invention. Thus, the techniques are not limited to any specific combination of hardware circuitry and software, or to any particular source for the instructions executed by the data processing system.

[0114] In the foregoing specification, specific exemplary embodiments have been described. It will be evident that various modifications may be made to those embodiments without departing from the broader spirit and scope set forth in the following claims. The specification and drawings are, accordingly, to be regarded in an illustrative sense rather than a restrictive sense.

Claims

CLAIMS1. A method for determining transitions between video clips of a plurality of video clips according to a particular target style, the method comprising: receiving the plurality of video clips; for each video clip of the plurality of video clips, video encoding the respective video clip by transforming the frames of the respective video clip into a corresponding video clip embedding, the size of the video clip embedding being independent from the number of frames in the respective video clip; concatenating the video clip embeddings to obtain a singular video embedding; transformer encoding the singular video embedding to obtain, for each video clip represented in the singular video embedding, a latent embedding; computing a set of predicted transition embeddings using a target style embedding and the latent embeddings encoded from the plurality of video clips, each predicted transition embedding of the set of predicted transition embeddings corresponding to a video clip of the plurality of video clips, the target style embedding determined from one or more style video clips corresponding to the particular target style, wherein the computing the set of predicted transition embeddings comprises, by sequentially processing each latent embedding: transformer decoding the respective latent embedding to obtain a corresponding estimated transition embedding; style conditioning the respective latent embedding using the target style embedding to produce a corresponding style conditioned latent embedding, wherein the style conditioning employs activation maximization; and transformer decoding the style conditioned latent embedding to update the estimated transition embedding, thereby obtaining a predicted transition embedding; and determining a plurality of transitions for the plurality of video clips using the set of predicted transition embeddings in a recurrent manner, wherein the determining the plurality of transitions comprises: computing the set of transitions using a plurality of pretrained transition embeddings.

2. The method of claim 1, wherein the plurality of video clips is a sequence of video clips ordered in time.

3. The method of claim 1 or claim 2, wherein the video encoding comprises, for each video clip of the plurality of video clips: initially encoding the respective video clip to obtain a corresponding initial video clip embedding, the size of the initial video clip embedding being independent from the length of the respective video clip; and positionally encoding the respective initial video clip embedding to obtain the corresponding video clip embedding that encodes information about the order of the respective video clip in the plurality of video clips.

4. The method of claim 3, wherein the video encoding comprises, for each video clip of the plurality of video clips: between initially encoding and positionally encoding, reducing the number of feature dimensions of the initial video clip embedding by applying a pooling layer to the initial video clip embedding.

5. The method of any of claims 1-4, wherein the style conditioning employing activation maximization comprises, while processing the respective latent embedding: computing a distance between the target style embedding and an average of the estimated transition embeddings determined for the preceding latent embeddings up to the respective embedding; and adjusting the respective latent embedding based on a scaled gradient of the computed distance to generate the style conditioned latent embedding; wherein employing the activation maximization causes minimizing the computed distance.

6. The method of claim 5, wherein adjusting the respective latent embedding based on a scaled gradient of the computed distance comprises subtracting a gradient of the computed distance, scaled by a learning rate, to generate the style conditioned latent embedding.

7. The method of any of claims 1-6, wherein the plurality of pretrained transition embeddings includes a plurality of sets of pretrained transition embeddings, wherein eachset of pretrained transition embeddings of the plurality of sets of pretrained transition embeddings is associated with a different transition style class.

8. The method of claim 7 , wherein the transition style class is one of direct_cut, pull_in, mix, pull_out, circle, open, windmill, cube, switch, left, pane, circle, right, black_fade, turn_page, clock_wipe, blinds, heart, squeeze, floodlight, down, kaleidoscope, memory, white_flash, memory, blur, gradient_wipe, superimpo\se, dissolve, star, and blanch.

9. The method of claim 7 or claim 8, further comprising determining the plurality of sets of pretrained transition embeddings, the determining of a set of pretrained transition embeddings associated with a particular transition style class comprising: receiving a plurality of sets of video transitions, wherein each set of video transitions of the plurality of sets of video transitions is associated with the particular transition style class; and for each set of video transitions of the plurality of sets of video transitions: transforming, by a neural network, the respective set of video transitions into a singular embedding; inputting the singular embedding into a linear layer; normalizing the output of the linear layer to obtain a pretrained transition embedding; and classifying the pretrained transition embedding to obtain the transition style class associated with the corresponding set of video transitions; combining the determined pretrained transition embeddings over the plurality of sets of video transitions to obtain the set of pretrained transition embeddings.

10. The method of any of claims 7-9, further comprising: determining, by using contrastive learning, the target style embedding from one or more style video clips, each of the one or more style video clips including transitions representing the particular target style .

11. The method of any of claims 7-9, further comprising: determining, by using contrastive learning and a large language model, the target style embedding from (i) one or more style video clips, each of the one or more style videoclips including transitions representing the particular target style, and (ii) textual information indicative of the transition style class.

12. The method of any of claims 7-9, further comprising determining the target style embedding in a zero-shot way by: receiving one or more style video clips, each of the one or more style video clips including transitions representing the same particular target style; for each of the one or more style video clips: video encoding the respective style video clip by transforming the frames of the respective style video clip into a corresponding singular style video clip embedding, the size of the style video clip embedding being independent from the number of frames in the respective style video clip; transformer encoding the respective style video clip embedding to obtain a corresponding latent style embedding; and transformer decoding the latent style embedding to obtain a corresponding style embedding; and determining the target style embedding by calculating a mean of the style embeddings obtained for each of the one or more style video clips.

13. The method of any of claims 1-12, further comprising: determining, according to the particular target style, a next video clip from a set of unordered video clips to be appended to a sequence of video clips ordered in time by computing the distance between the target style embedding and the predicted transition embedding; and determining a growing window of next video clips iteratively.

14. An apparatus comprising a processor configured to perform the method of any of claims 1-13.

15. A computer readable medium comprising instructions which when executed by a processor cause the processor to perform the method of any of claims 1-13.

Citation Information

Patent Citations

  • Video remixing system

    US20150194185A1

  • Systems and methods for automating video editing

    US20210272599A1

  • Video processing methods and apparatuses, electronic devices, storage mediums and computer programs

    US20220084313A1

  • Automated cinematographic editing tool

    WO2009055929A1

Cited By

  • Real-time interactive video generation method and system

    US12608757B2

  • Real-time interactive video generation method and system

    US20260017754A1