Apparatus and method for compiling a video clip on the basis of multiple video sequences and text segments
The device automates the production of high-quality, semantically and stylistically consistent video clips by using a transformer module to generate prediction vectors from text and video features, addressing inefficiencies in current news production systems.
Patent Information
- Application Number
- PCT/EP2025/053593
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-02-11
- Publication Date
- 2025-08-21
AI Technical Summary
Current video editing systems for news production are inefficient and time-consuming, requiring extensive manual intervention, and fail to produce semantically and stylistically consistent video clips that meet professional editing standards, especially when dealing with diverse source materials.
A device and method utilizing a transformer module to generate prediction vectors based on text and video features, selecting suitable video sequences from a feature database, and creating video clips that align with editorial text and stylistic standards, incorporating stylistic elements like camera movement and angle, through a cross-modal encoder and feature decoder.
Enables automated, high-quality video clip production that is semantically coherent with editorial text and stylistically consistent with professional editing standards, reducing manual intervention and improving editing efficiency in news production.
Smart Images

Figure EP2025053593_21082025_PF_FP_ABST
Abstract
Description
[0001] Apparatus and method for assembling a video clip based on multiple video sequences and text segments
[0002] The present invention relates to a device and a method for composing a video clip based on multiple video sequences and text segments of a text. The device comprises a text module for generating a text feature vector with semantic text features of a text segment. The device further has a transformer module for generating a prediction vector, a selection module for selecting a target vector of a suitable target video sequence from a feature database, an assignment module for creating a sequence list with target vectors for selected video sequences, and a video clip module for creating video clips from video scenes of a video database based on the sequence list.
[0003] Video clips consisting of individual scenes or sequences are used in many areas of daily life, for example in social media, as instructional guides or in the form of news reports on television.
[0004] The distribution of news in the form of video clips, so-called video news or news clips, has evolved significantly due to the growth of digital media and widespread internet access. This has led to a shift toward video-based news on platforms such as YouTube and social media. These engaging and accessible formats are preferred over traditional text-based news, especially by younger people. There is therefore a growing demand for the rapid creation of video content to capture audience interest in the online news space.
[0005] News broadcasters must therefore produce compelling video clips faster than ever to ensure their successful positioning in the market. This is due in part to the growing number of news sources and the changing media consumption patterns of their target audience. This development has intensified the need to produce news clips quickly—a requirement that conflicts with traditionally manual and time-consuming video editing processes.
[0006] Video news production involves complex, time-consuming steps that make it difficult to keep up with the pace of breaking news. The process begins with writing a news article, which is then converted into a voice-over script for the video. The production of the news clip involves selecting and sequencing shots that align with the sections of the article, observing the principles of rhythm (i.e., the length of the scenes affects the tone of the news clip), continuity (i.e., the content of successive scenes must be consistent), and editing (i.e., the style of successive scenes follows rules) to ensure high-quality output.
[0007] Apart from some advances in video production automation, current systems do not meet the efficiency and quality standards required for professional news broadcasts. Extensive manual intervention is still required in the creation of video clips, especially when a compilation is based on diverse source materials and the generated clips are to be semantically consistent with the editorial text and stylistically consistent with professional editing standards. Therefore, there is a great need to improve and streamline the video clip editing and production process and propose a system that can maintain the quality of the video clips while reducing manual intervention.
[0008] The present object is achieved by a device having the features of claim 1 and by a method having the features of claim 11.
[0009] In one aspect, the present invention relates to a device comprising a text module for generating a text feature vector with (semantic) text features of a text segment, a transformer module for generating a prediction vector for a video sequence based on the video vector of the previous video sequence, a selection module for selecting a target vector of a suitable target video sequence (also called target) from a feature bank (also called source shots) based on the prediction vector and optionally for generating or selecting a target vector corresponding to the target video sequence, an assignment module for creating a sequence list with target vectors for the selection of video sequences based on the target vectors and a video clip module for creating a video clip from video sequences of a video database based on the sequence list.The Transformer module includes a cross-modal encoder and a feature decoder and is trained to create the prediction vector based on the video vector and the text feature vector. The video vectors include video classification features that include or are stylistic devices or features.
[0010] In a further aspect, the present invention relates to a training device for training a transformer module and for creating video vectors for video sequences based on text segments of a text or news article. The training device comprises a CLIP encoder for creating text feature vectors with text features of a text segment, a CLIP encoder for creating content vectors with video content features for a video sequence, a VGG encoder for creating classification vectors with video classification features for a video sequence, and the transformer module with a cross-modal encoder and a feature decoder. The cross-modal encoder is designed to combine a text feature vector and a content vector as well as a classification vector for a video sequence into a feature vector and feed it to the feature decoder, wherein the feature vector preferably has twice the length of the content vector.The feature decoder is configured to create a video vector from the feature vector. The video vector comprises the features of a video sequence present at the input of the training device and is preferably half the length of the feature vector. The training device outputs the video vector at an output and makes it available for further processing. The feature vector can preferably include additional stylistic features, such as the camera movement or the camera angle. This allows additional stylistic parameters to be incorporated into the training.
[0011] Further aspects of the invention relate to a corresponding method and a computer program product with program code for carrying out the steps of the method when the program code is executed on a computer, as well as a storage medium on which a computer program is stored which, when executed on a computer, effects execution of the method described herein.
[0012] Preferred embodiments of the invention are described in the dependent claims. It is understood that the features mentioned above and those to be explained below can be used not only in the respectively specified combination, but also in other combinations or alone, without departing from the scope of the present invention. In particular, the method and the computer program product can be designed according to the embodiments described for the device in the dependent claims. According to the invention, a prediction vector is generated based on a text feature vector that comprises semantic text features of a text segment, and based on a video vector with features of a video sequence. The video vector of the video sequence is based on a previous video sequence and comprises its features.The prediction vector is the feature vector generated by the AI unit, the so-called transformer module, for the video sequence following the previous one and includes its features. The prediction vector therefore describes the desired subsequent video sequence.
[0013] Based on this prediction vector with features of a video sequence to be selected, which is to follow the previous video sequence and should match the content of the text features of the text segment relevant to the planned video sequence, the search for a suitable sequence is carried out. Based on the prediction vector, a target vector with features of a suitable target video sequence is selected, whereby the target vector is contained in a feature bank. The feature bank contains the video vectors of all available video sequences, i.e. the feature vectors that represent and correspond to the video sequences. From this "pool" of available video sequences, which are represented by their feature vectors, the so-called video vector or target vector, a comparison is made with the prediction vector. The goal is to select the most suitable target vector.The target vector that correlates best with the prediction vector is selected. This target vector is then written to a sequence list, preferably appended to an existing or empty sequence list. In this way, after several passes, the sequence list comprises a sequence of target vectors that includes features of selected and classified as suitable video sequences from the "pool" of video sequences, which are represented by the vectors of the feature database. As soon as all text feature vectors of the text segments of the entire news text have been processed, the sequence list contains all associated video sequences of the video clip as features in the form of target vectors that represent the individual video sequences of the video clip. In a video clip module, the video sequences are selected from a video database based on the target vectors in the feature list and combined to form a video clip.The resulting video clip then includes the individual video sequences that best represent and reproduce the stored text, for example a news text.
[0014] The individual recordings of camera footage or archive material from a corresponding database, which are available in the form of video segments, and the corresponding article segments of a text are the basic units for compiling news or video clips. A high-quality news clip with multiple recordings or video sequences, which is tailored to the underlying texts, especially news articles, exhibits the characteristics considered desirable. Firstly, it provides a visual representation of the semantics of the news text. Secondly, it implements appealing transitions between the individual video sequences that comply with the rules for video editing of news clips.
[0015] According to the invention, prediction vectors based on a video vector with features of a video sequence and a text feature vector with features of a text segment are created in the transformer module, which comprises a cross-modal encoder and a feature decoder. Advantageously, the video vector with features of the video sequence and the text feature vector can be combined in the cross-modal encoder to create a feature vector of twice the length. The feature decoder extracts the appropriate features from this to create a feature vector, the prediction vector, whose size corresponds to the size of the video vector or the text feature vector. Preferably, the video vector, the text feature vector, and the prediction vector are each of the same length.
[0016] The present invention has the advantage that a video clip can be produced reliably and automatically with high quality, even when the video sequences are based on different source materials. The invention enables the generated video clips to be semantically coherent with the editorial text on which the video clip is based, and furthermore, to stylistically conform to and meet professional editing standards. The invention therefore provides an efficient automated system for editing news clips or video clips with news, supporting journalists in their tasks and significantly reducing their workload.
[0017] Because previously known systems only focus on media areas outside of news production, such as feature film production, they cannot be directly applied to news production or documentaries, as significantly different aspects play a decisive role in these areas than in films. For journalists and video editors, there are certain stylistic means for producing high-quality news that cannot be adequately represented with existing systems. The present invention therefore focuses on the specific stylistic means of news production in order to automate not only content-consistent but also high-quality news production. Within the scope of the invention, it was also recognized that these stylistic means significantly increase and improve the quality of the news clips or videos produced.Their consideration was recognized as a decisive factor for the quality and acceptance of news videos and for increasing the viewer's information intake.
[0018] The present invention thus succeeds in using an embedding of imagery and video editing steps learned from an unlabeled video dataset containing multiple video sequences. The fully automated system reduces editing time and significantly improves the quality of the generated news clips. The video segments used are aligned with the content of the news article and the respective text segments of the news text, ensuring smooth transitions between recordings by adhering to established guidelines for video editing. This is achieved by processing the video vector of the previous video sequence containing the features, as well as the text features of a text segment determined for the next video sequence, in the form of a text feature vector when creating the prediction vector for the next video sequence.
[0019] The present invention therefore uses stylistic characteristics and features of a video sequence as video features to perform video-text matching. This clearly distinguishes the invention from known systems that only consider content-related features of a video or a sequence of images. While it is also known to use an "aesthetic score" to estimate the next video sequence, this only considers aspects such as sharpness or focus, which are technical or qualitative-technical characteristics of a video. However, these characteristics play a negligible role in professional camera footage and are almost always met. They simply do not include stylistic elements and aspects.
[0020] In a preferred embodiment of the invention, the video vector and the prediction vector each comprise text features, video content features, and video classification features. Thus, in addition to content and semantic features, the vectors also have classification features. These are suitable for making transitions between individual video sequences smoother and more pleasant for the viewer. In this way, it is possible to establish the principles of video editing technology automatically. The video classification features are preferably generated using a VGG encoder. Such an encoder can, for example, be integrated into the transformer module or the transformer unit. The VGG encoder is based on the concept of transfer learning, i.e., the ability to use existing knowledge developed to solve specific problems to solve new ones.The VGG encoder comprises a neural network of convolutional cells developed at the University of Oxford. The model achieved an accuracy of over 92% in a competition, the highest ever achieved. The neural network's convolution matrices use smaller convolution kernels (3x3) and are based on a large database of more than 14 million labeled images divided into over 100 classes.
[0021] For example, using a fine-tuned, specially trained "VILA encoder" (e.g., from Google), features of the scene composition can be generated and processed. Likewise, motion vectors of the scene can be generated preferentially using a motion encoder. The various video classification features can be combined preferentially into a single vector.
[0022] In a preferred embodiment, the classification features of the video vectors and the prediction vector as well as the target vector are features of the video sequence, which can include, for example, size features of the recording. The features are members of a list that includes shot sizes of the video sequence. This can, for example, be information about whether a wide view, a normal or medium view, a detailed view, or a portrait view is present. Further elements of the list of classification features include the camera angle used in the video sequence, statements about the movement of the camera in the video sequence, the direction of movement of the camera, movements of people shown in the video sequence, the direction of movement of people in the sequence, zoom settings, lighting settings, or similar settings used in the video sequence.Preferably, the scene composition with regard to the subject depicted is also taken into account, i.e. is the person in the center of the image and clearly recognizable. Image quality such as image blurriness, camera shake, etc. can also serve as a classification feature. Although these are important, they are not as important for the device according to the invention and for the creation of news videos as in the systems known from the prior art. In the field of news production, the camera images are recorded by professional cameramen; therefore, image quality plays a minor role. It can preferably be assumed. To date, the prior art has only considered image quality and not the other stylistic means that are essential for high-quality news production.
[0023] A preferred embodiment provides that the video content features of the video vectors, prediction vectors, and / or target vectors were generated using a CLIP encoder. A CLIP encoder makes it possible to extract content from images or video sequences, similar to text-based transformers. A CLIP model is a Contrastive Language Image pre-training model that can predict the most relevant captions or text for an image without the model being optimized for a specific task. It is based on approximately 400 million image-text pairs, with which it was trained to predict the captions belonging to a specific image. The CLIP encoder can preferably be integrated into the transformer module. However, it can also be implemented as a standalone module.
[0024] In a preferred embodiment, the text module also includes a CLIP encoder. The text feature vector is generated using this CLIP encoder. It includes the text features and content of the text segment to be assigned to the individual video sequence.
[0025] In a preferred embodiment, the transformer module is configured to comprise a plurality of submodules. Each of the submodules is configured to analyze and evaluate a predefined video classification feature of a video vector. This allows the video classification features to be taken into account with higher quality when creating the prediction vector.
[0026] Each of the submodules in the transformer module preferably comprises a neural network. This neural network is preferably a feed-forward network (FFN). Each individual feed-forward network is specifically trained for a modality or stylistic feature. This makes it possible to consider multiple stylistic features simultaneously, so that when creating a video vector and / or a prediction vector, multiple stylistic features of the video sequences are taken into account. This improves the quality and interplay between text and image or image sequence.
[0027] In a preferred embodiment, each individual submodule can generate a subfeature for classification, which is taken into account when creating the prediction vector. The prediction vector then comprises the sum of all subfeatures of the submodules.
[0028] In a preferred embodiment, the submodule comprises a neural network or a feed-forward network. Particularly preferably, a plurality of stylistic features or characteristics in the form of video classification features are processed in the feed-forward network.
[0029] Likewise preferred is an embodiment in which each feed-forward network is specialized for a given stylistic device or stylistic feature of a video classification feature. Thus, specific processing of the individual stylistic features takes place. Preferably, the individual results of the feed-forward networks of the individual sub-modules are combined and form the basis for the prediction vector. The prediction vector forms the basis for the video sequence, with the target vector or target video vector of a suitable target video sequence being selected from the feature database based on this prediction vector. In a preferred embodiment, the transformer module is designed to create the prediction vector based on multiple video vectors and multiple text feature vectors. In this case, multiple video classification features are taken into account in the individual video vectors.
[0030] In a preferred embodiment, the prediction vector is generated based on the video vectors of a majority of the available video sequences. Particularly preferably, the generation takes place based on the video vectors of all available video sequences. Simultaneously or alternatively, the generation of the prediction vector can be carried out based on all available text feature vectors. It is also conceivable for the generation of the prediction vector to be based on a majority of the available text feature vectors. Preferably, the prediction vector is created based on video vectors and text feature vectors, particularly preferably based on all available video vectors and text feature vectors. For example, an input vector can be used for this purpose that includes all segments of the available news text, as well as all scenes or video sequences of the available camera material.The transformer module thus receives all available information for its estimation and for generating the prediction vector, which in turn allows it to select the most suitable target vector for a target video sequence. The transformer module therefore considers all text segments and thus incorporates all text information into its estimation. This has the advantage of improving the otherwise typically iteratively generated contribution, which only considers the previous video sequence or scene, because all available text feature vectors and video vectors can now be incorporated into the process of creating the prediction vector and thus into the selection of the target vector. This significantly accelerates the entire process and simultaneously improves the quality of the generated news videos.
[0031] The entire target vector is generated taking into account all text and all previous scenes. This provides the Transformer module with all available information for its estimation. It considers all text segments and thus incorporates all text information into its estimation.
[0032] In this way, the scene sequence for the report as a whole can be estimated, including the interview content. Since these (additionally considered) text passages can contain important information, the quality of the estimation improves. Interview passages, such as those typical for news broadcasts or videos, can also be considered, which is why special interview tokens (using modal type embedding) can be assigned in the feature vector. The feature vector forms the basis for the video vector, on the basis of which the prediction vector is created.
[0033] In a preferred embodiment, the sequence list for selecting video sequences is created by appending a target vector to the sequence list. The sequence list is thus expanded with additional target vectors until all text segments of the text that is to be underlaid for the video clip have been processed. The sequence list is created in the assignment module.
[0034] In a special embodiment, the text element of a text comprises a single sentence from the given (entire) text. Of course, it is also possible for the text element to consist of several sentences, preferably up to five. This is useful if the sentences are related in content.
[0035] In a preferred embodiment of the device, the selection module is designed to compare the video vectors stored in the feature database with the prediction vector. Preferably, all video vectors in the feature bank are compared with a prediction vector from the selection module. This is particularly preferably done individually. The purpose of the comparison is to find a video vector from the feature bank that corresponds as closely as possible to the prediction vector and to incorporate this video vector into the sequence list. Preferably, the quality of the selection can be assessed when assessing the fit or matching of the prediction vector with one of the video vectors in the feature bank. The vector selected from the video vectors is the target vector that is fed to the sequence list. A quality function can preferably be used to determine the quality for the appropriate selection.For example, it is possible to assess quality based on cosine similarities.
[0036] In a preferred embodiment, the sequence list created in the assignment module can include interview vectors that correspond to video sequences that include or are an interview sequence. This makes it possible to include interview sequences in the video clip to be created. The interview vectors can be feature vectors that also include the transcript of the spoken dialogue in the interview sequence.
[0037] According to the device, the method according to the invention can also provide in a preferred embodiment that the generation of the prediction vector for the video sequence is carried out based on several video vectors and several text feature vectors.
[0038] According to one embodiment of the method, the prediction vector is preferably generated based on all available video vectors or on the video vectors of all available video sequences and / or based on all available text feature vectors. In general, the amount of available information increases with the number of processed video vectors and text segments, thus further improving the quality of subsequent clips. The same applies to the device.
[0039] In a preferred embodiment of the training device, the feature vector comprises stylistic means or stylistic features. These stylistic means are the same as those used in generating the prediction vector and selecting a target video sequence. A preferred embodiment of the training device provides that it is configured such that multiple text feature vectors and / or multiple content vectors and / or multiple classification vectors are used to form the feature vector, from which a video vector is created that has the features of an unknown video sequence.
[0040] The transformer module preferably comprises a neural network, e.g., a deep neural network or a feed-forward network (FFN). The transformer module can preferably comprise multiple submodules, each of which comprises a neural network or an FFN.
[0041] Preferably, the training device can be configured to train the submodules of the transformer module and, in doing so, to process the stylistic means and stylistic features for which the submodule is to be specifically designed and trained. Preferably, a feature vector (sub-feature vector) can be formed for each submodule.
[0042] In a preferred embodiment, the training device may additionally or optionally comprise a VILA encoder (Visual Language Encoder) for creating a vector or feature vector of the scene composition and preferably a score for the image quality in order to describe and characterize the video sequence.
[0043] A preferred embodiment of the training device may include a motion encoder to generate a feature vector of the camera motion for a video sequence. This "motion feature vector" can be incorporated into the feature vector for the video vector, for example, through summation or another vector calculation rule.
[0044] In a preferred embodiment of the training device, it can have multiple submodules for the stylistic features to be considered, wherein each submodule can preferably be specialized for a particular stylistic feature. In other words, each submodule or FFN is trained specifically for one modality or one particular stylistic feature. In this way, more stylistic features than just the setting size can be considered.
[0045] Instead of processing only two modalities, namely content and shot size, for example, by concatenating CLIP features and VGG features and then mapping them linearly to reduce size, the use of this large amount of information requires special consideration to avoid information loss due to excessive reduction. The entire information summarized in a feature vector is divided into individual sub-feature vectors, which can be processed in the multiple sub-modules with little or no practical loss of information. After processing in the sub-modules, the processed sub-feature vectors can be recombined into a single feature vector.
[0046] The vectors processed in the device, the training device, and the method are feature vectors and comprise features of the text segment or a video sequence. The text vector is a text feature vector with features of a text segment. The prediction vector is a feature vector generated by the device, in particular by the transformer module, with content features as well as classification features for the next video sequence to be selected. The video vector or video feature vector comprises the features of a video sequence, which can include both content features and classification features. The target vector is the feature vector found by comparing the prediction vector with video vectors from a feature bank. This is also referred to as the target feature vector or matching feature vector. Another term is target vector.
[0047] The invention is described and explained in more detail below using selected exemplary embodiments in conjunction with the accompanying drawings. Figure 1 shows the device according to the invention for compiling a video clip;
[0048] Figure 2 shows the process and use of the device when creating a video clip;
[0049] Figure 3 shows a transformer module for generating a video vector;
[0050] Figure 4 shows a training device according to the invention with a transformer module; and
[0051] Figure 5 is a schematic representation of the process steps of the method for compiling a video clip.
[0052] Figure 1 shows a device 10 for composing a video clip based on multiple video sequences and text segments of a text. The device 10 comprises a transformer module 20 for generating a prediction vector 22 for a video sequence based on a video vector of a previous video sequence. A text segment 12 is processed in a text module 30, in which a text feature vector 32 is generated from the text features of the text segment 12. This text feature vector 32 represents the semantic content of the text segment 12.
[0053] A selection module 40 processes the prediction vector 22, based on which the selection module 40 selects a target vector 42 of a suitable target video sequence from a feature bank 50. Feature vectors from individual video sequences are stored in the feature bank 50. From this feature bank 50, which can be a database or a corresponding table, the feature vectors that best match the prediction vector 22 generated in the transformer module 20 are selected by means of selection modules 40. The selection can be based on a quality criterion, for example, based on cosine similarities, as described by way of example in Formula 1.
[0054] Formula 1 :
[0055] Here || x ||2 is the L2-norm of a vector x. v p is the prediction vector 22 and v g the target vector 42. v T means the transpose of the vector v.
[0056] In the selection module 40, all target vectors 42 stored in the feature bank 50 are compared with the prediction vector 22. The target vector 42 that has the greatest cosine similarity to the prediction vector 22 is selected and stored in a sequence list 62 in an assignment module 60. The sequence list 62 includes all target vectors 42 of video sequences that form the basis for the video clip to be created.
[0057] In a video clip module 70, a video clip 72 is generated from video sequences 74 based on the sequence list 62. The individual video sequences 74 are the sequences described and characterized by the respective target vectors 42. Consequently, a video clip 72 is created in which several selected video sequences 74 are strung together. The individual video sequences 74 typically have a length of 30 seconds to 5 minutes, preferably 10 seconds to 3 minutes, and particularly preferably 30 seconds to 90 seconds.
[0058] Since the target vectors 42 of the sequence list 62 include the text feature vectors 32 for the respective individual text segments 12, a video clip 72 is produced in the form of a news clip, comprising several video sequences 74 as well as the corresponding text passages for the video sequence 74. In this way, a speaker can speak the corresponding text for the video clip 72. Either the underlying text segments 12 can be used, or free text oriented solely to the text segments 12 can be spoken. It is possible for the audio signal generated for the text segments 12 to be generated with the aid of artificial intelligence or an AI unit and assigned to the video sequences 74. The determined target vector 42 is fed not only to the assignment module 60, but also to the transformer module 20, where it is processed with the text feature vector 32 of the next text segment 12.In this way, the target vector 42 is taken into account as the video vector 44 of the previous video sequence 74 when selecting the next video sequence 74, i.e., the target vector 42. The associated processing is thus continued until all video sequences 74 and text segments 12 have been processed.
[0059] Figure 2 shows the cascal process for generating a video clip 72 and assembling the video clip 72 from multiple video sequences 74 based on the selected target vectors 42 of a sequence list 62. Instead of the representation chosen in Figure 1, in which the processing loop is run through multiple times, the individual processing steps are shown sequentially here, with the device 10 being shown multiple times in different stages 16, 17, 18, 19.
[0060] First, processing is started with a target vector 42, which represents a best-shot vector 14 and is specified. The best-shot vector 14 represents a predetermined video sequence that serves as the input sequence for the video clip 72. The best-shot vector 14 can, for example, be selected and specified manually or by a separate Kl unit.
[0061] In the transformer module 20, the best-shot vector 14 and the text feature vector 32 generated in the text module 30 are first processed based on the first text segment 12. A prediction vector 22 is generated here, which is processed with the target vectors 42 of the feature bank 50 in the selection module 40. The target vector 42 that best matches the prediction vector 22, i.e., has the greatest agreement or similarity, is selected.
[0062] The selected target vector 42 is inserted into the sequence list 62 present in the assignment module 60, with this target vector 42 being the first entry in the sequence list 62. The target vector from the first stage 16 is fed to the second stage 17 and processed there with the new text segment 12. The device 10 is thus used again in the second stage 17 to select a suitable target vector 42. This is then appended to the sequence list 62 in the assignment module 60, so that the sequence list 62 grows with each stage and each additional video sequence or each additional text segment 12.
[0063] In a third stage 18, the target vector 42 generated in the second stage 17 is used to generate a new target vector 42. This principle is continued up to an nth stage 19 until all text segments 12 have been processed. The text segments 12 originate from the text available for a video clip 72, in particular a news text.
[0064] At the end of the run through all stages 16 to 19, a video clip is created, which is generated in the video clip module 70 based on the sequence list 62 generated by the assignment module 60.
[0065] Figure 2 shows that the feature bank 50 can include an optional interview bank 52. The underlying news text, which is divided into several text segments 12, can also include interview segments 54 that match corresponding interview sequences 55 in the interview bank 52. In an equally optional interview module 56, the interview sequences 55 of the interview bank 52 are analyzed, for example, using a speech-to-text converter (speech to text), and then compared with the interview segments 54 of an interview segment database. The interview sequences 55 with the greatest similarity and the highest match to the examined interview segments 54 are selected and added to the sequence list 62 in the assignment module 60 in the form of an interview feature vector 58. In the video clip module 70, the corresponding interview sequence 55 of the interview bank 52 is selected and inserted into the video clip 72.This is optional and represents a preferred extension of the video clip 72. The effective compilation of a video clip 72 or news clip with multiple video segments is therefore based on pre-trained representations of the individual video sequences by feature vectors, so-called video vectors. Based on the coded text segments 12 in the form of text feature vectors 32 and the coded source material, i.e., the video vectors or target vectors 42 for individual video sequences 74, the workflow of a so-called inference pipeline described in Figure 2 results. In a preferred embodiment, previously marked interview sections in the news article and the corresponding interview recordings or interview sequences 55 can be filtered out, which were classified, for example, by a pre-trained interview classifier.Using an index-based text search, the individual interview-in points and interview-out points are inserted into the sequence list 62 with the target vectors 42. Initially, a best-shot vector 14 is generated that best represents the best-shot initial scene or video sequence present in the entire set of video clips 72. This selection is made according to the principle "start with the best scene," i.e., select the best video sequence as the starting point.
[0066] The following video sequences are selected using the current text segments 12 and the video vectors of the previous scene or the target vectors 42. For this purpose, a prediction vector 22 is generated for the next video sequence. The generated prediction vector 22 is then compared with all existing video vectors of the feature bank 50, each of which represents the target vectors 42. The target vector 42 that is most similar to the prediction vector is selected. The top candidates, i.e., the target vectors 42 that are most similar to the prediction vector, are inserted into the sequence list 62, based on which the video sequences are selected and combined to form the video clip 72.
[0067] Figure 3 shows the transformer module 20 in detail. The transformer module 20 includes a cross-modal encoder 24, which processes or combines a target vector 42 or a video vector 44 and a text feature vector 32 to form a feature vector 26. This includes both the text feature vector 32 and the video vector 44. The feature vector 26 is twice the length of the text feature vector 32 or the video vector 44. In a feature decoder 28 of the transformer module 20, the feature vector 26 is further processed, and a prediction vector 22 is generated. This includes the features for the next video sequence to be included in the video clip to be generated. The prediction vector 22 has the same length as the target vector 42 or the video vector 44. It is therefore preferably half the size of the feature vector 26.
[0068] Figure 4 shows the training device 80 according to the invention with the transformer module 20. The training device 80 is used to generate video vectors 44 from a plurality of video sequences 74 and a plurality of text segments 12, which are suitable for describing the video sequences 74 in the best possible way and for allowing a selection of video sequences 74 based on the features contained in the video vectors 44.
[0069] The text segments 12 are processed in the text module 30. They are parts of an overall text, for example, a news text or news article. Preferably, each text segment 12 comprises a main sentence of the news article. The news text is typically divided into individual text segments 12 in order to comply with the general video editing rule "one recording per main sentence." The text module 30 preferably includes a CLIP encoder 82, which is used to convert the content part of the text into a feature vector, here the text feature vector 32.
[0070] The video sequences 74 are categorized using a (different) CLIP encoder 82, whereby the content features of the video are extracted. Based on this, a content vector 84 with video content features is generated for the respective video sequence 74. Furthermore, the video sequences 74 are processed using a VGG encoder 86 to generate a classification vector 88. The classification vector 88 includes video classification features, which can include, for example, shot sizes of the video sequence 74, corresponding views included in the video sequence 74, camera angles, movements within the video or movements of the camera, as well as directions of movement of persons or objects within the video or video sequence 74 and / or the camera, zoom settings, or similar settings. The classification vector 88 thus represents the stylistic part of the video sequence 74.
[0071] Content vector 84 and classification vector 88 are combined to form video vector 44, which includes both the semantic essence of a video sequence 74 and its stylistic features.
[0072] Feature vector 26 and text feature vector 32 are further processed in the cross-modal encoder 24 to form feature vector 26. Further processing of feature vector 26 occurs in the feature decoder 28 to generate a prediction vector 22.
[0073] Thus, using text-image pairs or text-video sequence pairs, the transformer module 20 or the training device 80 can be trained to precisely predict the recording features that correlate with specific text segments 12 of news articles. This is done based on the available video segments. The use of the cross-modal encoder 24 and the feature decoder 28 facilitates the matching process.
[0074] The video sequences 74 used both for the training device 80 and for the device according to the invention according to Figures 1 to 3 can consist of current news clips based on a news article as well as archive material or archive video sequences. Each video sequence 74 preferably has a corresponding news article intended for dubbing. In addition, a list of manual editorial cuts often exists. The video vector 44 or prediction vector 22 based on the feature vector 26, which is available at the output of the training device 80, is fed to an evaluation module 90 to assess the quality of the generated prediction vector 22 or video vector 44. For the generation of video clips based on video segments and the text-based compilation of the video clips, finding video segments that are content-relevant and of appropriate size is important.To this end, the transformer module 20 of the training device 80 is trained to fine-tune the feature vector 26 using the text features from the text segment 12 and the features of the previous video sequence 74. The training pairs thus comprise a feature vector for the previous recording, a feature vector of the target recording, and a text of the text segment. The CLIP encoders 82 and VGG encoders 86 used generate content features and size features for the video, while the text features are also derived using a CLIP encoder 82. The encoded sequences generated here are processed by a stack of transformer coding layers, each layer having a consistent architecture.
[0075] The video vectors 44 generated in the transformer module 20 are then evaluated using a quality criterion in the evaluation module 90, wherein the cosine similarities between the generated video vector 44 and video vectors 44 of a database corresponding to the text segment are preferably assessed in a similarity module 92. Here, according to formula 2, each video vector 44 of a video sequence 74 is compared with other video sequences 74 and their feature vectors as negative samples. Contrastive learning is performed according to formula 2, wherein v p reflects the prediction vector 22 and represents Vj video vectors 44 as a comparison sample, T is a temperature parameter.
[0076] Formula 2: In an entropy module 94 of the evaluation module 90, the cross-entropy losses are used as a training target for quality assessment. Based on the multimodal transformer module 20 and the generated prediction vectors 22, the entropy module 94 uses three 1x3 CONV layers, followed by a ReLU activation. This entropy module 94 is trained to classify the stylistic setting size of the scenes.
[0077] Finally, sigmoidal activations are added to improve the prediction vectors 22. The cross-entropy losses can be
[0078] L s = — ( * log v p + (1 — v) * log(l — Vp)) Formula 3 with v p= prediction vector 22; v = target vector 42 or video vector 44. For this purpose, the combination of the individual cross-entropy losses across all video sequences 74 in the training set is set as the overall training goal. These entropy losses can be calculated by Formula 4, where L = total loss, L s = binary cross entropy loss (from formula 3) of the prediction vector 22 and L c = contrastive learning (from Formula 2). α and β are hyperparameters and control the robustness of the overall loss; N is the number of individual training sets in the overall training set.
[0079] The training set comprises the video sequences 74 and the text segments 12, which are present at the input of the training device 80 and the features generated therefrom at the input of the transformer module 20.
[0080] Figure 5 shows the basic flow of the method for compiling a video clip based on multiple video sequences and text segments of a text, in particular a news text. In a first step S10, a text feature vector with semantic text features of the text segment is generated. In a further step S12, a prediction vector for a video sequence is generated based on a video vector of a previous video sequence and based on the text feature vector of the current text segment.
[0081] Step S14 involves selecting a target vector of a target video sequence from a feature bank based on the prediction vector. In this step, the current text segment is linked to the previous video sequence to generate a feature vector for the next video sequence.
[0082] Based on this prediction vector, the target vector for the next, i.e. the subsequent video sequence, also called target video sequence, is generated in step S14.
[0083] Step S16 involves creating a sequence list with target vectors for selecting video sequences based on the target vectors. This is preferably done by appending a new target vector to the sequence list or to the target vectors already present in the sequence list.
[0084] In a step S18, video sequences are selected from a video database based on the sequence list to generate a video clip. The video database can be a database that can contain multiple video sequences from current films or videos, but also archive material, which can preferably also consist of shorter video sequences or short video clips. Alternatively, the video database can also be implemented on an external server or on multiple storage media, for example, by individual video clips that can be stored on the storage media or at different locations.
[0085] In a preferred embodiment, steps S10 to S16 are repeated multiple times, namely for each new text segment, to generate a sequence list with multiple target vectors. In a preferred embodiment, an interview feature vector can be generated in step S20. This is inserted into the sequence list. The interview feature vector corresponds to an interview sequence or to a video sequence comprising an interview sequence.
[0086] Stylistic features within the meaning of this invention are, for example, features that encompass the overall content-related consistency of text or text segments and scenes or video sequences or images. For this purpose, the content of both the text or text sequences and the scenes, images, or video sequences can be represented in the form of characteristic quantities or coded numbers, for example, in the form of numerical values or in the form of colors or color combinations.
[0087] A match (of coded features) can be determined if the coding of the text as well as the scenes and video sequences are the same or largely the same, or if the feature vectors resulting from the respective text segments and video scenes show a high degree of similarity at individual elements of the vector.
[0088] Further stylistic means within the meaning of the invention relate to the sequence of shot sizes of a video scene or video sequence. These include, for example, the long shot, medium shot, portrait shot, or detail shot. The stylistic means in this context provides information about how the corresponding stylistic sizes or features are changed. It has proven advantageous if one video sequence shows a long shot, the subsequent video sequence includes a medium shot or portrait shot, and the subsequent video sequence includes a detail shot.
[0089] Another stylistic feature for video classification or video classification attribute involves camera movement. Here, it can be considered that the camera movement is relatively smooth and that the same camera movement does not occur repeatedly in different sequences. For example, it has proven disadvantageous within the scope of the invention if the camera performs a panning movement in the same direction several times in a row. This gives the viewer the impression that they are rotating around their own axis, which is undesirable and undesirable, especially in news broadcasts and news videos.
[0090] Another stylistic feature or characteristic used as a video classification feature is scene quality. This includes features such as image sharpness, exposure, or camera shake or unpredictable camera movements.
[0091] Other stylistic features that should be considered include scene composition. These features express whether the described subject or object contained in the corresponding text segments also appears in the video sequence. In particular, the location of the object or subject can be considered and expressed, or whether the object or subject is located in the center of the image or at a peripheral position within the image.
[0092] The consideration of stylistic characteristics or features in the video classification features, which are taken into account when creating the prediction vector, results in the successive video sequences and the corresponding text segments matching each other and fitting together well in the sense that, especially for news videos and information videos, a calm and informative image and scenario is created that, on the one hand, matches and corresponds to the underlying text and the individual text segments and, on the other hand, evokes an atmosphere in the viewer in which they can best absorb the information content generated by the images and text. The stylistic means include the content-related consistency of text and scenes, the sequence of shot sizes (long shot, medium shot, close-up), the sequence of camera movements (still, pan, or zoom), the scene composition (position of the subject), and the image quality.
[0093] The invention has been comprehensively described and explained with reference to the drawings and the description. The description and explanation are to be understood as exemplary and not restrictive. The invention is not limited to the disclosed embodiments. Other embodiments or variations will become apparent to those skilled in the art upon use of the present invention and upon careful analysis of the drawings, the disclosure, and the following claims.
[0094] In the claims, the words "comprising" and "having" do not exclude the presence of further elements or steps. The undefined article "a" or "an" does not exclude the presence of a plurality. A single element or unit can perform the functions of several of the units recited in the claims. An element, unit, device, and system can be partially or completely implemented in hardware and / or software. The mere mention of some measures in several different dependent claims should not be understood to mean that a combination of these measures cannot also be used advantageously. A computer program can be stored / distributed on a non-volatile data carrier, for example on an optical memory or on a solid-state drive (SSD).A computer program may be distributed together with hardware and / or as part of hardware, for example, via the Internet or via wired or wireless communication systems. Reference signs in the patent claims are not to be construed as limiting.
Claims
Patent claims 1. A device for composing a video clip (72) based on multiple video sequences (74) and text segments (12) of a text; comprising a text module (30) for generating a text feature vector (32) with text features of a text segment (12); a transformer module (20) for generating a prediction vector (22) for a video sequence (74); a selection module (40) for selecting a target vector (42) of a suitable target video sequence from a feature bank (50) based on the prediction vector (22); an assignment module (60) for creating a sequence list (62) with target vectors (42) for selecting video sequences (74) based on the target vectors (42); a video clip module (70) for creating a video clip (72) from video sequences (74) of a video database based on the sequence list (62); wherein the transformer module (20) comprises a cross-modal encoder (24) and a feature decoder (28);and the transformer module (20) is configured to create the prediction vector (22) based on a video vector (44) of the previous video sequence (74) and on a text feature vector (32); the video vectors (44) comprise video classification features; stylistic feature.
2. Device according to claim 1, characterized in that the video vector (44) and the prediction vector (42) comprise text features, video content features and video classification features.
3. Device according to the preceding claim, characterized in that the video classification features are generated by means of a VGG encoder (86), wherein the VGG encoder (86) is preferably integrated in the transformer module (20).
4. Device according to claim 2 or 3, characterized in that the video classification features comprise features of the video sequence (74), wherein the features are elements of a list relating to setting variables of the video sequence (74) such as long shot, detailed view, portrait view or medium view, camera angle used in the video sequence (74), movement of the camera in the video sequence (74), direction of movement of the camera, movement of persons in the video sequence (74), direction of movement of persons in the video sequence (74), zoom settings and similar settings of the video sequence (74).
5. Device according to claim 2, characterized in that the video content features are generated by means of a CLIP encoder (82), wherein the CLIP encoder (82) is preferably integrated in the transformer module (20).
6. Device according to one of the preceding claims, characterized in that the transformer module (20) comprises a plurality of submodules, each submodule having a neural network and being designed to analyze and evaluate a predetermined video classification feature of a video vector (44) in order to take the video classification feature into account when creating the prediction vector (22), wherein the submodule preferably comprises a Feed-forward network and wherein the video classification feature preferably comprises a stylistic feature.
7. Device according to one of the preceding claims, characterized in that the transformer module (20) is designed to create the prediction vector (22) based on a plurality of video vectors (44) and a plurality of text feature vectors (32), preferably based on the video vectors (44) of all available video sequences (74) and / or on all available text feature vectors (32).
8. Device according to one of the preceding claims, characterized in that the text module (30) comprises a CLIP encoder (82), wherein the text feature vector (32) is generated by means of the CLIP encoder (82).
9. Device according to one of the preceding claims, characterized in that the creation of the sequence list (62) for the selection of video sequences (74) is carried out by appending a target vector (42) to the sequence list (62).
10. Device according to one of the preceding claims, characterized in that the text segment (12) comprises a sentence of the predetermined text.
11. Device according to one of the preceding claims, characterized in that the selection module (40) is designed to compare the video vectors (44) stored in the feature bank (50) with the prediction vector (22), preferably individually in each case, and particularly preferably the quality of the selection of the target vector (42) is assessed, very preferably by means of a quality function.
12. Device according to one of the preceding claims, characterized in that in the sequence list (62) interview feature vectors (58) corresponding to video sequences (74) comprising an interview sequence (55).
13. Training device for training a transformer module (20) and for creating video vectors (44) for video sequences (74) based on text segments (12) of a text or news article; comprising a CLIP encoder (82) for creating text feature vectors (32) with text features of a text segment (12); a CLIP encoder (82) for creating content vectors (84) with video content features for a video sequence (74); a VGG encoder (86) for creating classification vectors (88) with video classification features for a video sequence (74);the transformer module (20) with a cross-modal encoder (24) and a feature decoder (28), wherein the cross-modal encoder (24) is designed to combine a text feature vector (32) and a content vector (84) as well as a classification vector (88) for a video sequence (74) into a feature vector (26) and to feed it to the feature decoder (28), wherein the feature vector (26) preferably has a greater length than the content vector (84), very preferably twice the length; the feature vector (26) comprises stylistic features or features; the feature decoder (28) is designed to create a video vector (44) from the feature vector (26), which video vector has the features of a video sequence (74) present at the input of the training device (80) and preferably has half the length of the feature vector (26); wherein the training device (80) outputs the video vector (44) at an output and makes it available for further processing.
14. Training device according to the preceding claim, characterized in that the training device is designed to form the feature vector (26) from a plurality of text feature vectors (32) and / or a plurality of content vectors (84) and / or a plurality of classification vectors (88).
15. A method for composing a video clip (72) based on multiple video sequences (74) and text segments (12) of a text; comprising the following steps: a) generating a text feature vector (32) with text features of a text segment (12); b) generating a prediction vector (22) for a video sequence (74) based on a video vector (44) of a previous video sequence (74) and based on the text feature vector (32) of the text segment (12); c) selecting a target vector (42) of a target video sequence from a feature bank (50) based on the prediction vector (22); d) creating a sequence list (62) with target vectors (42) for selecting video sequences (74) based on the target vectors (42) by appending a target vector (42) to the sequence list (62); e) Creating a video clip (72) from video sequences (74) of a video database based on the sequence list (62).
16. Method according to the preceding claim, characterized in that steps a) to d) are carried out several times in order to generate the sequence list (62).
17. The method according to claim 15 or 16, characterized in that an interview feature vector (58) is added to the sequence list (62), which corresponds to a video sequence (74) comprising an interview sequence (55).
18. The method according to one of claims 15 to 17, characterized in that the generation of the prediction vector (22) for the video sequence (74) is based on a plurality of video vectors (44) and a plurality of text feature vectors (32), preferably based on the video vectors (44) of all available video sequences (74) and / or on all available text feature vectors (32).
19. A computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the method according to claim 11.
Citation Information
Patent Citations
Short video news generation system based on large-scale pre-training model
CN117041458A
Generating digital video summaries utilizing aesthetics, relevancy, and generative neural networks
US20190377955A1