Moment prediction generation

Advanced machine learning techniques integrate features from multiple media modalities to automate the search for relevant moments in large media libraries, addressing inefficiencies and inconsistencies in manual methods and deterministic algorithms.

WO2025144660A1PCT designated stage expired Publication Date: 2025-07-03DOLBY LABORATORIES LICENSING CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/060801
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-19
Filing Date
2024-12-18
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Manual search processes for identifying relevant moments in large media libraries are time-consuming, inefficient, and prone to subjectivity and inconsistency, while deterministic algorithms struggle with the complexity and variability of audio and visual content, failing to integrate different media types and maintain context.

Method used

Utilizing advanced machine learning techniques to extract features from multiple media modalities (audio, video, and text) and integrate them into a cohesive representation, capturing interactions and correlations, followed by a moment prediction model to identify relevant moments.

Benefits of technology

Provides accurate, efficient, and scalable search results by automating the analysis of diverse media content, reducing subjectivity and enhancing user experience with precise and contextually relevant outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060801_03072025_PF_FP_ABST
    Figure US2024060801_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method includes generating, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality. The extracted features include a lower-dimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality. The method includes generating, at a multi-modal fusion model, an integrated representation based on the extracted features. The integrated representation includes attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality. The method includes generating, at a moment prediction model, a moment prediction based on the integrated representation. The moment prediction identifies a moment in the input media.
Need to check novelty before this filing date? Find Prior Art

Description

MOMENT PREDICTION GENERATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to Spanish (ES) Patent Application No. P202331088 filed December 27, 2023, and claims the benefit of U.S. Provisional Application No. 63 / 561,206 filed March 4, 2024 and U.S. Provisional Application No. 63 / 722,181 filed November 19, 2024 . The entire disclosures of the above applications are incorporated by reference.FIELD

[0002] The present disclosure relates to media content analysis and, more particularly, to techniques for searching multi-modal media content using artificial intelligence.BACKGROUND

[0003] Manually searching through media content (such as, for example, audio and visual content) to identify moments related to search queries can be impractical for large media libraries due to the significant amount of time required to review extensive amounts of content. For example, the sheer volume of content can make it difficult for a human searcher to analyze every piece of content systematically and comprehensively. Thus, manual search processes are not only timeconsuming but also mentally taxing, which may lead to decreased efficiency and accuracy.Furthermore, the inherent subjectivity of human searchers means that different individuals might interpret the same search query or piece of media content in different ways, leading to inconsistent results. Searchers may also miss relevant moments due to lapses in attention or misunderstanding the context of the search query or the media content, resulting in incomplete or idiosyncratic findings that do not fully align with the intended search query.

[0004] Automated solutions may be desirable to address these concerns as automation offers the potential for consistent, efficient, and comprehensive search and analysis of large media libraries. However, programming conventional deterministic algorithms to understand media content for automated search and analysis can be difficult due to the inherent complexity and variable of media content (e.g., audio and visual content). These types of content can contain a vast array of information that includes subtle nuances and context-dependent elements, which may be challengingfor fixed logical rules (e.g., deterministic algorithms) to quantify or capture using fixed logical rules. Furthermore, because media content may unfold over time, deterministic algorithms may be required to handle temporal dependencies within the content. However, deterministic algorithms often struggle to integrate different media types and maintain the necessary context to understand sequences of events. For example, the high-dimensional nature of audio and visual data may need sophistical feature extraction before they can be efficiently analyzed, but deterministic algorithms may lack the adaptability to efficiently identify and prioritize relevant features. Additionally, the ambiguity and var iability of human language and visual information make it impractical to create an exhaustive set of rules to cover every possible scenario.

[0005] Furthermore, deterministic algorithms may also be difficult to scale across different content types and contexts. For example, the complexity of deterministic rules may need to increase exponentially to address each new type (e.g., modality) or context of content. Additionally, deterministic algorithms may lack the learning capability to generalize from examples and therefore may be rigid and unable to adapt to new or unexpected variations in content. Nor are deterministic methods robust to the noise and errors commonly found in real- world audio and visual media, which may contribute to inaccurate results. By contrast, machine learning techniques excel at recognizing complex patterns and relationships within high-dimensional data (such as audio and video content), capturing the contextual information within the data, and adapting to and learning from new content. Thus, solutions described herein use machine learning techniques to provide scalable, flexible, and adaptive solutions to the challenges that deterministic algorithms face in understanding and searching diverse libraries of media content.SUMMARY

[0006] Systems, apparatuses, methods, and techniques described in this specification provide technical solutions to these challenges (among others) by using advanced machine learning techniques to understand and search media content (such as, for example, audio, video, and / or text data). Machine learning techniques may be used to extract distinct features from each type of media, capturing the unique and relevant aspects of each type of content. These extracted features may then be integrated into a single, cohesive representation that captures the interactions between different data types, allowing for a more comprehensive understanding of the content. The integrated representation may then be subsequently analyzed using machine learning techniques, which maygenerate outputs indicating moments in the searched media content that are relevant to the search query.

[0007] By extracting rich and detailed features from each media content modality, the systems, apparatuses, and techniques described in this specification may perform a comprehensive analysis of the media content, recognizing complex patterns and contextual information that a single-modality approach might miss. The integration of multi-modal features ensures that complementary information from different data types is effectively combined, which may lead to more accurate and insightful interpretations. Furthermore, the use of automated machine learning techniques ensures consistency and objectivity, providing users with precise and contextually relevant results that may improve the accuracy of the searches by removing the subjectivity and idiosyncrasies introduced by manual human searches. Additionally, the machine learning techniques described herein are scalable and flexible, capable of handling large, diverse media libraries, and are adaptable to various types of media content. Accordingly, systems, apparatuses, methods, and techniques described herein significantly enhance the user experience by delivering accurate and meaningful search results quickly and efficiently, even when applied to large batches of diverse media content.

[0008] A computer-implemented method includes generating, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality. The extracted features include a lower-dimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality. The method includes generating, at a multi-modal fusion model, an integrated representation based on the extracted features. The integrated representation includes attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lowerdimensional numerical representation of the second modality. The method includes generating, at a moment prediction model, a moment prediction based on the integrated representation. The moment prediction identifies a moment in the input media.

[0009] In other features, the input media includes a third modality, the first modality includes visual media, the second modality includes text media, and the third modality includes audio media. In other features, the feature extraction model includes a visual feature extraction model configured to extract visual features from the first modality, a text feature extraction model configured to extract text features from the second modality, and an audio feature extraction model configured to extract audio features from the third modality. In other features, the multi-modal fusion modelincludes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

[0010] In other features, the combined representation includes interactions between the extracted visual features and the extracted audio features. In other features, the multi-modal fusion model includes a cross-attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features. In other features, the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features. In other features, the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

[0011] In other features, the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media. In other features, the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation. In other features, the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

[0012] In other features, the moment prediction model includes a span predictor model configured to generate the span prediction based on the decoded representation and a binary classifier model configured to generate the validity prediction based on the decoded representation. In other features, the method, further includes generating, based on the moment prediction, trimmed media including the identified moment in the input media. In other features, a non-transitory computer-readable medium includes executable instructions that, when executed by an electronic processor, causes the electronic processor to perform the method.

[0013] A system includes memory hardware storing instructions and processor hardware configured to execute the instructions. Executing the instructions causes the processor hardware to generate, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality, the extracted features including a lowerdimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality, generate, at a multi-modal fusion model, an integrated representation based on the extracted features, the integrated representation including attentionscores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality, and generate, at a moment prediction model, a moment prediction based on the integrated representation. The moment prediction identifies a moment in the input media.

[0014] In other features, the input media includes a third modality, the first modality includes visual media, the second modality includes text media, and the third modality includes audio media. In other features, the feature extraction model includes a visual feature extraction model configured to extract visual features from the first modality, a text feature extraction model configured to extract text features from the second modality, and an audio feature extraction model configured to extract audio features from the third modality. In other features, the multi-modal fusion model includes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

[0015] In other features, the combined representation includes interactions between the extracted visual features and the extracted audio features. In other features, the multi-modal fusion model includes a cross-attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features. In other features, the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features. In other features, the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

[0016] In other features, the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media. In other features, the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation. In other features, the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

[0017] In other features, the moment prediction model includes a span predictor model configured to generate the span prediction based on the decoded representation, a binary classifier model configured to generate the validity prediction based on the decoded representation. In other features, executing the instructions further causes the processor hardware to generate, based on themoment prediction, trimmed media including the identified moment in the input media. In other features, the moment is identified by one or more temporal markers identifying a temporal segment of media.

[0018] Other examples, embodiments, features, and aspects will become apparent by consideration of the detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0019] FIG. 1 is a block diagram illustrating an example computing system for analyzing media content, according to some embodiments.

[0020] FIG. 2 is a block diagram illustrating example data flow between a feature extraction model, a multi-modal fusion model, and a moment prediction model, according to some embodiments.

[0021] FIG. 3 is a block diagram illustrating example data flow of a visual feature extraction model, an audio feature extraction model, and a text feature extraction model, according to some embodiments.

[0022] FIG. 4 is a block diagram illustrating example data flow of an audio-visual fusion model and a cross-attention transformer model, according to some embodiments.

[0023] FIG. 5 is a block diagram illustrating example data flow of a transformer encoder model, a transformer decoder model, a span predictor model, and a binary classifier model, according to some embodiments.

[0024] FIG. 6 is a block diagram illustrating data flow between the trim composition application and the post edit application, according to some embodiments.

[0025] FIG. 7 is a message sequence chart illustrating example interactions between components of a computing system, according to some embodiments.

[0026] FIG. 8 is a message sequence chart illustrating example interactions between components of a computing system, according to some embodiments.

[0027] FIG. 9 is a flowchart illustrating an example process for generating a moment prediction, according to some embodiments.

[0028] In the drawings, reference numbers may be reused to identify similar and / or identical elements.DETAILED DESCRIPTION

[0029] FIG. 1 is a block diagram illustrating an example computing system 100 for analyzing media content, according to some embodiments. As illustrated in FIG. 1, some examples of the system 100 include a media analysis platform 102, one or more user devices 104 (such as user device 104-1 and user device 104-2), and / or a communications system 106. Although two user devices 104 are shown in the example of FIG. 1, the system 100 may include any number of user devices 104. In some examples, one or more of the user devices 104 and / or the communications system 106 are omitted from the system 100. In various implementations, the media analysis platform 102, the user devices 104, and / or other components described herein may communicate via the communications system 106.

[0030] In some examples, the media analysis platform 102 is a computing platform deployed across a single server or one or more servers. The media analysis platform 102 may use advanced machine learning techniques to understand and search media content and identify key moments in the content. A “moment” may refer to a specific temporal segment of media identified by one or more temporal markers (such as, for example, timestamps, durations, intervals, etc.). For example, a moment may be characterized by distinctive start and end points in the media. In various implementations, a moment is defined by a timestamp or an identification of a precise temporal location in the media and / or a width or a surrounding duration. For example, a moment may be represented by a pair of timestamps that mark the beginning and end of a segment of media. In other examples, a moment may be represented by a timestamp representing a center point along with a defined width representing temporal segments before and / or after the center point. In various implementations, the user devices 104 include one or more computing platforms, such as smartphones, tablet computers, laptop computers, desktop computers, computer servers, etc.

[0031] In some examples, the communications system 106 includes one or more networks, such as a General Packet Radio Service (GPRS) network, a Time-Division Multiple Access (TDMA)network, a Code-Division Multiple Access (CDMA) network, a Global System of Mobile Communications (GSM) network, an Enhanced Data Rates for GSM Evolution (EDGE) network, a High-Speed Packet Access (HSPA) network, an Evolved High-Speed Packet Access (HSPA+) network, a Long Term Evolution (LTE) network, a Worldwide Interoperability for Microwave Access (WiMAX) network, a 5th-generation mobile network (5G), an Internet Protocol (IP) network, a Wireless Application Protocol (WAP) network, or an IEEE 802.11 standards network, as well as any suitable combination of the above networks. In various implementations, the communications system 106 includes an optical network, a local area network, and / or a global communication network, such as the Internet.

[0032] As illustrated in the example of FIG. 1, the media analysis platform 102 may include system resources 108, a communications interface 110, non-transitory computer-readable storage media, such as, for example, storage 112 and / or a media library 114. The non-transitory computer- readable storage media may contain instructions that, when executed, cause one or more electronic processors (such as, for example, electronic processors of the system resources 108) or neural processors to perform various functions described herein. The system resources 108 may include one or more electronic processors, one or more graphics processing units, volatile computer memory, non-volatile computer memory, and / or one or more system buses interconnecting various components of the media analysis platform 102. The communications interface 110 may include hardware and / or software components that communicate with other devices, platforms, and / or systems over the communications system 106. In various implementations, the communications interface 110 includes one or more transceivers for sending and / or receiving data over the communications system 106.

[0033] In some examples, the system 100 may include a media library 116 deployed at a location remote from the media analysis platform 102. The media analysis platform 102 may communicate with the media library 116 via the communications system 106. In various implementations, one or more of the user devices 104, the media library 114, and / or the media library 116 may include media content, such as, for example, audio content, video content, and / or metadata. Examples of metadata include titles, summaries, content-type annotations, viewer annotations, and / or engagement data. Content-type annotations may include data indicating a category of user-generated content (such as, for example, instructional videos, tutorials, egocentric videos, vlogs, product reviews, reaction videos, gaming videos, podcasts, educational videos, liveevents such as sporting events and concerts, etc.). In various implementations, content-type annotations are user-generated and / or generated by a content classifier model.

[0034] Metadata associated with media content can be dynamic or static, depending on its nature and intended purpose. Dynamic metadata may change over time, reflecting variations in user interaction and / or content performance. For example, engagement data — such as view counts, likes, and / or user comments, etc. — may fluctuate based on audience behavior. By contrast, static metadata may remain constant or substantially constant over time, providing foundational information about the content. Examples of static metadata may include titles, summaries, and / or original content-type annotations. Accordingly, metadata may capture both enduring (e.g., static metadata) and evolving (e.g., dynamic metadata) characteristics of media content, providing a range of technical benefits to the machine learning applications described herein.[00351 For example, static metadata, such as titles and summaries, may serve as reliable anchors for initial classification tasks, search indexing, and similar processes. Dynamic metadata, by contrast, may provide real-time or near-real-time feedback and behavioral insights, which can help machine learning models adapt to changes in user interactions. For instance, engagement metrics — such as view counts and / or reactions, etc. — may supply valuable context for models performing tasks such as content recommendation or trend analysis. Thus, static and / or dynamic metadata may provide both stable and adaptive signals to various machine learning models described herein, enhancing their ability to perform nuanced tasks.

[0036] Examples of viewer annotations may include user-generated input that provide additional context or insights for the media content. In various implementations, viewer annotations include comments, highlights, tags, notes, reactions, bookmarks, transcriptions, translations, subtitles, visual annotations, etc. Comments may include text-based remarks or observations about parts of the media content. Highlights may include key moments or segments in the media content. Tags may include keywords or phrases added by viewers to categorize and / or describe the content. Notes may include detailed annotations or explanations added to sections of the content. Reactions may include emoticons, emojis, or brief reactions (such as thumbs up or thumbs down) indicating the viewer’s sentiment towards the content. Bookmarks may mark moments in the content that users deem important. Transcriptions and subtitles may include text versions of spoken dialogue or audio content, for example, with time stamps. Translations may include text versions of the content indifferent languages. Visual annotations may include drawings or sketches over video or image content to highlight or illustrate portions of the content.

[0037] In various implementations, the storage 112 may include a feature extraction model 118, a missing modality generation model 120, a multi-modal fusion model 122, a moment prediction model 124, a trim composition application 126, and / or a post edit application 128. FIG. 2 is a block diagram 200 illustrating example data flow between the feature extraction model 118, the multimodal fusion model 122, and the moment prediction model 124, according to some embodiments. Referring collectively to FIGS. 1 and 2, the feature extraction model 118 may receive input media 202 (such as, for example, raw or preprocessed representations of audio media, visual media, and / or text media) from various sources (such as, for example, one or more of the user devices 104, the media library 114, and / or the media library 116) and generate extracted features 204 based on the input media 202. For example, the feature extraction model 118 may extract relevant features from the input media 202 and generate a lower-dimensional representation of the input media 202 that preserves important characteristics of the input media 202 for processing by downstream machine learning models. In various implementations, the lower-dimensional representations may be embeddings, which may be numerical representations (such as, for example, vectors, matrices, and / or tensors) of the input media 202 in a lower-dimensional space. Embeddings may simplify and reduce the dimensionality of the input media 202 while preserving important information (such as, for example, the relationships between data points), transforming the input media 202 into a representation that is more efficient for downstream machine learning models to process and understand.

[0038] In some examples, the feature extraction model 118 performs a series of preprocessing steps to convert the input media 202 into a format suitable for analysis. Then, the feature extraction model 118 may perform feature extraction operations on the preprocessed input media 202. For example, the preprocessing steps may ensure that the input media 202 is standardized, denoised, and in a format ready for feature extraction. For visual media such as videos, preprocessing may include decomposing the video into individual frames, resizing frames to a standard resolution, applying color normalization, and performing noise reduction. The preprocessed video frames may be represented as tensors, with each frame stored as a matrix of pixel values, and the entire sequence of frames forming a 3D or 4D tensor. For audio media, preprocessing may include sampling the audio at a consistent rate, performing noise reduction, and converting the audio into representations suchas spectrograms or Mel-Frequency Cepstral Coefficients (MFCCs). The preprocessed audio can be represented as vectors or matrices. For text media, preprocessing may include tokenizing the text, normalizing it, and mapping tokens to numerical values. The preprocessed text may be represented as vectors or matrices, with each token converted into a high-dimensional vector.

[0039] The feature extraction model 118 may include one or more feature extraction models, each configured to process a specific modality of input media and extract relevant features from the modality for downstream machine learning applications. Generally, the feature extraction models may receive raw or preprocessed input media and generate lower-dimensional representations (such as, for example, lower-dimensional numerical representations or embeddings) that capture relevant characteristics of the input media. The lower-dimensional representations may serve as compact and efficient representations for downstream tasks such as classification, fusion, prediction, etc.

[0040] In various implementations, the feature extraction model 118 may include one or more feature extraction models, each tailored to a particular modality of input media. For example, the feature extraction model 118 may include a visual feature extraction model 130 that processes visual media (such as, for example, images and / or videos) to extract relevant features such as object presence, motion patterns, scene context, etc. In various implementations, the feature extraction model 118 includes an audio feature extraction model 132 that processes audio media to extract relevant features such as pitch, rhythm, speech content, etc. In some examples, the feature extraction model 118 includes a text feature extraction model 134 that processes text-based media (such as, for example, captions, transcriptions, etc.) to extract relevant features such as semantic information, syntactic information, etc. Each feature extraction model may convert its respective input into lower-dimensional representations such as embeddings for further analysis by downstream machine learning models.

[0041] While FIG. 1 illustrates an example where the feature extraction model 118 includes visual, audio, and text feature extraction models 130-134, other implementations may include any combination of feature extraction models depending on the requirements of the system 100 and the available input media. For instance, if audio media is unavailable or unnecessary, only visual and text feature extraction models may be used. In scenarios where a modality is unavailable but still desired, the feature extraction model 118 may generate features for the “missing” modality using available modalities. For example, if text features are required but not present, the system 100 may employ a missing modality generation model 120 to synthesize the text modality from visual and / oraudio data. This flexible configuration allows the feature extraction model 118 to handle any combination of input modalities. As a specific example of this flexibility, FIG. 3 illustrates an implementation where the feature extraction model 118 includes visual, audio, and text feature extraction models.

[0042] FIG. 3 is a block diagram 300 illustrating example data flow of the visual feature extraction model 130, the audio feature extraction model 132, and the text feature extraction model 134, according to some embodiments. Referring collectively to FIGS. 1 and 3, in various implementations, the input media 202 includes any combination of input visual media 302, input audio media 306, and input text media 310. In some examples, the input visual media 302 and / or the input audio media 306 include the media to be searched (or preprocessed representations of the media to be searched). In various implementations, the input text media 310 includes the search query and / or metadata. The input visual media 302, the input audio media 306, and / or the input text media 310 may retrieved from and / or provided by any of the user devices 104, the media library 114, and / or the media library 116. In some examples, one or more of the input visual media 302, the input audio media 306, and the input text media 310 may be missing, and the missing modality generation model 120 generates the missing modality.

[0043] In examples where the input media 202 includes input visual media 302 and / or input audio media 306 but does not include input text media 310, the missing modality generation model 120 may include a dense captioning system that generates input text media 310 based on the input visual media 302 and / or input audio media 306. The generated input text media 310 may include a series of detailed textual descriptions corresponding to moments in the input visual media 302 and / or the input audio media 306. In various implementations, the dense captioning system extracts features from the input visual media 302 and / or the input audio media 306 and processes the extracted features using a cross-modal attention mechanism, which finds correlations between the auditory and / or visual features. A sequence generation model (such as, for example an LSTM network and / or a transformer model) takes the integrated multi-modal features and produces sequences of words that form detailed textual descriptions based on the context of moments in the audio and / or visual media.

[0044] In some examples where the input media 202 includes input visual media 302 and / or input audio media 306 but does not include input text media 310, the missing modality generation model 120 may use the title of the audio and / or visual media (e.g., from the metadata) as the inputtext media 310. In some examples where the input media 202 includes input audio media 306 (but is missing corresponding input visual media 302) or includes input visual media 302 (but is missing corresponding input audio media 306), the missing modality generation model 120 may use models such as a CLIP model, a CLAP model, an ImageBind model, etc., to learn joint representations across multiple modalities such as audio and visual media. For example, when video media is missing corresponding audio media, the missing modality generation model 120 may use any of the previously described models to generate a corresponding audio embedding and / or a corresponding audio sample.

[0045] In various implementations, the feature extraction model 118 may treat the missing modality as a learnable parameter that needs to be learned during training of the feature extraction model 118. For example, the feature extraction model 118 may be trained to generate extracted features that include estimated embeddings representing the missing modality.

[0046] The visual feature extraction model 130 may receive input visual media 302 (for example, preprocessed video media represented as three- or four-dimensional tensors), process video frames from the input visual media 302 to identify and extract key features and combine the extracted features to form lower-dimensional embeddings (such as, for example, vectors, matrices, or tensors) that retains essential visual information. The visual feature extraction model 130 may output the lower-dimensional embeddings as extracted visual features 304. The visual feature extraction model 130 may include any combination of convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTMs), transformer models, and autoencoders. CNNs may capture spatial hierarchies and patterns in visual data through a layered structure of convolutions and pooling. RNNs and LSTMs may capture temporal dependencies and sequential information in video data. For example, convolutional LSTMs combine spatial feature extraction layers (e.g., convolutional layers) with temporal sequence modeling units (e.g., LSTM units), capturing spatiotemporal patterns like motion and object trajectories across frames. 3D CNNs may capture temporal features by extending convolutions into the temporal dimension, learning motion dynamics and temporal changes in addition to spatial features.

[0047] Transformer models may capture long-range dependencies and contextual relationships in video data through self-attention mechanisms. Autoencoders can learn compressed representations of input data, which include essential features that enable reconstruction. For example, convolutional autoencoders may capture the most significant spatial features required torestore an input image, learning to ignore noise and redundant information. Variational autoencoders may capture probabilistic representations of features, enabling the model to learn diverse and robust features that account for variations in the data. SlowFast networks may use a dual-pathway approach to capture different types of temporal features. For example, the “slow” pathway may capture detailed spatial semantics and slow-moving actions while the “fast” pathway captures rapid motion and temporal dynamics.

[0048] In various implementations, the visual feature extraction model 130 receives input visual media 302 along with preprocessed text associated with the input visual media 302 (such as, for example, titles, subtitles, captions, descriptions, annotations, etc.) as inputs to a contrastive language-image pretraining (CLIP) model (such as the VideoCLIP model), which outputs a lowerdimensional representation including features suitable for downstream object identification applications (e.g., to identify people, vehicles, etc. in the visual media). The CLIP model may use a jointly trained image encoder and text encoder to generate a joint embedding space where images and their corresponding text descriptions are closely aligned. The image encoder may include one or more CNNs that produces a higher-dimensional feature vector and / or one or more vision transformers (ViTs) that process images and generate fixed-size feature vectors. The text encoder may include one or more transformer models that process textual inputs and generate fixed-size text embeddings. The image and text encoders may be jointly trained using a contrasting learning objective that minimizes the distance between embeddings of matching image-text pairs and maximizes the distance between non-matching pairs.

[0049] In some examples, the visual feature extraction model 130 receives input visual media 302 as inputs to a SlowFast network. As previously described, the SlowFast network may include a “slow” pathway and a “fast” pathway to capture different aspects of the temporal dynamics in video data. The “slow” pathway may process video frames at a low temporal resolution, capturing detailed spatial information but with slower motion. The “fast” pathway may process video frames at a high temporal resolution, capturing fast-moving actions and / or changes in the frames. Lateral connections between the “slow” pathway and the “fast” pathway may fuse the features between the pathways at various stages. The SlowFast network may also include temporal convolutional networks that merge features output from both pathways at multiple stages and aggregate the features to create a comprehensive spatiotemporal representation. Final classificationlayers may pool the spatiotemporal representation and output probabilities for each class, providing a final prediction of actions being performed in the video.

[0050] In various implementations, the visual feature extraction model 130 receives input visual media 302 as inputs an HSEmotion model. The HSEmotion model may detect and analyze facial emotions in the video frames and output a lower-dimensional representation that may be used in facial emotion recognition applications (e.g., to detect periods in the visual media where subjects display heightened valence or arousal). For example, the HSEmotion model may generate an output indicating the detailed temporal dynamics of emotions, such as an arousal arc indicating the state of excitement over time and / or a discrete heatmap showing the most prominent emotion for each time frame. The HSEmotion model may include an input processing model, a feature extraction backbone, and an emotion classification head. The input processing model may include a facial detection network that detects faces in images and video frames and preprocesses the faces to align them to a standard orientation and size. The feature extraction backbone may include a CNN that processes the preprocessed face images and extracts hierarchical features from the preprocessed face images. The emotion classification head may include fully connected layers and / or softmax layers that flatten features extracted by the CNN and output probabilities associated with each discrete emotion class.

[0051] In some examples, the visual feature extraction model 130 receives input visual media 302 as inputs and performs optical character recognition on the visual media to generate and output extracted visual features 304 including text embeddings. The visual feature extraction model 130 may use models such as the Efficient and Accurate Scene Text Detector (EAST) model or the Connectionist Text Proposal Network (CTPN) to detect regions of interest containing text in the video data. The visual feature extraction model 130 may use a convolutional recurrent neural network (CRNN) and / or transformer model to convert text in the detected text regions into textual data. The visual feature extraction model 130 may use a language model to generate embeddings for the detected text.

[0052] In various implementations, the input visual media 302 may be preprocessed to divide input video into n segments. In some examples, the visual feature extraction model 130 generates extracted visual features 304 for each of the n segments. Accordingly, the visual feature extraction model 130 may output extracted visual features 304 that represent the n segments of the input video.

[0053] The audio feature extraction model 132 may receive input audio media 306 as inputs (for example, preprocessed audio media represented as vectors, matrices, or tensors), process the audio to identify and extract key features and combine the extracted features to form lower-dimensional embeddings (such as, for example, vectors, matrices, or tensors) that retain significant characteristics of the audio signal. The audio feature extraction model 132 may output the lowerdimensional embeddings as extracted audio features 308. The audio feature extraction model 132 may include any combination of CNNs, RNNs, LSTMs, gated recurrent units (GRUs), CRNNs, transformer models, and autoencoders. CNNs may capture spatial features from representations such as spectrograms or Mel-Frequency Cepstral Coefficients (MFCCs). For example, 2D CNNs may capture spectral and temporal features by applying 2D convolutions to spectrograms or MFCCs to capture patterns such as formants, harmonics, and temporal changes. ID CNNs may capture temporal features (such as amplitude and frequency changes over time) directly from raw audio waveforms.

[0054] RNNs, LSTMs, and GRUs may capture temporal dependencies and sequential information in audio signals, such as rhythm and / or melody in music, and / or phonetic and / or prosodic features in speech. Transformer models may capture global context and long-range dependencies in audio sequences, such as the intricate patterns in music and speech spanning longer durations. Autoencoders may learn compressed representations of audio data, focusing on essential features for reconstruction. For example, convolutional autoencoders may capture essential spectral features required to reconstruct the audio signals (e.g., filtering out noise and redundant information). Variational autoencoders may output probabilistic latent representations of audio features, capturing a range of possible variations in the audio data.

[0055] In various implementations, the audio feature extraction model 132 receives input audio media 306, processes the input audio media 306 to extract features that may be used to identify key audio events (e.g., laughter, shouting, etc.), and outputs extracted features as a lower-dimensional representation of the audio events suitable for use in audio event detection applications. The audio feature extraction model 132 may include a pre-trained audio neural network (PANN) that converts the audio media into a time-frequency representation (e.g., a Mel- spectrogram, a log Mel- spectrogram, etc.), detects patterns and extracts features from the time-frequency representation, and classifies the extracted features into predefined categories of audio events. The PANN may also produce a probability distribution indicating a likelihood of each event being present in the inputaudio. The audio feature extraction model may include an audio spectrogram transformer (AST) model that converts audio inputs into a time-frequency representation, uses the self-attention mechanisms of a transformer model to capture global dependencies and relationships between different parts of the time-frequency representation, and generates a probability distribution over possible audio events, indicating a likelihood of each audio event being present.

[0056] In some examples, the audio feature extraction model 132 receives input audio media 306 as inputs, processes the input audio media 306 to extract features that may be used in automatic speech recognition applications to the produce a text transcript of the audio media, and outputs a lower-dimensional representation of the audio and / or outputs the text transcript as extracted audio features 308. The audio feature extraction model 132 may include a transformerbased encoder-decoder architecture. The encoder may process a time-frequency representation of the audio and capture the temporal and frequency features of the audio. The decoder may use the encoder’s output to predict text tokens.

[0057] In various implementations, the audio feature extraction model 132 receives input audio media 306, processes the input audio media 306 to extract features that may be used in speech emotion recognition applications to detect heightened valence or arousal in audio segments, and outputs a lower-dimensional representation of the audio as extracted audio features 308. The audio feature extraction model 132 may include a CNN to extract high-level features from time-frequency audio representations, MFCCs, etc. The audio feature extraction model 132 may include an RNN to capture temporal dependencies in the audio media. The audio feature extraction model 132 can include output layers including softmax activation functions for discrete emotion classification, linear’ activation functions for continuous valence and arousal scores, sigmoid activation functions for binary classification tasks (e.g., predicting high vs. low arousal), etc.

[0058] In various implementations, the input audio media 306 may be preprocessed to divide input audio into n segments. In some examples, the audio feature extraction model 132 generates extracted audio features 308 for each of the n segments. Accordingly, the audio feature extraction model 132 may output extracted audio features 308 that represent the n segments of the input audio.

[0059] The text feature extraction model 134 may receive input text media 310 (for example, preprocessed text represented as vectors or matrices), process the text to identify and extract key features, and combine the extracted features to form lower-dimensional embeddings (such as, forexample, vectors or matrices) that retain essential semantic information. The lower-dimensional embeddings may be output as extracted text features 312. The input text media 310 may include (or represent) a user-generated search query aimed at identifying specific moments within the audio and / or visual media that are relevant to the query. The text feature extraction model 134 may include any combination of transformer models, recurrent neural networks (RNNs), long short-term memory networks (LSTMs), contrastive language-image pretraining (CLIP) models, contrastive languageaudio pretraining (CLAP) models, language models, GRUs, etc. Transformer models may capture long-range dependencies and contextual relationships in text data through self-attention mechanisms. RNNs and LSTMs may capture sequential information and temporal dependencies in text data, such as sentence structure and context. For example, transformers can capture the contextual meaning of words within a sentence, while LSTMs can track the sequential flow of ideas across sentences.

[0060] CLIP models may capture the semantic alignment between text media and visual media. Contrastive language- audio pretraining (CLAP) models may include an audio encoder and a text encoder, which may be jointly trained to map audio signals and their corresponding textual descriptions into a common embedding space. CLAP models may capture the semantic alignment between text media and audio media. Language models such as bidirectional encoder representations from transformers (BERT) may capture the bidirectional context of words in a sentence (e.g., for understanding the nuanced meanings between surrounding words). Language models such as generative pre-trained transformer (GPT) models may capture the generative properties of text.

[0061] In various implementations, the missing modality generation model 120 generates one or more missing input media modalities. In some examples where the input media 202 includes input visual media 302 and / or input audio media 306 but does not include input text media 310, the missing modality generation model 120 may include a dense captioning system that generates input text media 310 based on the input visual media 302 and / or the input audio media 306. The generated input text media 310 may include a series of detailed textual descriptions corresponding to moments in the input visual media 302 and / or the input audio media 306. In some examples, the dense captioning system extracts features from the input visual media 302 and / or the input audio media 306 and processes the extracted features using a cross-modal attention mechanism, which finds correlations between the auditory and / or visual features. A sequence generation model (suchas, for example an LSTM network and / or a transformer model) takes the integrated multi-modal features and produces sequences of words that form detailed textual descriptions based on the context of moments in the audio and / or visual media.

[0062] In some examples where the input media 202 includes input visual media 302 and / or input audio media 306 but does not include input text media 310, the missing modality generation model 120 may use the title of the audio and / or visual media (e.g., from metadata) as the input text media 310. In some examples where the input media includes audio media (but is missing corresponding visual media) or includes visual media (but is missing corresponding audio media), the missing modality generation model 120 may use models such as a CLIP model, a CLAP model, an ImageBind model, etc., to learn joint representations across multiple modalities such as audio and visual media. For example, when video media is missing corresponding audio media, the missing modality generation model 120 may use any of the previously described models to generate a corresponding audio embedding and / or a corresponding audio sample.

[0063] In various implementations, the feature extraction model 118 may treat the missing modality as a learnable parameter that needs to be learned during training of the feature extraction model 118. For example, the feature extraction model 118 may be trained to generate extracted features that include estimated embeddings representing the missing modality.

[0064] The multi-modal fusion model 122 may integrate audio, visual, and text features into a single representation by combining the features at the input level. The multi-modal fusion model 122 may capture the relationships between these different types of data, forming a unified representation that encompasses information from all the modalities. The integrated representation may be used by downstream machine learning applications. Referring collectively to FIGS. 1 and 2, in various implementations, the multi-modal fusion model 122 receives extracted features 204 output from the feature extraction model 118 as inputs and combines the multimodal features (such as, for example, the extracted visual features 304, extracted audio features 308, and / or extracted text features 312) at the input level using methods such as concatenation, element- wise product, and / or direct sum, and / or combines the extract audio and / or visual features using a bilinear attention mechanism to produce audio-visual features. The multi-modal fusion model 122 may further combine the extracted audio-visual features with the extracted text features to produce a final multimodal representation.

[0065] As illustrated in FIG. 1, in some examples, the multi-modal fusion model 122 includes an audio-visual fusion model 136 and a cross-attention transformer model 138. FIG. 4 is a block diagram 400 illustrating example data flow of the audio-visual fusion model 136 and the crossattention transformer model 138, according to some embodiments. Referring collectively to FIGS. 1 and 4, in various implementations, the audio-visual fusion model 136 uses bilinear operations to create a combined representation that captures the interactions between the input visual media 302 and the input audio media 306. For example, the audio- visual fusion model 136 may receive the extracted visual features 304 and the extracted audio features 308 output by the feature extraction model 118 as inputs and generate a combined representation 402 that captures the relationships and interactions between the visual and audio features. The combined representation 402 may offer a richer, more comprehensive understanding than the individual features alone.

[0066] In various implementations, the extracted visual features 304 include embeddings or vectors representing video frames and the extracted audio features 308 include embeddings or vectors representing audio data. The audio-visual fusion model 136 may perform bilinear operations that combine visual and audio features and capture interactions between the visual and audio features. For example, if a is an audio vector and v is a visual vector, then the bilinear operation may compute an interaction matrix that captures how elements in the audio vector a interact with elements in the visual vector v. The audio-visual fusion model 136 may output vectors, matrices, and / or tensors that encapsulate information from both the visual and audio components as the combined representation 402. In examples where the extracted visual features 304 represent n segments of input video and the extracted audio features 308 represent n segments of input audio, the combined representation 402 may represent n segments of combined audio- visual features.

[0067] The cross-attention transformer model 138 may receive the combined representation 402 output by the audio-visual fusion model 136 and / or extracted text features 312 output by the text feature extraction model 134 as inputs and output an integrated representation 206 that integrates visual, audio, and / or text information, highlighting the most relevant features according to a cross- modal attention mechanism. For example, the cross-attention transformer model may apply attention mechanisms to the combined representation 402 and extracted text features 312 to find correlations between the audio-visual features and the text features. In various implementations, the crossattention transformer model computes attention scores to determine the relevancy of each audio-visual feature to the text features. The attention scores may indicate how relevant each part of the combined representation is to each part of the text features.

[0068] In the integrated representation 206, the attention scores may be assigned to respective parts of the combined representation as a weight indicating the relative importance of each feature (such as, for example, with respect to the contents of the extracted text features). Thus, in examples where the combined representation 402 represents n segments of combined audio-visual features, the integrated representation 206 may include a weight (e.g., a probability score) assigned to each of the n segments of the combined representation 402, which indicates the relevancy of each feature to a search query represented by the extracted text features 312.

[0069] Referring collectively to FIGS. 1 and 2, the moment prediction model 124 may receive the integrated representation 206 from the multi-modal fusion model 122 as input and output a moment prediction 208. The moment prediction 208 may identify or predict relevant or important moments in the input media 202. In various implementations, the input media 202 includes or represents a search query from a user (e.g., a text query) and the moment prediction 208 identifies moments in the input visual media 302 and / or the input audio media 306 that are relevant to the search query. In some examples, the input media 202 includes or represents metadata (such as any of the previously described metadata) and the moment prediction 208 identifies moments in the input visual media 302 and / or the input audio media 306 that are relevant to the metadata. For example, the metadata may include a title and / or a summary, and the moment prediction 208 identifies moments relevant to the title and / or the summary. The metadata may include user engagement data such as viewing statistics, eye-gaze tracking data, blink rate data, and / or other data quantifying the engagement value of different segments, etc., and the moment prediction 208 identifies moments where user engagement is low or high.

[0070] In various implementations, the moment prediction 208 include indications of the locations of relevant moments in the input visual media 302 and / or the input audio media 306, moment relevance scores, highlight saliency scores, summary relevance scores, etc. Moment relevancy scores may describe how relevant the indicated locations are to the input text media 310 (e.g., to the search query). Highlight saliency scores may denote the “highlight-worthiness” (e.g., the importance or interest level) of the indicated locations. Summary relevance scores may indicate how relevant the indicated locations are to the overall summary or topic of the input visual media 302 and / or the input audio media 306.

[0071] The architecture of the moment prediction model 124 may include localization networks, which can be implemented using transformers, CNNs, RNNs, or other models. These networks may take the multi-modal fused features of the integrated representation 206 as inputs and output bounding boxes indicating the location of relevant moments in the input visual media 302 and / or the input audio media 306. In various implementations, the bounding boxes may identify a center of each moment and / or a width or span of the moment. Thus, the bounding boxes may indicate the segments of input visual media 302 and / or the input audio media 306 that contain relevant features (e.g., features relevant to the search query). In various implementations, the moment prediction model 124 includes a neural network designed to predict whether a timestamp (e.g., of the input visual media 302 and / or the input audio media 306) corresponds to a relevant moment, forming a two-class classification problem. In some examples, the moment prediction model 124 includes highlight saliency and / or summary relevance predictor heads for generating highlight saliency and / or summary relevance scores.

[0072] In some examples, the moment prediction model 124 is trained on annotated datasets, such as, for example, datasets including human-annotated data and / or engagement data. Human- annotated data may include user-generated content from various media platforms and / or include annotations indicating interesting and / or query-relevant segments. Engagement data (such as any of the previously described forms of engagement data) may provide signals quantifying the engagement value of different segments. The training process may employ one or more loss functions to optimize the performance of the moment prediction model 124. Segmentation loss may classify bounding boxes or timestamps as being relevant or not relevant. Boundary prediction loss may classify bounding boxes or timestamps as being bounding a relevant moment or not bounding a relevant moment. Bipartite set matching loss may include proposing bounding boxes and matching the bounding boxes to ground truth bounding boxes using L-l based regression loss. Contrastive loss, including intra-modal (e.g., between similar and dissimilar videos) and inter-modal (e.g., between matched and unmatched query-video pairs), may be used to refine the moment prediction model 124.

[0073] In various implementations, the task structure for specific input video types may be enforced using one or more pre-training tasks. Enforcing task structures for video types may include organizing and guiding the training of the moment prediction model 124 to handle the characteristics and challenges of specific input video types. Examples of pretraining tasks includeautomatic speech recognition (ASR) transcription-based pre -training, dense captioning-based pretraining, cross-modal retrieval-based pre-training, and / or unsupervised summary relevance pretraining. ASR transcription-based pre-training may leverage ASR transcriptions from models such as Whisper (e.g., at about two to about ten second intervals), providing sufficient data for moment localization tasks (e.g., identifying relevant moments within video and / or audio). Dense captioningbased pre-training may use a dense captioning model to generate granular captions (e.g., at about two to about ten second intervals) with corresponding start and end timestamps. Outputs of these pre-training tasks may initialize the final fine-tuning tasks (for example, using annotated training datasets).

[0074] The performance of the moment prediction model 124 may be assessed by measuring the overlap between moment predictions 208 generated by the moment prediction model 124 and annotated training data. For example, the overlap may be quantified using metrics such as intersection over union (loU) and / or mean average precision (mAP) to evaluate the accuracy and precision of the predictions. For example, the moment prediction model 124 may be trained until the metrics meet or exceed a threshold.

[0075] As illustrated in FIG. 1, some examples of the moment prediction model 124 include a transformer encoder model 140, a transformer decoder model 142, a span predictor model 144, and / or a binary classifier model 146. FIG. 5 is a block diagram 500 illustrating example data flow of the transformer encoder model 140, the transformer decoder model 142, the span predictor model 144, and the binary classifier model 146, according to some embodiments. Referring collectively to FIGS. 1 and 5, the transformer encoder model 140 may receive the integrated representation 206 from the cross- attention transformer model 138. The transformer encoder model 140 may capture relationships and dependencies between various features of the integrated representation 206 to generate and output an encoded representation 502. The encoded representation 502 may include a sequence of vectors, matrices, and / or tensors that represent the input features and the relationships between the input features in a transformed space, making the encoded representation 502 suitable for further processing by downstream machine learning models. The architecture of the transformer encoder model 140 may include one or more layers of selfattention mechanisms and / or feed-forward neural networks, allowing the transformer encoder model 140 to encode complex relationships between the multimodal features of the integrated representation 206.

[0076] The transformer decoder model 142 may receive the encoded representation 502 and / or moment queries 504 as inputs. Moment queries 504 may serve as candidate temporal regions (e.g., candidate time segments or intervals) within video and / or audio media that may be likely to contain moments of interest or relevance. Moment queries 504 may guide the transformer decoder model 142 to focus on specific time intervals within the encoded representation 502, improving the efficiency and / or accuracy of the transformer decoder model 142. Moment queries 504 may be manually generated or generated through automated techniques. For example, pre-training tasks such as ASR transcriptions and dense captioning can identify potential moment queries 504.Metadata, such as user engagement data may indicate potentially interesting regions of the input visual media 302 and / or the input audio media 306 and may identify potential moment queries 504. Additionally, moment queries 504 may be generated based on sudden changes, significant sounds, specific visual patterns, etc.

[0077] The transformer decoder model 142 may generate and output a decoded representation 506 that includes detailed information about the moments within the encoded representation 502, as guided by the moment queries 504. The architecture of the transformer decoder model 142 may include one or more self-attention layers, cross-attention layers with encoded representation, and / or feed-forward networks. These architectural features may allow the transformer decoder model 142 to process the encoded features and moment queries effectively, resulting in a decoded representation 506 that encapsulates moment-specific information derived from the integrated input features.

[0078] The transformer decoder model 142 may provide the output decoded representation 506 to the span predictor model 144 and / or the binary classifier model 146 as inputs. The span predictor model 144 may use the decoded representation 506 to identify and predict spans of relevant moments within the input media. The span prediction 508 may include or represent one or more bounding boxes [mc, mw], each identifying a center mcof a predicted moment and / or a width or span mwof the predicted moment. Thus, each bounding box [mc, mw] may bound or define a moment in the input visual media 302 and / or input audio media 306. In various implementations, the transformer decoder model 142 includes an architecture that employs regression techniques to determine the center mcand / or width or span mwof each predicted moment, providing temporal boundaries for relevant moments in the input visual media 302 and / or the input audio media 306.

[0079] The binary classifier model 146 may use the decoded representation 506 to determine whether predicted moments (e.g., indicated by the span prediction 508) are in the foreground or background of the respective scene. A predicted moment being in the foreground may indicate that the moment is a valid prediction while a predicted moment being in the background may indicate that the predicted moment is not a valid prediction. Thus, the binary classifier model 146 may output a validity prediction 510 indicating whether each span prediction 508 is valid or not valid. In various implementations, the binary classifier model 146 may include a binary classifier layer that outputs probabilities of a moment being in the foreground or background. Accordingly, in some examples, the moment prediction model 124 outputs a moment prediction 208 that includes a span prediction 508 identifying relevant moments in the input visual media 302 and / or input audio media 306 and a validity prediction 510 that identifies whether each identified moment is valid or not valid.

[0080] In various implementations, the trim composition application 126 may automatically trim input media 202 based on the moment prediction 208 output by the moment prediction model 124. In some examples, the post edit application further refines the trimmed media output by the trim composition application 126. FIG. 6 is a block diagram 600 illustrating data How between the trim composition application 126 and the post edit application, according to some embodiments. Referring collectively to FIGS. 1 and 6, the trim composition application 126 may receive the input media 202, the moment prediction 208, and / or an indication of trim type 602 as inputs. In examples where the input media 202 includes a prcproccsscd representation of video and / or audio media, the input media 202 also includes the original video and / or audio media used to generate the preprocessed representation, and the trim composition application 126 edits the original video and / or audio media.

[0081] The trim type 602 may include an indication of the type of processing the trim composition application 126 should perform on the input media 202. For example, the trim type 602 may include moment retrieval, highlight generation, and summary. In examples where the trim type 602 indicates that the trim composition application 126 should perform moment retrieval, the trim composition application 126 generates trimmed media 604 including moments related to the search query.

[0082] In implementations where the trim type 602 indicates that the trim composition application 126 should perform highlight generation, the moment prediction 208 may include arelevance score for each moment in the input media 202, highlighting the most relevant moments according to the search query. The trim composition application 126 may output trimmed media 604 that includes the top k relevant moments (as indicated by the moment prediction 208). In various implementations, k may be a user-defined value. Thus, the output trimmed media 604 may include segments trimmed from the input media 202 including the most relevant segments to the search query.

[0083] In examples where the trim type 602 indicates that the trim composition application 126 should perform highlight generation, the moment prediction 208 may include a highlight saliency score. In various implementations, the moment prediction 208 includes both a highlight saliency score and a moment relevance score, particularly if the input text media 310 includes a representation of a search query (further specifying the highlights that should be generated). The trim composition application 126 may output trimmed media 604 that includes the top k relevant moments (as indicated by the moment prediction 208). Thus, the output trimmed media 604 may include segments trimmed from the input media 202 including highlights (e.g., the most engaging and / or interesting segments from the input media 202).

[0084] In implementations where the trim type 602 indicates that the trim composition application 126 should perform summarization, the moment prediction 208 may include a summary relevance score for each segment, indicating its importance to the overall content of the media. The trim composition application 126 may output trimmed media 604 that includes the top k relevant moments (as indicated by the moment prediction 208). Thus, the trimmed media 604 may include segments trimmed from the input media 202 that provide a summary of the input media 202 (e.g., a concise and comprehensive overview capturing essential points within user- specified time constraints).

[0085] The post edit application 128 may receive the trimmed media 604 as an input, perform additional edits and refinements to enhance the trimmed media 604, and output final trimmed media 606. For example, users can make further adjustments to the trimmed media 604 by adding transitions and / or inserting advertisements at strategic points (such as, for example, relevant moments as indicated by the moment prediction 208). In various implementations, the post edit application 128 may allow users to refine the segments selected by the trim compositionapplication 126 to produce the trimmed media 604. For example, the user may modify the trim boundaries and / or add segments not initially selected by the trim composition application 126.

[0086] In various implementations, the post edit application 128 includes a multi-modal dense captioning transformer and uses the multi-modal dense captioning transformer to add captions to selected moment segments in the trimmed media 604, which may allow users to quickly preview different segments. The post edit application 128 may use the multi-modal dense captioning transformer to generate titles, captions, and descriptions for the trimmed media 604, generating various hierarchies of textual information describing the trimmed media 604.

[0087] In some examples, the post edit application 128 suggests and / or inserts advertisement clips at the most engaging parts of the trimmed media 604 (e.g., as indicated by highlight saliency scores from the moment prediction 208). In various implementations, the post edit application 128 allows users to customize transitions, output video trim lengths, and / or saliency score thresholds, improving the overall viewing experience of the trimmed media 604.

[0088] FIG. 7 is a message sequence chart 700 illustrating example interactions between components of the computing system 100, according to some embodiments. In the message sequence chart 700, a user device 104 communicates with the media analysis platform 102 and selects media for processing (at operation 702). In various implementations, the user device 104 selects video media and / or audio media from the media library 114 for processing. In some examples, the user device 104 selects video media and / or audio media from the media library 116 for processing, and the media library 116 transmits the selected media to the media analysis platform 102 (at operation 704). In some examples, the user device 104 transmits video media and / or audio media to the media analysis platform 102 for processing (at operation 706).

[0089] In the message sequence chart 700, the user device 104 transmits a search query to the media analysis platform 102 (at operation 708). As previously described, the search query may include a textual description of moments in the selected video and / or audio media for the media analysis platform 102 to identify. In the message sequence chart 700, the media analysis platform 102 generates moment predictions 208 using the selected video and / or audio files and the search query as input media 202 (for example, as input visual media 302, input audio media 306, and input text media 310, respectively) at operation 710. In the message sequence chart 700, the media analysis platform 102 transmits the generated moment predictions 208 and / or an indication ofthe moment predictions 208 to the user device 104 (at operation 712). Thus, the user device 104 may provide and / or select video and / or audio media to be searched, provide a search query to the media analysis platform 102, and the media analysis platform 102 provides search results identifying moments in the searched media that are relevant to the search query.

[0090] FIG. 8 is a message sequence chart 800 illustrating example interactions between components of the computing system 100, according to some embodiments. In the message sequence chart 800, a user device 104 communicates with the media analysis platform 102 and selects media for processing (at operation 802). In various implementations, the user device 104 selects video media and / or audio media from the media library 114 for processing. In some examples, the user device 104 selects video media and / or audio media from the media library 116 for processing, and the media library 116 transmits the selected media to the media analysis platform 102 (at operation 804). In some examples, the user device 104 transmits video media and / or audio media to the media analysis platform 102 for processing (at operation 806).

[0091] In various implementations, the user device 104 selects a trim type (such as, for example, any of the previously described trim types) and transmits an indication of trim type to the media analysis platform 102 (at operation 808). In some examples, the user device 104 transmits a search query to the media analysis platform 102 (at operation 810). As previously described, the search query may include a textual description of moments in the selected video and / or audio media for the media analysis platform 102 to identify. In the message sequence chart 800, the media analysis platform 102 generates moment predictions 208 using the selected video and / or audio files and / or the search query as input media 202 (for example, as input visual media 302, input audio media 306, and input text media 310, respectively) at operation 812.

[0092] In the message sequence chart 800, the media analysis platform 102 generates trimmed media (such as, for example, trimmed media 604 and / or final trimmed media 606) based on the input media 202, moment prediction 208, and / or trim type 602 at operation 814. In various implementations, the media analysis platform 102 transmits the trimmed media to the user device 104 (at operation 816). Thus, the media analysis platform 102 may automatically generate trimmed media based on moment predictions 208.

[0093] FIG. 9 is a flowchart illustrating an example process 900 for generating a moment prediction, according to some embodiments. In the example process 900, any combination of inputmedia is provided to a feature extraction model, such as the feature extraction model 118 (at block 902). The feature extraction model may extract features based on the input media (for example, according to any combination of the previously described techniques). The input media may include a first modality and a second modality (such as any combination of visual, audio, and text media), and the extracted features may include a lower-dimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality.

[0094] In the example process 900, an integrated representation may be generated based on the extracted features (at block 904). For example, the extracted features may be provided to a multimodal fusion model (such as multi-modal fusion model 122) and the multi-model fusion model may generate the integrated representation based on the extracted features (for example, according to any combination of the previously described techniques). The integrated representation may include attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality.

[0095] In the example process 900, a moment prediction may be generated based on the integrated representation (at block 906). For example, the integrated representation may be provided to a moment prediction model (such as moment prediction model 124) and the moment prediction model may generate the moment prediction based on the integrated representation (for example, according to any combination of the previously described techniques). The moment prediction may identify a moment in the input media.

[0096] Various implementations of the system 100 can provide technical benefits across a range of applications. For example, the system 100 may be used to automatically detect highlights in sports events, live concerts, and vlog content. The input media 202 may include long sequences of videos recorded by content creators covering various events over extended periods. The moment prediction 208 may identify key moments likely to interest viewers. The system 100 may automatically generate short-form content ranging from 10 seconds to 3 minutes and / or mediumform content from 3 to 20 minutes based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 efficiently shrinks long recordings into concise, engaging formats suitable for short format online platforms.

[0097] The system 100 may be used to curate product reviews. The input media 202 may include user-generated content (UGC) product reviews, which are valued for their authenticity anddiversity. The moment prediction 208 may identify the most critical segments, such as final product recommendations. The system 100 may automatically generate streamlined review videos by focusing on these key segments based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 enhances the decision-making process for consumers by providing them with concise and relevant product insights.

[0098] The system 100 may be used to preview user-generated content (UGC). The input media 202 may include long videos produced by content creators. The moment prediction 208 may identify key points that summarize the entire video. The system 100 may automatically generate a quick content preview based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 saves creators time and effort by highlighting the essential points, allowing them to understand the content without manually sifting through the entire footage.

[0099] The system 100 may be used for ad clip-based monetization. The input media 202 may include various types of videos where advertisements can be placed. The moment prediction 208 may identify the most engaging parts of the video based on highlight saliency scores. The system 100 may automatically generate ad placements during these highlighted segments based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 increases viewer engagement and click-through rates by strategically positioning ads during the most attention-grabbing moments.

[0100] The system 100 may be used to generate highlights for podcasts. The input media 202 may include long-form podcast content, which can be either audio-only or have both audio and video. The moment prediction 208 may identify relevant sections for medium-length highlights (about 10 minutes) and short-length highlights (10 seconds to 3 minutes). The system 100 may automatically generate these highlights based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 helps creators extract specific topics and sound bites, making podcast content more accessible and engaging for listeners.

[0101] The system 100 may be used to enhance reaction videos. The input media 202 may include videos where creators react to pre-existing content such as gaming, news, or other videos. The moment prediction 208 may identify moments of heightened emotion and engaging commentary. The system 100 may automatically generate edited reaction videos focusing on thesekey moments based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 helps creators produce captivating reaction videos that resonate with viewers, particularly on platforms that reward such content.

[0102] The system 100 may be used for instructional videos or tutorials. The input media 202 may include educational content such as tutorials, lectures, and demonstrations. The moment prediction 208 may identify segments containing essential instructions and task-relevant content. The system 100 may automatically generate trimmed videos that include all necessary instructions for the tutorial based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 ensures the completeness and correctness of instructional content and facilitates efficient learning by allowing users to access specific topics within long videos.

[0103] The system 100 may be used to create long format professional content, such as movie trailers. The input media 202 may include various video clips relevant to the content being promoted. The moment prediction 208 may identify the most compelling and relevant scenes. The system 100 may automatically generate movie trailers and promotional content based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 helps create engaging trailers that attract and retain viewer interest.

[0104] The system 100 may be used for anomaly detection in surveillance videos. The input media 202 may include long surveillance footage. The moment prediction 208 may identify moments of unusual activity through specific queries (e.g., “person commits [specific crime and / or action]”). The system 100 may automatically generate highlighted segments of interest based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606). Thus, the system 100 improves the efficiency and effectiveness of surveillance operations by detecting and highlighting significant activities.

[0105] The system 100 may be used to create highlights from gaming live streams. The input media 202 may include live stream footage from gaming platforms like YouTube and Twitch. The moment prediction 208 may identify exciting moments based on gameplay dynamics and viewer engagement. The system 100 may automatically generate highlight videos from these segments based on the moment prediction 208 (e.g., as trimmed media 604 and / or final trimmed media 606).Thus, the system 100 ensures the creation of high-quality highlight videos that capture the most thrilling aspects of gaming live streams.ADDITIONAL ENUMERATED EXAMPLES

[0106] The following paragraphs provide examples of systems, methods, and devices implemented in accordance with this specification.

[0107] Example 1. A computer-implemented method comprising: generating, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality, the extracted features including a lower-dimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality; generating, at a multi-modal fusion model, an integrated representation based on the extracted features, the integrated representation including attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lowerdimensional numerical representation of the second modality; and generating, at a moment prediction model, a moment prediction based on the integrated representation, wherein the moment prediction identifies a moment in the input media.

[0108] Example 2. The computer-implemented method of example 1, wherein: the input media includes a third modality; the first modality includes visual media; the second modality includes text media; and the third modality includes audio media.

[0109] Example 3. The computer- implemented method of example 2, wherein the feature extraction model includes: a visual feature extraction model configured to extract visual features from the first modality; a text feature extraction model configured to extract text features from the second modality; and an audio feature extraction model configured to extract audio features from the third modality.

[0110] Example 4. The computer-implemented method of example 3, wherein the multi-modal fusion model includes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

[0111] Example 5. The computer-implemented method of example 4, wherein the combined representation includes interactions between the extracted visual features and the extracted audio features.

[0112] Example 6. The computer- implemented method of example 5, wherein the multi-modal fusion model includes a cross-attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features.

[0113] Example 7. The computer-implemented method of example 6, wherein the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features.

[0114] Example 8. The computer-implemented method of any one of examples 1-7, wherein the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

[0115] Example 9. The computer-implemented method of example 8, wherein the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media.

[0116] Example 10. The computer-implemented method of example 9, wherein the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation.

[0117] Example 11. The computer- implemented method of example 10, wherein the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

[0118] Example 12. The computer-implemented method of example 11, wherein the moment prediction model includes: a span predictor model configured to generate the span prediction based on the decoded representation; and a binary classifier model configured to generate the validity prediction based on the decoded representation.

[0119] Example 13. The computer-implemented method of any one of examples 1-12, further comprising generating, based on the moment prediction, trimmed media including the identified moment in the input media.

[0120] Example 14. The computer- implemented method of any one of examples 1-13, wherein the moment is identified by one or more temporal markers identifying a temporal segment of media.

[0121] Example 15. A non-transitory computer-readable medium comprising executable instractions that, when executed by an electronic processor, causes the electronic processor to: generate, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality, the extracted features including a lowerdimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality, generate, at a multi-modal fusion model, an integrated representation based on the extracted features, the integrated representation including attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality, and generate, at a moment prediction model, a moment prediction based on the integrated representation, wherein the moment prediction identifies a moment in the input media.

[0122] Example 16. The non-transitory computer-readable medium of example 15, wherein: the input media includes a third modality; the first modality includes visual media; the second modality includes text media; and the third modality includes audio media.

[0123] Example 17. The non-transitory computer-readable medium of example 16, wherein the feature extraction model includes: a visual feature extraction model configured to extract visual features from the first modality; a text feature extraction model configured to extract text features from the second modality; and an audio feature extraction model configured to extract audio features from the third modality.

[0124] Example 18. The non-transitory computer-readable medium of example 17, wherein the multi-modal fusion model includes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

[0125] Example 19. The non-transitory computer-readable medium of example 18, wherein the combined representation includes interactions between the extracted visual features and the extracted audio features.

[0126] Example 20. The non-transitory computer-readable medium of example 19, wherein the multi-modal fusion model includes a cross- attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features.

[0127] Example 21 . The non-transitory computer-readable medium of example 20, wherein the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features.

[0128] Example 22. The non-transitory computer-readable medium of any one of examples 15- 21, wherein the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

[0129] Example 23. The non-transitory computer- readable medium of example 22, wherein the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media.

[0130] Example 24. The non-transitory computer-readable medium of example 23, wherein the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation.

[0131] Example 25. The non-transitory computer-readable medium of example 24, wherein the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

[0132] Example 26. The non-transitory computer-readable medium of example 25, wherein the moment prediction model includes: a span predictor model configured to generate the span prediction based on the decoded representation; and a binary classifier model configured to generate the validity prediction based on the decoded representation.

[0133] Example 27. The non-transitory computer-readable medium of any one of examples 25-26, wherein when executed by an electronic processor, the instructions cause the electronicprocessor to generate, based on the moment prediction, trimmed media including the identified moment in the input media.

[0134] Example 28. The non-transitory computer-readable medium of any one of examples 17- 27, wherein the moment is identified by one or more temporal markers identifying a temporal segment of media.

[0135] Example 29. A system comprising: memory hardware storing instructions; and processor hardware configured to execute the instructions, wherein executing the instructions causes the processor hardware to: generate, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality, the extracted features including a lower-dimensional numerical representation of the first modality and a lowerdimensional numerical representation of the second modality, generate, at a multi-modal fusion model, an integrated representation based on the extracted features, the integrated representation including attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality, and generate, at a moment prediction model, a moment prediction based on the integrated representation, wherein the moment prediction identifies a moment in the input media.

[0136] Example 30. The system of example 29, wherein: the input media includes a third modality; the first modality includes visual media; the second modality includes text media; and the third modality includes audio media.

[0137] Example 31. The system of example 30, wherein the feature extraction model includes: a visual feature extraction model configured to extract visual features from the first modality; a text feature extraction model configured to extract text features from the second modality; and an audio feature extraction model configured to extract audio features from the third modality.

[0138] Example 32. The system of example 31, wherein the multi-modal fusion model includes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

[0139] Example 33. The system of example 32, wherein the combined representation includes interactions between the extracted visual features and the extracted audio features.

[0140] Example 34. The system of example 33, wherein the multi-modal fusion model includes a cross-attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features.

[0141] Example 35. The system of example 34, wherein the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features.

[0142] Example 36. The system of any one of examples 29-35, wherein the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

[0143] Example 37. The system of example 36, wherein the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media.

[0144] Example 38. The system of example 37, wherein the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation.

[0145] Example 39. The system of example 38, wherein the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

[0146] Example 40. The system of example 39, wherein the moment prediction model includes: a span predictor model configured to generate the span prediction based on the decoded representation; and a binary classifier model configured to generate the validity prediction based on the decoded representation.

[0147] Example 41. The system of any one of examples 29-40, wherein executing the instructions further causes the processor hardware to generate, based on the moment prediction, trimmed media including the identified moment in the input media.

[0148] Example 42. The system of any one of examples 29-41, wherein the moment is identified by one or more temporal markers identifying a temporal segment of media.

[0149] The foregoing description is merely illustrative in nature and does not limit the scope of the disclosure or its applications. The broad teachings of the disclosure may be implemented in many different ways. While the disclosure includes some particular examples, other modifications will become apparent upon a study of the drawings, the text of this specification, and the following claims. In the written description and the claims, one or more processes within any given method may be executed in a different order — or processes may be executed concurrently or in combination with each other — without altering the principles of this disclosure. Similarly, instructions stored in a non-transitory computer-readable medium may be executed in a different order — or concurrently — without altering the principles of this disclosure. Unless otherwise indicated, the numbering or other labeling of instructions or method steps is done for convenient reference and does not necessarily indicate a fixed sequencing or ordering.

[0150] It should also be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components may be utilized in various implementations. Aspects, features, and instances may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one instance, the electronic based aspects of the invention may be implemented in software (for example, stored on non- transitory computer-readable medium) executable by one or more processors. As a consequence, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components may be utilized to implement the invention. For example, “control units” and “controllers” described in the specification can include one or more electronic processors, one or more memories including a non-transitory computer-readable medium, one or more input / output interfaces, and various connections (for example, a system bus) connecting the components.

[0151] Unless the context of their usage unambiguously indicates otherwise, the articles “a,” “an,” and “the” should not be interpreted to mean “only one.” Rather, these articles should be interpreted to mean “at least one” or “one or more.” Likewise, when the terms “the” or “said” are used to refer to a noun previously introduced by the indefinite article “a” or “an,” the terms “the” or “said” should similarly be interpreted to mean “at least one” or “one or more” unless the context of their usage unambiguously indicates otherwise.

[0152] It should also be understood that although certain drawings illustrate hardware and software located within particular devices, these depictions are for illustrative purposes only. In some embodiments, the illustrated components may be combined or divided into separate software, firmware, and / or hardware. For example, instead of being located within and performed by a single electronic processor, logic and processing may be distributed among multiple electronic processors. Regardless of how they are combined or divided, hardware and software components may be located on the same computing device or may be distributed among different computing devices connected by one or more networks or other suitable connections or links.

[0153] Thus, in the claims, if an apparatus or system is claimed, for example, as including an electronic processor or other element configured in a certain manner, for example, to make multiple determinations, the claim or claim element should be interpreted as meaning one or more electronic processors (or other element) where any one of the one or more electronic processors (or other element) is configured as claimed, for example, to make some or all of the multiple determinations collectively. To reiterate, those electronic processors and processing may be distributed.

[0154] Spatial and functional relationships between elements — such as modules — are described using terms such as (but not limited to) “connected,” “engaged,” “interfaced,” and / or “coupled.” Unless explicitly described as being “direct,” relationships between elements may be direct or include intervening elements. The phrase “at least one of A, B, and C” should be construed to indicate a logical relationship (A OR B OR C), where OR is a non-exclusive logical OR, and should not be construed to mean “at least one of A, at least one of B, and at least one of C.” The term “set” does not necessarily exclude the empty set. For example, the term “set” may have zero elements. The term “subset” does not necessarily require a proper subset. For example, a “subset” of set A may be coextensive with set A, or include elements of set A. Furthermore, the term “subset” does not necessarily exclude the empty set.

[0155] In the figures, the directions of arrows generally demonstrate the flow of information — such as data or instructions. The direction of an arrow does not imply that information is not being transmitted in the reverse direction. For example, when information is sent from a first element to a second element, the arrow may point from the first element to the second element. However, the second element may send requests for data to the first element, and / or acknowledgements of receipt of information to the first element. Furthermore, while the figures illustrate a number of componentsand / or steps, any one or more of the components and / or steps may be omitted or duplicated, as suitable for the application and setting.

[0156] Additionally, operations (such as processes, decisions, inputs, outputs, actions, messages, interactions, events, and / or any other operations) shown in the flowcharts and / or message sequence charts may be illustrated once each and in a particular order in the drawings. However, in various implementations, the operations may be reordered and / or repeated as may be suitable. In some examples, different operations may be performed in parallel, as may be appropriate.

[0157] The term computer-readable medium does not encompass transitory electrical or electromagnetic signals or electromagnetic signals propagating through a medium — such as on an electromagnetic carrier wave. The term “computer-readable medium” is considered tangible and non-transitory. The functional blocks, flowchart elements, and message sequence charts described above serve as software specifications that can be translated into computer programs by the routine work of a skilled technician or programmer.

Claims

CLAIMSWhat is claimed is:

1. A computer- implemented method comprising: generating, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality, the extracted features including a lowerdimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality; generating, at a multi-modal fusion model, an integrated representation based on the extracted features, the integrated representation including attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality; and generating, at a moment prediction model, a moment prediction based on the integrated representation, wherein the moment prediction identifies a moment in the input media.

2. The computer-implemented method of claim 1, wherein: the input media includes a third modality; the first modality includes visual media; the second modality includes text media; and the third modality includes audio media.

3. The computer-implemented method of claim 2, wherein the feature extraction model includes: a visual feature extraction model configured to extract visual features from the first modality; a text feature extraction model configured to extract text features from the second modality; and an audio feature extraction model configured to extract audio features from the third modality.

4. The computer-implemented method of claim 3, wherein the multi-modal fusion model includes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

5. The computer-implemented method of claim 4, wherein the combined representation includes interactions between the extracted visual features and the extracted audio features.

6. The computer-implemented method of claim 5, wherein the multi-modal fusion model includes a cross-attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features.

7. The computer-implemented method of claim 6, wherein the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features.

8. The computer-implemented method of any one of claims 1-7, wherein the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

9. The computer-implemented method of claim 8, wherein the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media.

10. The computer-implemented method of claim 9, wherein the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation.

11. The computer-implemented method of claim 10, wherein the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

12. The computer-implemented method of claim 11, wherein the moment prediction model includes: a span predictor model configured to generate the span prediction based on the decoded representation; and a binary classifier model configured to generate the validity prediction based on the decoded representation.

13. The computer-implemented method of any one of claims 1-12, further comprising generating, based on the moment prediction, trimmed media including the identified moment in the input media.

14. The computer-implemented method of any one of claims 1-13, wherein the moment is identified by one or more temporal markers identifying a temporal segment of media.

15. A non-transitory computer-readable medium comprising executable instructions that, when executed by an electronic processor, causes the electronic processor to perform the method of any one of claims 1-14.

16. A system comprising: memory hardware storing instructions; and processor hardware configured to execute the instructions, wherein executing the instructions causes the processor hardware to: generate, at a feature extraction model, extracted features based on input media, the input media including a first modality and a second modality, the extracted features including a lowerdimensional numerical representation of the first modality and a lower-dimensional numerical representation of the second modality, generate, at a multi-modal fusion model, an integrated representation based on the extracted features, the integrated representation including attention scores capturing correlations between the lower-dimensional numerical representation of the first modality and the lower-dimensional numerical representation of the second modality, and generate, at a moment prediction model, a moment prediction based on the integrated representation, wherein the moment prediction identifies a moment in the input media.

17. The system of claim 16, wherein: the input media includes a third modality; the first modality includes visual media; the second modality includes text media; and the third modality includes audio media.

18. The system of claim 17, wherein the feature extraction model includes: a visual feature extraction model configured to extract visual features from the first modality; a text feature extraction model configured to extract text features from the second modality; and an audio feature extraction model configured to extract audio features from the third modality.

19. The system of claim 18, wherein the multi-modal fusion model includes an audio-visual fusion model configured to generate a combined representation based on the extracted visual features and the extracted audio features.

20. The system of claim 19, wherein the combined representation includes interactions between the extracted visual features and the extracted audio features.

21. The system of claim 20, wherein the multi-modal fusion model includes a cross-attention transformer model configured to generate the integrated representation based on the combined representation and the extracted text features.

22. The system of claim 21, wherein the integrated representation includes a score indicating a relevance of a feature of the combined representation to a feature of the extracted text features.

23. The system of any one of claims 16-22, wherein the moment prediction includes a span prediction identifying a location of the identified moment in the input media.

24. The system of claim 23, wherein the moment prediction includes a validity prediction indicating whether the identified moment pertains to a foreground of the input media or to a background of the input media.

25. The system of claim 24, wherein the moment prediction model includes a transformer encoder model configured to generate an encoded representation capturing relationships and dependencies between features of the integrated representation.

26. The system of claim 25, wherein the moment prediction model includes a transformer decoder model configured to generate a decoded representation of the encoded representation.

27. The system of claim 26, wherein the moment prediction model includes: a span predictor model configured to generate the span prediction based on the decoded representation; and a binary classifier model configured to generate the validity prediction based on the decoded representation.

28. The system of any one of claims 16-27, wherein executing the instructions further causes the processor hardware to generate, based on the moment prediction, trimmed media including the identified moment in the input media.

29. The system of any one of claims 16-28, wherein the moment is identified by one or more temporal markers identifying a temporal segment of media.

Citation Information

Patent Citations

  • Recognizing salient video events through learning-based multimodal analysis of visual features and audio-based analytics

    US20160004911A1