Method, System, and Storage Medium for Generating a Highlights Video from Video and Text Inputs

Through text analysis and neural network model, combined with time anchoring and multi-modal feature extraction, key events in video are automatically identified, solving the problem that generating videos at highlight moments in the existing technology requires manual editing, and achieving efficient and accurate video processing.

CN116170651BActive Publication Date: 2025-07-08BAIDU USA LLC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202210979659.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-23
Filing Date
2022-08-16
Publication Date
2025-07-08
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

The prior art requires a lot of manual editing when generating videos in highlight moments, and it is difficult to accurately identify and extract key events in the video.

Method used

Using computer-implemented methods, through text parsing, time anchoring, neural network model and multi-modal feature extraction, key events in the video are automatically identified and highlighted moment videos are generated.

Benefits of technology

It realizes automatic and precise generation of highlight-time videos, reduces manual editing work, and improves video processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116170651B_ABST
    Figure CN116170651B_ABST
Patent Text Reader

Abstract

The present disclosure provides systems, methods, and data sets for automatically and precisely generating highlight videos or summary videos of content. In one or more embodiments, the input includes text (e.g., an article) of key events (e.g., goals, player movements, etc.) in an event (e.g., a game, a concert, etc.) and one or more videos of the event. In one or more embodiments, the output is a short video of one or more events in the text, where the video may include commentary and / or other audio (e.g., music) of the highlight events, which may also be automatically synthesized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This patent application is a partial continuation - in - part of co - pending and commonly - owned U.S. Patent Application 17 / 393,373, and claims the priority of U.S. Patent Application 17 / 393,373, which was filed on August 3, 2021, titled "AUTOMATICALLY AND PRECISELY GENERATING HIGHLIGHT VIDEOS WITH ARTIFICIAL INTELLIGENCE", and lists Zhiyu Cheng, Le Kang, Xin Zhou, Hao Tian, and Xing Li as inventors (Case No.: 28888 - 2450(BN201118USN1)), which claims the priority of co - pending and commonly - owned U.S. Patent Application 63 / 124,832, which was filed on December 13, 2020, titled "AUTOMATICALLY AND PRECISELY GENERATING HIGHLIGHT VIDEOS WITH ARTIFICIAL INTELLIGENCE", and lists Zhiyu Cheng, Le Kang, Xin Zhou, Hao Tian, and Xing Li as inventors (Case No.: 28888 - 2450P(BN201118USN1 - Provisional)); each of the above - mentioned patent documents is incorporated herein by reference in its entirety for all purposes. Technical Field

[0003] The present disclosure generally relates to systems and methods for machine learning, which can provide improved computer performance, features, and uses. More specifically, the present disclosure relates to systems and methods for automatically generating abstracts or highlights of content. Background Art

[0004] With the rapid development of Internet technology and emerging tools, video content generated online (such as videos of sports-related or other events) is growing rapidly at an unprecedented rate. Especially during the COVID-19 pandemic, since fans were not allowed to attend events in person (such as stadiums or theaters), the viewership of online videos has skyrocketed. Creating highlight videos or other event-related videos usually involves manual work of manually editing the original untrimmed videos. For example, the most popular sports videos usually consist of short clips of a few seconds, and precisely understanding the video and the key events on-site is very challenging for machines. Combining the large amount of original content that exists, breaking down the original content into appropriate highlight videos is very time-consuming and costly. Moreover, for viewers, given the limited time for watching content, it is important to be able to obtain compressed content that appropriately captures the prominent elements or events.

[0005] Therefore, there is a need for a system and method that can automatically and precisely generate refined or compressed video content such as highlight videos. Summary of the Invention

[0006] One aspect of the present disclosure provides a computer-implemented method, including: given input text that mentions an event in an activity, using a text parsing module to parse the input text to identify the event mentioned in the input text; and using a text-to-speech (TTS) module to convert the input text into TTS-generated audio; given at least a portion of the activity and an input video of the identified event: performing time anchoring to associate the running time of the input video with the running time of the activity; by using time information parsed from the input text, from additional sources related to the activity, or from both, and the relevant time obtained through time anchoring, identifying the approximate time when the event occurs during the activity, and then generating an initial clip from the input video that includes the event; extracting features from the initial video clip; using the extracted features and a trained neural network model to obtain the final time value of the event in the initial video clip; in response to the running time of the initial video clip not being consistent with the running time of the TTS-generated audio, generating a final video clip by editing the initial video clip to have a running time consistent with the running time of the TTS-generated audio; and in response to the running time of the initial video clip being consistent with the running time of the TTS-generated audio, using the initial video clip as the final video clip; and combining the TTS-generated audio with the final video clip to generate an event highlight video.

[0007] Another aspect of the present disclosure provides a system, comprising: one or more processors; and a non-transitory computer-readable medium including one or more sets of instructions that, when executed by at least one of the one or more processors, cause the following steps to be performed, the steps including: given input text that mentions an event in an activity, parsing the input text using a text parsing module to identify the event mentioned in the input text; and converting the input text into TTS-generated audio using a text-to-speech (TTS) module; given at least a portion of the activity and input video of the identified event: performing time anchoring to associate the running time of the input video with the running time of the activity; identifying an approximate time during the activity when the event occurs by using time information parsed from the input text, from additional sources related to the activity, or from both, and the related time obtained through time anchoring, and thereby generating an initial clip from the input video including the event; extracting features from the initial video clip; using the extracted features and a trained neural network model to obtain a final time value of the event in the initial video clip; in response to the running time of the initial video clip being inconsistent with the running time of the TTS-generated audio, generating a final video clip by editing the initial video clip to have a running time consistent with the running time of the TTS-generated audio; and in response to the running time of the initial video clip being consistent with the running time of the TTS-generated audio, using the initial video clip as the final video clip; and combining the TTS-generated audio with the final video clip to generate an event highlight video.

[0008] Another aspect of the present disclosure provides a non - transitory computer - readable medium including one or more instruction sequences that, when executed by at least one processor, cause the following steps to be performed. The steps include: given input text that mentions an event in an activity, parsing the input text using a text parsing module to identify the event mentioned in the input text; and converting the input text into TTS - generated audio using a text - to - speech (TTS) module; given an input video of at least a portion of the activity and the identified event: performing time anchoring to associate the running time of the input video with the running time of the activity; identifying an approximate time during the activity when the event occurs by using time information parsed from the input text, from additional sources related to the activity, or from both, and the related time obtained through time anchoring, and then generating an initial clip from the input video including the event; extracting features from the initial video clip; using the extracted features and a trained neural network model to obtain a final time value of the event in the initial video clip; in response to the running time of the initial video clip not being consistent with the running time of the TTS - generated audio, generating a final video clip by editing the initial video clip to have a running time consistent with the running time of the TTS - generated audio; and in response to the running time of the initial video clip being consistent with the running time of the TTS - generated audio, using the initial video clip as the final video clip; and combining the TTS - generated audio with the final video clip to generate an event highlight video. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Embodiments of the present disclosure will be referred to, examples of which may be shown in the drawings. These drawings are illustrative and not restrictive. Although the present disclosure is generally described in the context of these embodiments, it should be understood that it is not intended to limit the scope of the present disclosure to these particular embodiments. Items in the figures may not be drawn to scale.

[0010] Figure 1 An overview of a highlight generation system according to an embodiment of the present disclosure is described;

[0011] Figure 2 An overview method for training a generation system according to an embodiment of the present disclosure is described;

[0012] Figure 3 An overall overview of a dataset generation process according to an embodiment of the present disclosure is described;

[0013] Figure 4 Comments and tags of some cloud - sourced text data according to an embodiment of the present disclosure are summarized;

[0014] Figure 5 Summarizes the untrimmed game videos collected according to embodiments of the present disclosure;

[0015] Figure 6 Shows an embodiment of a user interface designed for event time annotation of videos for humans according to embodiments of the present disclosure;

[0016] Figure 7 Shows a method for associating event time and video running time according to embodiments of the present disclosure;

[0017] Figure 8 Shows an example of identifying timer digits in a game video according to embodiments of the present disclosure;

[0018] Figure 9 Depicts a method for generating clips from an input video according to embodiments of the present disclosure;

[0019] Figure 10 Shows feature extraction according to embodiments of the present disclosure;

[0020] Figure 11 Shows a pipeline for feature extraction according to embodiments of the present disclosure;

[0021] Figure 12 Shows a neural network model that can be used for feature extraction according to embodiments of the present disclosure;

[0022] Figure 13 Describes feature extraction using a slow-fast neural network model according to embodiments of the present disclosure;

[0023] Figure 14 Depicts a method for audio feature extraction and interest event time prediction according to embodiments of the present disclosure;

[0024] Figure 15A Shows an example of the original audio waveform according to embodiments of the present disclosure, and Figure 15B Shows its corresponding average absolute value feature according to embodiments of the present disclosure;

[0025] Figure 16 Shows a method for predicting the time of interest events in a video according to embodiments of the present disclosure;

[0026] Figure 17 Shows a pipeline for time localization according to embodiments of the present disclosure;

[0027] Figure 18 Describes a method for predicting the likelihood of interest events in a video clip according to embodiments of the present disclosure;

[0028] Figure 19 Shows a pipeline for action location prediction according to an embodiment of the present disclosure;

[0029] Figure 20 Describes a method for predicting the likelihood of an interesting event in a video clip according to an embodiment of the present disclosure;

[0030] Figure 21 Shows a pipeline for final time prediction using an integrated neural network model according to an embodiment of the present disclosure;

[0031] Figure 22 Describes the goal location results according to an embodiment of the present disclosure compared to another method;

[0032] Figure 23 Shows the goal location results of 3 clips according to an embodiment of the present disclosure, and the integrated learning achieved the best results;

[0033] Figure 24A and Figure 24B Describes a system for generating a summary or highlight video from video input and text summary input according to an embodiment of the present disclosure;

[0034] Figure 25 Describes a method for extracting information from input text according to an embodiment of the present disclosure;

[0035] Figure 26 Shows a method for generating a player database according to an embodiment of the present disclosure;

[0036] Figure 27 Describes a method for combining video clips and corresponding audio clips to create a summary video according to an embodiment of the present disclosure;

[0037] Figure 28 Describes a simplified block diagram of a computing device / information processing system according to an embodiment of the present disclosure. Detailed Description

[0038] In the following description, for purposes of explanation, specific details are set forth to provide an understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these details. Additionally, those skilled in the art will recognize that the embodiments of the present disclosure described below may be implemented in various ways, such as a process, apparatus, system, device, or method on a tangible computer-readable medium.

[0039] The components or modules shown in the figures are examples of exemplary embodiments of the present disclosure and are intended to avoid obscuring the present disclosure. It should also be understood that throughout the discussion, components may be described as separate functional units that may include sub-units, but those skilled in the art will recognize that various components or portions thereof may be divided into separate components or may be integrated together, including, for example, in a single system or component. It should be noted that the functions or operations discussed herein may be implemented as components. Components may be implemented in software, hardware, or a combination thereof.

[0040] In addition, the connections between components or systems in the figures are not limited to direct connections. Instead, the data between these components may be modified, reformatted, or otherwise changed by intermediate components. In addition, additional or fewer connections may be used. It should also be noted that the terms "coupled", "connected", "communicatively coupled", "joined", "interface", or any of their derivatives should be understood to include direct connections, indirect connections through one or more intermediate devices, and wireless connections. It should also be noted that any communication such as a signal, response, reply, confirmation, message, query, etc. may include one or more information exchanges.

[0041] Reference to "one or more embodiments", "preferred embodiments", "an embodiment", "embodiments", etc. in the specification means that a particular feature, structure, characteristic, or function described in connection with the embodiments is included in at least one embodiment of the present disclosure and may be included in multiple embodiments. In addition, the above phrases that appear in various places in the specification do not necessarily all refer to the same one or more embodiments.

[0042] The use of certain terms throughout this specification is for illustrative purposes and should not be construed as limiting. A service, function, or resource is not limited to a single service, function, or resource; the use of these terms may refer to a grouping of related services, functions, or resources, which may be distributed or aggregated. The terms "include", "including", "comprise", and "comprising" should be understood as open-ended terms, and any subsequent lists are examples and are not meant to be limited to the items listed. A "layer" may include one or more operations. The words "optimal", "optimize", "optimization", etc. refer to an improvement in a result or process and do not require that the specified result or process has reached an "optimal" or peak state. The use of memory, database, information repository, data storage, table, hardware, cache, etc. here may be used to refer to one or more system components into which information may be input or otherwise recorded.

[0043] In one or more embodiments, the stop conditions may include: (1) a set number of iterations have been performed; (2) a processing time amount has been reached; (3) convergence (e.g., the difference between consecutive iterations is less than a first threshold); (4) divergence (e.g., performance degradation); and (5) an acceptable result has been reached.

[0044] Those skilled in the art should recognize that: (1) certain steps may be optionally performed; (2) the steps may not be limited to the specific order described herein; (3) certain steps may be performed in a different order; and (4) certain steps may be performed simultaneously.

[0045] Any headings used herein are for organizational purposes only and should not be used to limit the scope of the specification or claims. Each reference / document mentioned in this patent document is incorporated herein by reference in its entirety.

[0046] It should be noted that any experiments and results provided herein are provided by way of illustration and are performed under specific conditions using one or more specific embodiments; therefore, these experiments and their results should not be used to limit the scope of the disclosure of this patent document.

[0047] It should also be noted that although the embodiments described herein may be in the context of a sports event (such as football), aspects of the present disclosure are not limited thereto. Accordingly, aspects of the present disclosure may be applied to or adapted for use in other environments.

[0048] A. General Introduction

[0049] 1. General Overview

[0050] Embodiments are presented herein for automatically, massively, and precisely generating highlight videos. For illustration purposes, a football game will be used. However, it should be noted that the embodiments herein can be used for or adapted for use in other sports and non-sports events, such as concerts, performances, speeches, demonstrations, news, shows, video games, games, sports events, animations, social media posts, movies, etc. Each of these activities can be referred to as an occurrence event or event, and the highlight moments of the occurrence event can be referred to as interesting events, happenings, or highlight moments.

[0051] Using a large-scale multi-modal dataset, an existing deep learning model is created and trained to detect one or more events (such as goals) in the game—although interesting events (such as penalty kicks, injuries, brawls, red cards, corner kicks, penalty shots, etc.) can also be used. Embodiments of an integrated learning module are also provided herein to improve the performance of interesting event localization.

[0052] Figure 1Describes an overview of a highlight moment generation system according to an embodiment of the present disclosure. In one or more embodiments, large-scale cloud source text data and untrimmed soccer game videos are collected and fed into a series of data processing tools to generate candidate long clips (e.g., 70 seconds, although other time lengths may also be used) that contain the main game interest events (e.g., goal events). In one or more embodiments, a novel interest event localization pipeline precisely locates the moments of the events in the clips. Finally, embodiments can construct one or more customized highlight moment videos / stories around the detected highlight moments.

[0053] Figure 2 Describes an overview method for training a generation system according to an embodiment of the present disclosure. To train the generation system, a large-scale multimodal dataset of event-related data must be generated or obtained to be used as training data (205). Since the video running time may not correspond to the time in the event, in one or more embodiments, for each video in a set of training videos, time anchoring is performed to associate the video running time with the event time (210). Then, metadata (e.g., comments and / or tags) and the associated time obtained through time anchoring can be used to identify the approximate time of the interest events to generate clips including the interest events from the videos (215). By using clips instead of the entire video, the processing requirements are greatly reduced. For each clip, features are extracted (220). In one or more embodiments, a set of pre-trained models can be used to obtain the extracted features, which can be multimodal.

[0054] In one or more embodiments, for each clip, a neural network model is used to obtain the final time value of the interest event (225). In an embodiment, the neural network model can be an integration module that receives features from a set of models and outputs the final time value. Given the predicted final time value for each clip, the predicted final time value is compared with its corresponding true value to obtain a loss value (230); and the loss value can be used to update the model (235).

[0055] Once trained, the generation system can be output and given an input event video, the generation system is used to generate a highlight moment video.

[0056] 2. Related Work

[0057] In recent years, artificial intelligence has been applied to the analysis of video content and the generation of videos. In sports analysis, many computer vision technologies have been developed to understand sports broadcasts. In particular, in football, researchers have proposed algorithms for identifying key match events and player actions, algorithms for analyzing passing feasibility using the body orientation of players, algorithms for detecting events by combining audio and video streams, algorithms for identifying group activities on the field using broadcast streams and trajectory data, algorithms for aggregating deep frame features to locate major match events, and algorithms for leveraging the temporal context information around actions to process the inherent temporal patterns representing these actions.

[0058] For various video understanding tasks, deep neural networks are trained with large-scale datasets. Recent challenges include finding the temporal boundaries of activities or localizing events in the temporal domain. In football video understanding, some people define the goal event as the moment the ball crosses the goal line.

[0059] In one or more embodiments, this definition of a goal is adopted, and existing deep learning models and methods as well as audio stream processing technologies are utilized, and an integrated learning module is adopted in the embodiments to accurately locate events in football video clips.

[0060] 3. Some contributions of the embodiments

[0061] In this patent document, embodiments of an automatic highlight generation system that can accurately identify the occurrence of events in videos are proposed. In one or more embodiments, the system can be used to generate highlight videos in large quantities without conventional manual editing work. Some contributions provided by one or more embodiments include, but are not limited to, the following:

[0062] — Create a large-scale multimodal football dataset including cloud source text data and high-definition videos. Moreover, in one or more embodiments, various data processing mechanisms are applied to parse, clean, and annotate the collected data.

[0063] — Align multimodal data from multiple sources and generate candidate long video clips by cutting the original video into 70-second clips using the parsed tags from cloud source comment data.

[0064] — Embodiments of an event localization pipeline are proposed herein. The embodiments extract high-level feature representations from multiple perspectives and apply temporal localization methods to help localize events in the clips. Additionally, the embodiments are designed with an integrated learning module to improve the performance of event localization. It should be noted that although the occurring event can be a football game and the interesting event can be a goal, the embodiments can be used for or adapted to other occurring events and other interesting events.

[0065] — Experimental results show that for localizing goal events in clips, the tested embodiments achieve a precision close to 1 (0.984) with a 5 - second tolerance, which is better than existing work and establishes a new state of the art. This result helps accurately capture the goal moment and precisely generate highlight videos.

[0066] 4. Patent Document Layout

[0067] This patent document is organized as follows: Section B presents the creation of the dataset and how data is collected and annotated. In Section C, embodiments of the method for constructing an embodiment of the highlight generation system are given, as well as embodiments of how to accurately localize goal events in football video clips using the proposed method. In Section D, the experimental results are outlined and discussed. It should be reiterated that only as an example is a football game used as the overall content and a goal used as an event within that content, and those skilled in the art should recognize that aspects herein can be applied to other content areas (including outside the realm of games) and other events.

[0068] B. Embodiments of Data Processing

[0069] To train and develop the system embodiments, a large - scale multi - modal dataset was created. Figure 3 An overall overview of the dataset generation process according to embodiments of the present disclosure is described. In one or more embodiments, one or more comments and / or tags (305) associated with the video of the event are collected. For example, football game comments and tags (such as corner kick, goal, block, header, etc.) can be crawled from websites or other sources (see, for example, Figure 1 the tags and comments 105 in ) to obtain data. Additionally, videos (305) associated with the metadata (i.e., comments and / or tags) are also collected. For the embodiments herein, high - definition (HD) untrimmed football game videos from various sources are collected. The start time of the game is annotated in the untrimmed original video using Amazon Mechanical Turk (AMT) (315). In one or more embodiments, metadata (such as comment and / or tag information) can be used to help identify the approximate time of the interesting event to generate clips (such as clips of goals) from the video including the interesting event (320). Finally, Amazon Mechanical Turk (AMT) is used to identify the exact time of the interesting event (such as a goal) in the processed video clip. During the process of training embodiments of the goal localization model, the annotated goal time can be used as the ground truth.

[0070] 1. Embodiments of Data Collection

[0071] In one or more embodiments, a crawling sports website obtained more than 1,000,000 comments and tags, covering more than 10,000 football matches from various leagues from the 2015 to 2020 seasons. Figure 4 Summarizes comments and tags in some cloud source text data according to embodiments of the present disclosure.

[0072] Comments and tags provide a wealth of information for each match. For example, they include the match date, team names, league, match event times (e.g., in minutes), event tags (such as goals, shots, corner kicks, substitutions, fouls, etc.), and the names of the associated players. These comments and tags from the cloud source data can be translated or recognized as rich metadata for the original video processing embodiments as well as the highlight video generation embodiments.

[0073] More than 2,600 high-definition (720P or above) untrimmed football match videos were also collected from various online sources. The matches are from various leagues from 2014 to 2020. Figure 5 Summarizes the collected untrimmed match videos according to embodiments of the present disclosure.

[0074] 2. Data annotation embodiments

[0075] In one or more embodiments, the untrimmed original video is first sent to Amazon Mechanical Turk (AMT) workers to annotate the match start time (defined as the moment the referee blows the whistle to start the match), and then the cloud source match comments and tags are parsed to obtain the goal times in minutes for each match. By combining the goal minute tags and the match start time in the video, candidate 70-second clips containing goal events are generated. Next, in one or more embodiments, these candidate clips are sent to AMT for annotating the goal times in seconds. Figure 6 Illustrates a user interface embodiment designed for AMT for goal time annotation according to embodiments of the present disclosure.

[0076] For goal time annotation on AMT, each HIT (Human Intelligence Task, a worker task) contains one (1) candidate clip. Each HIT is assigned to five (5) AMT workers, and the median timestamp value is collected as the ground truth label.

[0077] C. Method embodiments

[0078] In this section, details of embodiments of each of the five modules of the highlight generation system are given. As a brief overview, the first module embodiment in Section C.1 is a match time anchoring embodiment that checks the time integrity of the video and maps any time in the match to the time in the video.

[0079] The second module embodiment in Section C.2 is a coarse interval extraction embodiment. This module is the main difference relative to the event localization pipeline typically studied. In the embodiment of this module, intervals of 70 seconds are extracted (although intervals of other sizes can be used), where specific events are located by leveraging text metadata. Compared to a normal end-to-end visual event localization pipeline, this method has advantages for at least three reasons. First, the clips extracted with metadata contain more context information and can be used in different dimensions. With metadata, the clips can be used as time clips (such as highlight videos of a game) or can be used together with other clips of the same team or player to generate team, player, and / or season highlight videos. The second reason is robustness, which stems from the low event ambiguity of text data. Third, by analyzing shorter clips of the events of interest rather than the entire video, many resources (processing, processing time, memory, energy consumption, etc.) are saved.

[0080] The embodiment of the third module in the system embodiment is multi-modal feature extraction. Video features are extracted from multiple perspectives.

[0081] The embodiment of the fourth module is precise time localization. Extensive studies of the techniques for how to design and implement the feature extraction and time localization embodiments are provided in Sections C.3 and C.4 respectively.

[0082] Finally, the embodiment of the integrated learning module is described in Section C.5.

[0083] 1. Game time anchoring embodiment

[0084] The event clock in an event video is sometimes irregular. The main reason seems to be that at least some of the event video files collected from the Internet contain corrupted timestamps or frames. It has been observed that in video collection, approximately 10% of the video files contain time corruption that offsets a portion of the video in time, sometimes by more than 10 seconds. Some of the severe corruptions observed include missing frames of over 100 seconds. In addition to errors in the video files, some unexpected rare events may have occurred during the event / incident, and the event clock had to stop for several minutes before they resumed. If it is video content corruption or a game interruption, the time irregularity can be regarded as a time jump forward or backward. To precisely locate the clips of any event specified by the metadata, in one or more embodiments, time jumps are detected and calibrated accordingly. Thus, in one or more embodiments, an anchoring mechanism is designed and used.

[0085] Figure 7A method for event time and video running time association according to an embodiment of the present disclosure is shown. In one or more embodiments, OCR (Optical Character Recognition) is performed on video frames at intervals of 5 seconds (although other intervals can also be used) to read the game clock (705) displayed in the video. The game start time (710) in the video can be derived from the recognized game clock. Whenever a time jump occurs, in one or more embodiments, the record of the game time after the time jump is maintained and is called a time anchor (710). Using the time anchor, in one or more embodiments, any time in the game can be mapped to the time in the video (i.e., the video running time) (715), and any clip specified by the metadata can be accurately extracted. Figure 8 An example of identifying timer digits in a game video according to an embodiment of the present disclosure is shown.

[0086] As Figure 8 shown, the timer digits 805 - 820 can be identified and associated with the video running time. Embodiments can collect multiple recognition results over time and can self - correct based on spatial smoothness and temporal continuity.

[0087] 2. Coarse interval extraction embodiments

[0088] Figure 9 A method for generating clips from an input video according to an embodiment of the present disclosure is depicted. In one or more embodiments, metadata (905) from cloud - sourced game commentary and tags is parsed, and the metadata includes time stamps in minutes for goal events. Combining the game start time detected by an embodiment of the OCR tool (discussed above), the original video can be edited to generate x - second (e.g., 70 - second) candidate clips containing events of interest. In one or more embodiments, the extraction rule can be described by the following equations:

[0089] t {clipStart} = t (gameStart} + 60 * t (goalMinute} - tolerance(1)

[0090] t (clipEnd} = t {clipStart} +(base clip length + 2 * tolerance)(2)

[0091] In one or more embodiments, given the goal minute t {goalMinute} and the game start time t {gameStart) , from the video at t {clipStart}Second extraction clip. In one or more embodiments, the duration of a candidate clip can be set to 70 seconds (where the base clip length is 60 seconds and the tolerance is 5 seconds, although it should be noted that different values and different formulas can be used), as this covers extreme cases when the interesting event occurs very close to the goal minute, and it can also tolerate small deviations in the game start time detected by OCR. In the next section, method embodiments for locating the goal second (the moment the ball crosses the goal line) in the candidate clip are given.

[0092] 3. Multi-modal feature extraction embodiments

[0093] In this section, three embodiments for obtaining high-level feature representations from candidate clips are disclosed.

[0094] a) Feature extraction embodiment using a pre-trained model

[0095] Figure 10 Feature extraction according to an embodiment of the present disclosure is shown. Given video data, in one or more embodiments, time frames are extracted (1005), and if necessary to match the input size, the size of the time frames is adjusted in the spatial domain (1010) and fed into a deep neural network model to obtain a high-level feature representation. In one or more embodiments, a ResNet-152 model pre-trained on an image dataset is used, but other networks can also be used. In one or more embodiments, time frames are extracted at the native frames per second (fps) of the original video and then downsampled to 2 fps, i.e., a ResNet-152 feature representation of 2 frames per second of the original video is obtained. ResNet is a very deep neural network that outputs a feature representation of 2048 dimensions per frame with a fully connected 1000-layer. In one or more embodiments, the output of the layer before the softmax layer can be used as the extracted high-level feature. Note that ResNet-152 can be used to extract high-level features from a single image; it does not inherently embed temporal context information. Figure 11 A pipeline 1100 for extracting high-level features according to an embodiment of the present disclosure is shown.

[0096] b) SlowFast feature extractor embodiment

[0097] As part of a video feature extractor, in one or more embodiments, a slowfast network architecture such as that proposed by Feichtenhofer et al. (Feichtenhofer, C., Fan, H., Malik, J., & He, K., “Slowfast Networks for Video Recognition”, Proceedings of The IEEE International Conference on Computer Vision (pp. 6202 - 6211) (2019), the entire content of which is incorporated herein by reference) or that proposed by Xiao et al. (Xiao et al., “Audiovisual SlowFast Networks for Video Recognition”, arxiv.org / abs / 2001.08740v1 (2020), which is incorporated herein by reference in its entirety) may be used; although it should be noted that other network architectures may be used. Figure 12 Graphically depicts a neural network model that can be used to extract features according to an embodiment of the present disclosure.

[0098] Figure 13 Describes feature extraction using a slowfast neural network model according to an embodiment of the present disclosure. In one or more embodiments, the slowfast network (1305) is initialized with pre - trained weights using a training data set. The network can be fine - tuned as a classifier (1310). The second column in Table 1 below shows the event classification results using a benchmark network with a test data set. In one or more embodiments, the feature extractor is used to classify 4 - second clips into 4 categories: 1) away from the event of interest (e.g., a goal), 2) just before the event of interest, 3) the event of interest, and 4) just after the event of interest.

[0099] Several techniques can be implemented to find the best classifier, which is evaluated by the top - 1 error percentage. First, apply as Figure 12The network constructed therein adds audio as an additional path to the SlowFast network with Audio and Video (AVSlowfast). The visual part of the network can be initialized with the same weights. It can be seen that the direct joint training of visual and audio features harms performance. This is a common problem found when training multi-modal networks. In one or more embodiments, a technique of adding different loss functions for visual and audio modalities respectively is applied, and the entire network is trained with a multi-task loss. In one or more embodiments, the cross-entropy loss on the audiovisual results and a linear combination of each audiovisual branch can be used. The linear combination can be a weighted combination, where the weights can be learned or can be selected as hyperparameters. The best top-1 error result shown in the bottom row of Table 1 is obtained.

[0100] Table 1. Results of Event Classification

[0101] Algorithm Error % of the first 1 SlowFast 33.27 Audio only 60.01 AVSlowfast 40.84 AVSlowfast multitasking 31.82

[0102] In one or more embodiments of the goal localization pipeline, the feature extractor part of the network (AVSlowfast with multi-task loss) can be utilized. Thus, the goal is to reduce the top-1 error, which corresponds to stronger features.

[0103] c) Average Absolute Value Audio Feature Embodiment

[0104] By listening to the sound track of an event (e.g., a game without live commentary), one can usually simply determine when an interesting event occurs based on the volume of the audience. Inspired by this observation, a simple method for directly extracting key information about interesting events from audio is developed.

[0105] Figure 14 A method for audio feature extraction and interesting event time prediction according to an embodiment of the present disclosure is depicted. In one or more embodiments, the absolute value of the audio waveform is taken and downsampled to 1 Hertz (Hz) (1405). This feature representation can be referred to as the average absolute value feature because it represents the average sound amplitude per second. Figure 15A and Figure 15B respectively show examples of the original audio waveform of a clip and its average absolute value feature according to an embodiment of the present disclosure.

[0106] For each clip, the maximum value 1505 (1410) of the average absolute value audio feature 1500B can be located. By locating the maximum value (e.g., maximum value 1505) of the average absolute value audio feature and its corresponding time (e.g., time 1510) for the clips in the test dataset, a precision of 79% is achieved for time localization (with a 5-second tolerance).

[0107] In one or more embodiments, an average absolute value audio feature (e.g., Figure 15B 1500B in Figure 15B ) can be used as a prediction of the likelihood of an event of interest in time within a clip. As will be discussed below, this average absolute value audio feature can be a feature input into an integration model that predicts the final time at which an event of interest occurs within the clip.

[0108] 4. Action Localization Embodiments

[0109] In one or more embodiments, to precisely locate the moment of a goal in a soccer game video, the time context information around that moment is combined to learn what is happening in the video. For example, before a goal event occurs, a player will shoot (or head the ball), and the ball will move towards the goal. In some cases, offensive and defensive players gather in the penalty area and not far from the goal. After a goal event, the goal scorer usually runs to the sideline, hugs teammates, and there are also celebrations among the audience and the coach. Intuitively, these patterns in the video can help the model learn what is happening and locate the moment of the goal event.

[0110] Figure 16 A method for predicting the likelihood of an event of interest in a video clip according to an embodiment of the present disclosure is described. In one or more embodiments, to construct a time localization model, a temporal convolutional neural network (1605) that uses the extracted visual features as input is used. In one or more embodiments, the input features can be features extracted from one or more of the existing models discussed above. For each frame, it outputs a set of intermediate features that mix the temporal information across frames. Then, in one or more embodiments, the intermediate features are input into a segmentation module (1610) that generates segmentation scores, which are evaluated by a segmentation loss function. A cross-entropy loss function can be used for the segmentation loss function:

[0111]

[0112] where t i is the ground truth label and p i is the softmax probability of the i-th classification.

[0113] In one or more embodiments, the segmentation scores and the intermediate features are concatenated and fed into an action localization module (1615), which generates localization predictions (e.g., predictions of the likelihood of an event of interest occurring at each time point within the range of the clip) (1620), and the predictions can be evaluated by an action localization loss function similar to YOLO. An L2 loss function can be used for the action localization loss function:

[0114]

[0115] Figure 17 A pipeline for time localization according to an embodiment of the present disclosure is shown. In one or more embodiments, the temporal CNN may include convolutional layers, the segmentation module may include convolutional layers and batch normalization layers, and the action localization model may include pooling layers and convolutional layers.

[0116] In one or more embodiments, considering the context information of time, the model embodiments are trained with segmentation and action localization loss functions as described by Cioppa et al. (A., Deliège, A., Giancola, S., Ghanem, B., Droogenbroeck, M. V., Gade, R., & Moeslund, T., “A Context-Aware Loss Function for Action Spotting in Soccer Videos”, 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13123-13133, the entire content of which is incorporated herein by reference). In one or more embodiments, the segmentation loss is used to train the segmentation module, where each frame is associated with a score indicating the likelihood that the frame belongs to an action classification, while the action localization loss is used to train the action localization module, where the temporal position of the predicted action classification is predicted.

[0117] At least one major difference between the embodiments herein and the method of Cioppa et al. is that the embodiments herein process short clips, while Cioppa et al. use the entire game video as input, so when implemented in real time, it takes much longer to process the video and extract features.

[0118] In one or more embodiments, the extracted feature input may be features extracted from the ResNet model discussed above or the AVSlowFast multi-task model discussed above. Alternatively, for the AVSlowFast multi-task model, the segmentation part of the action localization model may be removed. Figure 18A method for predicting the likelihood of an interesting event in a video clip according to an embodiment of the present disclosure is described. In one or more embodiments, a temporal convolutional neural network receives features extracted from an AVSlowfast multi-task model as input (1805). For each frame, it outputs a set of intermediate features that mix temporal information across frames. Then, in one or more embodiments, the intermediate features are input into an action localization module (1810), which generates a localization prediction (e.g., a prediction of the likelihood of an interesting event occurring at each time point within the range of the clip) (1815), and this prediction can be evaluated by an action localization loss function. Figure 19 A pipeline for action localization prediction according to an embodiment of the present disclosure is shown.

[0119] 5. Integrated learning embodiments

[0120] In one or more embodiments, a single predicted time of an interesting event in the clip can be obtained from each of the three models described above (e.g., selecting the maximum value). One of the predictions can be used, or the predictions can be combined (e.g., averaged). Alternatively, an integrated model can be used to combine information from each model in order to obtain a final prediction of the interesting event in the clip.

[0121] Figure 20 A method for predicting the likelihood of an interesting event in a video clip according to an embodiment of the present disclosure is described, and Figure 21 A pipeline for final time prediction according to an embodiment of the present disclosure is shown. In one or more embodiments, the final accuracy can be enhanced in an integrated manner that pools the outputs of the three models / features described in the above subsections. In one or more embodiments, the outputs of all three previous models can be combined with a positional encoding vector as the input to an integration module (2005). The combination can be done using concatenation. For example, 4 d-dimensional vectors become a 4×d matrix. For the ResNet and AVSlowfast multi-task models, the input can be the likelihood prediction output from their action localization models in Section 4 above. Also, for audio, the input can be the average absolute value audio feature of the clip (e.g., Figure 15B ). In one or more embodiments, the positional encoding vector is a 1-D vector representing the temporal length (i.e., index) of the clip.

[0122] In one or more embodiments, the core of the integration module is an 18-layer 1-D ResNet network with a regression head. Essentially, the integration module learns a mapping from multi-dimensional input features, including multi-modalities, to the final temporal location of the event of interest in the clip. In one or more embodiments, a final temporal value prediction (2010) is output from the integration model and can be compared with the ground truth time to calculate a loss. The losses for various clips can be used to update the parameters of the integration model.

[0123] 6. Inference Embodiment

[0124] Once trained, the entire highlight generation system as shown in Figure 1 can be deployed. In one or more embodiments, the system may also include an input that allows the user to select one or more parameters regarding the generated clip. For example, the user can select a specific player, a range of games, one or more events of interest (e.g., goals and penalties), and the number of clips for making the highlight video (or the length of time for each clip and / or the entire highlight-edited video). Then, the highlight generation system can access the video and metadata and produce a highlight-edited video by concatenating clips. For example, the user may want 10 seconds for each clip of an event of interest. Thus, in one or more embodiments, a customized highlight video generation module can select 8 seconds before and 2 seconds after the event of interest based on the final predicted time for the clip. Or, as shown in Figure 1 the key events in a player's career can be events of interest and they can be automatically identified and compiled into a "story" of the player's career. Audio and other multimedia features that can be selected by the user can be added to the video by the customized highlight video generation module. Those skilled in the art will recognize other applications of the highlight generation system.

[0125] D. Experimental Results

[0126] It should be noted that these experiments and results are provided by way of illustration and are conducted under specific conditions using one or more specific embodiments; therefore, neither these experiments nor their results should be used to limit the scope of the disclosure of this patent document.

[0127] 1. Goal Localization

[0128] For a fair comparison with existing work, the test model embodiment is trained with candidate clips containing goals extracted from games in the training set of the dataset, and validated / tested with candidate clips containing goals extracted from games in the validation / test set of the dataset.

[0129] Figure 22Shows the main results: For localizing goals in a 70-second clip, Example 2205 of the test significantly outperforms the prior art method 2210, which is known as the Context-Aware method for localizing goals in soccer.

[0130] Also shows the intermediate prediction results obtained by using three different features described in Section C.3 or C.4, and the final result is predicted by the integration learning module described in Section C.5. Figure 23 The goal localization results for 3 clips are stacked. As Figure 23 shown, the final prediction output of the integration learning module embodiment is the best in terms of its proximity to the ground truth label (shown by the dashed ellipse).

[0131] 2. Partial Notes

[0132] As Figure 22 shown, the embodiment can achieve an accuracy close to 1 (0.984) with a 5-second tolerance. This result is phenomenal as it can be used to correct mislabeling from text and synchronize with customized audio commentary. It also helps to precisely generate highlights and thus gives the user / editor the option to customize their video around the exact goal moment. The pipeline embodiment can be naturally extended to capture the moments of other events (such as corner kicks, free kicks, and penalty kicks).

[0133] Again, using a soccer game as the overall content and using goals as the events in that content is merely exemplary, and those skilled in the art will recognize that aspects herein can be applied to other content domains (including those outside the domain of sports) and other events.

[0134] E. Alternative Embodiments for Generating Highlight Videos from Text and Video Inputs

[0135] As described above, the application of the systems and methods disclosed above is the ability to generate highlight videos. For example, it would be very beneficial to be able to automatically generate event highlight videos such as sports highlight videos with commentary. As mentioned above, the demand for video content is increasing, and generating video content takes longer than generating article-based content. Previously, the generation of videos, for example, was a manual process where video editors spent a lot of time. The embodiments herein make it easier to generate highlight videos, where input text, such as a matching summary article, is used when generating the corresponding video content. In one or more embodiments, artificial intelligence / machine learning is used to find, match, or generate video clips corresponding to the text, where the text is related to parts of the video; thus, once the text article is written, the video can be generated, where the process of generating such highlight videos is simplified to writing the article and using an embodiment of an automated system to edit the original video to generate the highlight video.

[0136] Figure 24A and Figure 24B A system for generating highlight videos of one or more events from video and text in accordance with embodiments of the present disclosure is described. As shown, the system 2400 may receive, as input, an article or text segment 2408 describing a game and one or more videos 2402 of the entire game. Note that, for illustrative purposes, the event is a sports game, but it should be noted that other events (e.g., speeches, rallies, news broadcasts, concerts, etc.) may also be used.

[0137] Returning Figure 24A , one task of the disclosed system is to identify one or more correct portions in the video 2402 that correspond to the elements indicated in the input text 2408. Figure 25 A method for extracting information from input text in accordance with embodiments of the present disclosure is described. As shown, an input text (e.g., Figure 24A the text 2408 in) including text related to one or more highlight events in an event is received (2505). For example, the input text may be an article outlining a game, and the article will form the basis of the final summary video output by the system. To assist in video segment selection, the input text is parsed to identify interesting events and related data (if any) (2510). In one or more embodiments, the parsing may be rule-based (e.g., pattern matching, keyword matching, etc.), may employ a machine learning model (e.g., a neural network model trained to extract), or utilize both to extract and classify key data, such as: the number of events, the minutes / time at which the events occur, the type of event / action, players, and other identifying information (who, what, where, when, etc.). Some examples of template matching are provided below:

[0138] - A corner kick by the Arsenal team, the ball was lost by Adam Weber: It contains the keyword "corner kick" and is classified as a corner kick action, and the extracted player is Adam Weber.

[0139] - Paul Grom (Brighton team) won a free kick in the defensive half: This is classified as a free kick action, and the player is Paul Grom.

[0140] - Goal, Arsenal team 1 point, Brighton team 0 point. Nicolas Pem (Arsenal team) shot with his right foot from the center of the penalty area to the lower right corner, and Clark Kamers assisted him with a cross pass: This contains the keywords "goal" and "shot", so this will be classified as both a goal and a shot event. The player is Nicolas Pem.

[0141] In one or more embodiments, the information extracted including key events and related data (if any) is provided to a data set (e.g., Figure 24A data set 2424 in). This data set is used by a video model (e.g., Figure 24A video deep learning model 2428 in) to generate video clips for a video summary.

[0142] As Figure 24A shown, the data can be combined with text data extracted from other sources. For example, text data 2404 can be collected from online publications, social media (e.g., Facebook, Twitter, etc.), high-frequency or live report texts, news feeds, forums, user groups, etc. In one or more embodiments, the same or similar rule-based models or neural network models can be used to parse and classify this text data, or different rule-based models or neural network models closer to the input text data can be used to parse and classify this text data. Generally, the web content or user-generated content of high-frequency (live) text data describing game actions includes time information, which can also be extracted and used to help identify the correct time in the video. This additional text data 2404 can be parsed and classified into a set of sentence groups, and each sentence contains minutes, actions, and possibly other data (such as player data, time, etc.). Each sentence or sentence group describes an event during the game. Therefore, the text parsing / action classification models 2412 and 2416 can be the same or different models. Note that the source of the information 2404 provides additional information to help extract the correct events from the input summary text 2408

[0143] Note that additional data inputs can also be used to assist in parsing and / or video summary generation. As an example, consider that system 2400 can use a player data database 2406 (where the system can web crawl player information or use user-generated player information content). A parser 2414 can also be used to parse the information, and the parser 2414 can use the same or a similar rule-based model or neural network model, or can use a different rule-based model or neural network model that is closer to the input text data. For example, consider Figure 26 the method shown.

[0144] In one or more embodiments, the parser module 2414 can normalize table information and store entity information (such as teams, players, actors, bands, organizations, etc.). For example, a player list of a team (which can be obtained at the team website) can include a player list (2605) by player number and name. The parser can extract rows and obtain player names and jersey numbers (2610), thus obtaining a database with the numbers and names of the players of that team. The output is a player data set 2418, which can be used to help supplement the parsed data (i.e., text data 2420 with seconds classified by action and having possible player data and data collection (such as minutes, actions, and possibly other relevant data) 2422). It should be noted that similar actions can also be used for other entities (such as performers, actors, hosts, etc.).

[0145] Returning to Figure 24A , given the input video 2402, the time detection module 2410 correlates the running time of the video with the game time. In one or more embodiments, time detection can be performed as discussed above regarding time anchoring, but other methods can also be used.

[0146] As Figure 24A shown in the embodiments in, system 2400 can include a minute and action matching module 2426, which can also help identify the time information of events. For example, for each group of sentences describing an event, module 2426 can match the sentences with high-frequency text data to obtain the seconds when the action occurs. In one or more embodiments, the model can also associate additional relevant data, such as players. Information from other data sets (i.e., data set 2420 and data set 2422) can be provided to module 2426, and module 2426 uses this information for matching.

[0147] For example, in the matching summary, in the 4th minute, Arsenal got a corner kick, and the real-time stream data includes the following:

[0148] 3'30” Foul by Bukayo

[0149] 4'5” Adam Weber passes the ball

[0150] 4'12” Nicolas Pem scores a goal

[0151] 4'25” Pascal is fouled by Adam

[0152] 4'37” Corner kick for Arsenal, the ball is lost by Adam Weber

[0153] 5'10” Arsenal substitutes, Arsenal.Sam Bukayo replaces Emil Bowe

[0154] In this case, the interesting match summary data appears at the 4th minute, and the action is a corner kick. In one or more embodiments, the match model embodiment filters the live data and retains all the 4th minute data, that is:

[0155] 4'5” Adam Weber passes the ball

[0156] 4'12” Nicolas Pem scores a goal

[0157] 4'25” Pascal is fouled by Adam

[0158] 4'37” Corner kick for Arsenal, the ball is lost by Adam Weber

[0159] Then, through the above text parsing module, these actions are known: passing the ball, scoring a goal, losing the ball, and corner kick. Therefore, the module 2426 matches to the "4'37” corner kick for Arsenal, the ball is lost by Adam Weber", if this is an interesting event. In one or more embodiments, this information can also be provided to the action module 2430, and the action module 2430 uses the time information to generate the corresponding video clip.

[0160] As shown in the embodiment Figure 24A shown, the assembled data set 2424 includes the approximate minute time when one or more events occur and the classification of these events. In the shown embodiment, the data set 2424 can be an assembly of information from the time detection module 2410, text data information 2420, and data collection information 2422. Then, this information can be used by the video deep learning model 2428 to determine a more accurate time in the video in order to generate the corresponding video clip containing the interesting event. The model 2428 can be one or more of the models discussed above or an integration of one or more of the models discussed above.

[0161] For example, in one or more embodiments, for each group of sentences describing an event, the sentences may include the minute in which the action occurs. The system may extract one minute (or more) of video from the input video and, in one or more embodiments, use one or more deep learning video understanding models to identify the second in which the event / action occurs. As described above, the model may be an integration of one or more of the models discussed above or all or a subset of one or more of the models discussed above.

[0162] In one or more embodiments, the output of module 2428 is a set 2430 of the identified actions / events and the times (in seconds) at which those events / actions occur in the video. This information may be combined with the information from matching module 2426 (and redundant events and times may be removed). This final set 2430 of time and event information may be used to generate a video clip 2434 that includes the action / event of interest ( Figure 24B ). The clip may be a set amount of time (e.g., from x seconds before the event to y seconds after the event, and the clip time span may vary according to the type of action or the length of the event detected by video deep learning model 2428) and / or may have a length corresponding to the length of the audio from the text-to-speech module for the corresponding text that is related to the event in the video clip.

[0163] In one or more embodiments, the input text 2408 or a group of one or more sentences describing an event / action may be input into a text-to-speech (TTS) system 2432, and the TTS system 2432 converts the input text into audio. In one or more embodiments, the text may be an assembly of audio segments generated by converting sentences to audio using TTS.

[0164] Those skilled in the art should recognize that any of many TTS systems may be used. For example, several works address the problem of synthesizing speech from a given input text via neural networks, including but not limited to:

[0165] —DeepSpeech 1 (which is disclosed in U.S. Patent Application 15 / 882,926 (Docket No. 28888-2105), commonly assigned, titled "SYSTEMS AND METHODS FOR REAL-TIME NEURAL TEXT-TO-SPEECH", filed on January 29, 2018, and U.S. Provisional Patent Application 62 / 463,482 (Docket No. 28888-2105P), titled "SYSTEMS AND METHODS FOR REAL-TIME NEURAL TEXT-TO-SPEECH", filed on February 24, 2017. Each of the above patent documents is incorporated herein by reference in its entirety. For convenience, these disclosures may be referred to as "DeepSpeech 1" or "DV1");

[0166] —DeepSpeech 2 (which is disclosed in U.S. Patent Application 15 / 974,397 (Docket No. 8888-2144), commonly assigned, titled "SYSTEMS AND METHODS FOR MULTI-SPEAKER NEURAL TEXT-TO-SPEECH", filed on May 8, 2018, and U.S. Provisional Patent Application 62 / 508,579 (Docket No. 28888-2144P), titled "SYSTEMS AND METHODS FOR MULTI-SPEAKER NEURAL TEXT-TO-SPEECH", filed on May 19, 2017. Each of the above patent documents is incorporated herein by reference in its entirety. For convenience, these disclosures may be referred to as "DeepSpeech 2" or "DV2");

[0167] —DeepSpeech 3 (which is disclosed in U.S. Patent Application 16 / 058,265 (Docket No. 28888-2175), commonly assigned, titled "SYSTEMS AND METHODS FOR NEURAL TEXT-TO-SPEECH USING CONVOLUTIONAL SEQUENCE LEARNING", filed on August 8, 2018, and U.S. Provisional Patent Application 62 / 574,382 (Docket No. 2888-2175P), titled "SYSTEMS AND METHODS FOR NEURAL TEXT-TO-SPEECH USING CONVOLUTIONAL SEQUENCE LEARNING", filed on October 19, 2017, and lists Sercan as the inventor Ar1k, Wei Ping, Kainan Peng, Sharan Narang, Ajay Kannan, Andrew Gibiansky, Jonathan Raiman, and John Miller (each of the above patent documents is incorporated herein by reference in its entirety (for convenience, its disclosure may be referred to as "DeepSpeech 3" or "DV3"));

[0168] — Examples disclosed in the co-owned U.S. Patent 10,872,596 (Case No. 28888-2269) authorized on December 22, 2020 (this patent document is incorporated herein by reference in its entirety); and

[0169] — Examples disclosed in the co-owned U.S. Patent 11,017,761 (Case No. 28888-2326) authorized on May 25, 2021 (this patent document is incorporated herein by reference in its entirety).

[0170] Returning to Figure 24B , given video clip 2434 and corresponding audio 2436, the video and audio clips can be combined into a game highlight video 2440. Figure 27 Depicts a method for generating a combined video according to an embodiment of the present disclosure.

[0171] As described above, for each event, audio (2705) can be generated using TTS and a set of one or more sentences about the event, where TTS converts the set of one or more sentences into audio of a specific length. Having determined the exact or approximate time at which the event occurs in the video, a video clip (2710) including the event can be extracted from the full video, and its length can be selected such that it is long enough for the correspondingly generated audio. For example, if the audio generated by TTS for a sentence or group of sentences about an event (e.g., a corner kick) requires a certain amount of time, the corresponding video can be edited to match the length of the audio for the event in the video clip (e.g., the video can have a few seconds before the audio and a few seconds after the audio). Finally, the video clip and the corresponding audio can be combined into a multimedia video about the event. In one or more embodiments, tools (such as FFmpeg tools) can be used to combine the video clip and the audio.

[0172] It should be noted that the input text 2408 can refer to multiple events / actions and can associate and combine multiple video segments with corresponding audio. For example, in one or more embodiments, video clips are cascaded into a single final video and the audio is synchronized with the corresponding events. In one or more embodiments, tools (such as the FFmpeg tool) can be used to combine the video segments and audio.

[0173] F. Computational System Embodiments

[0174] In one or more embodiments, aspects of this patent document can be directed to one or more information processing systems (or computational systems), can include one or more information processing systems (or computational systems), or can be implemented on one or more information processing systems (or computational systems). An information processing system / computational system can include any operable tool or collection of tools to compute, estimate, determine, classify, process, send, receive, retrieve, originate, route, exchange, store, display, communicate, manifest, detect, record, reproduce, process, or utilize any form of information, intelligence, or data. For example, a computational system can be or can include a personal computer (e.g., a laptop computer), a tablet computer, a mobile device (e.g., a personal digital assistant (PDA), a smart phone, a phablet, a tablet, etc.), a smart card, a server (e.g., a blade server or a rack server), a network storage device, a camera, or any other suitable device, and can vary in size, shape, performance, function, and price. A computational system can include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, read only memory (ROM), and / or other types of memory. Additional components of a computational system can include one or more drives (e.g., a hard disk drive, a solid state drive, or both), one or more network ports for communicating with external devices and various input and output (I / O) devices (e.g., a keyboard, a mouse, a stylus, a touch screen, and / or a video display). A computational system can also include one or more buses for transmitting communications between the various hardware components.

[0175] Figure 28 A simplified block diagram of an information processing system (or computational system) according to an embodiment of the present disclosure is depicted. It should be understood that the functions shown for system 2800 can be used to support various embodiments of a computational system, although it should be understood that a computational system can be configured differently and include different components, including having Figure 28 fewer or more components as shown.

[0176] As Figure 28As shown, the computing system 2800 includes one or more central processing units (CPUs) 2801, which provide computing resources and control the computer. The CPU 2801 can be implemented with a microprocessor or the like, and the computing system 2800 may also include one or more graphics processing units (GPUs) 2802 and / or floating-point coprocessors for mathematical calculations. In one or more embodiments, one or more GPUs 2802 may be incorporated within the display controller 2809, such as a part of a graphics card. The system 2800 may also include a system memory 2819, which may include RAM, ROM, or both.

[0177] As Figure 28 shown, a plurality of controllers and peripherals may also be provided. The input controller 2803 represents an interface to various input devices 2804 such as a keyboard, mouse, touch screen, and / or stylus. The computing system 2800 may also include a storage controller 2807 for engaging with one or more storage devices 2808, each of the one or more storage devices 2808 including a storage medium such as a magnetic tape or disk or an optical medium, which may record instruction programs for an operating system, utilities, and applications, and the instruction programs may include embodiments of programs implementing various aspects of the present disclosure. The storage device 2808 may also be used to store processed data or data to be processed according to the present disclosure. The system 2800 may also include a display controller 2809 for providing an interface to a display device 2811, and the display device 2811 may be a cathode ray tube (CRT) monitor, a thin film transistor (TFT) monitor, an organic light-emitting diode, an electroluminescent panel, a plasma panel, or any other type of monitor. The computing system 2800 may also include one or more peripheral device controllers or interfaces 2805 for one or more peripheral devices 2806. Examples of peripheral devices may include one or more printers, scanners, input devices, output devices, sensors, etc. The communication controller 2814 may be connected to one or more communication devices 2815, which enables the system 2800 to connect to remote devices through any one of a variety of networks or through any suitable electromagnetic carrier signal including infrared signals, and the networks include the Internet, cloud resources (e.g., Ethernet cloud, Fibre Channel over Ethernet (FCoE) / Data Center Bridging (DCB) cloud, etc.), local area network (LAN), wide area network (WAN), storage area network (SAN). As shown in the depicted embodiment, the computing system 2800 includes one or more fans or fan trays 2818 and one or more cooling subsystem controllers 2817, and the cooling subsystem controller 2817 monitors the thermal temperature of the system 2800 (or its components) and operates the fan / fan tray 2818 to help regulate the temperature.

[0178] In the system shown, all of the major system components can be connected to bus 2816, which may represent more than one physical bus. However, the various system components may or may not be physically close to each other. For example, input data and / or output data may be transmitted remotely from one physical location to another. Additionally, programs implementing aspects of the present disclosure may be accessed from a remote location (e.g., a server) via a network. Such data and / or programs may be conveyed via any of a variety of machine-readable media, including, for example: magnetic media (such as hard disks, floppy disks, and magnetic tape); optical media (such as compact discs (CDs) and holographic devices); magneto-optical media; and hardware devices specially configured to store or execute program code (such as application specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices).

[0179] Aspects of the present disclosure may be encoded on one or more non-transitory computer-readable media having instructions for one or more processors or processing units to perform steps. It should be noted that one or more non-transitory computer-readable media should include volatile and / or non-volatile memory. It should be noted that alternative embodiments are possible, including hardware embodiments or software / hardware embodiments. Hardware-implemented functions may be implemented using ASICs, programmable arrays, digital signal processing circuitry, etc. Accordingly, the term "means" in any claim is intended to cover both software and hardware implementations. Similarly, the term "computer-readable media" as used herein includes software and / or hardware, or combinations thereof, on which an instruction program is included. Given these implementation alternatives, it should be understood that the drawings and the accompanying description provide the functional information needed by those skilled in the art to write program code (i.e., software) and / or fabricate circuitry (i.e., hardware) to perform the required processing.

[0180] It should be noted that embodiments of the present disclosure may also relate to computer products having a non - transitory, tangible computer - readable medium having computer code thereon for performing various computer - implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of the present disclosure, or they may be of the type well - known or available to those skilled in the relevant art. Examples of tangible computer - readable media include, for example: magnetic media (such as hard disks, floppy disks, and magnetic tapes); optical media (e.g., optical discs (CDs) and holographic devices); magneto - optical media; and hardware devices specially configured to store or execute program code (such as application - specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non - volatile memory (NVM) devices (such as 3D XPoint - based devices), and ROM and RAM devices). Examples of computer code include machine code such as that produced by a compiler, and files containing high - level code that is executed by a computer using an interpreter. Embodiments of the present disclosure may be implemented, in whole or in part, as machine - executable instructions that may be located in program modules executed by a processing device. Examples of program modules include libraries, programs, routines, objects, components, and data structures. In a distributed computing environment, program modules may be physically located in local, remote, or both settings.

[0181] Those skilled in the art will recognize that the computing system or programming language is not critical for the practice of the present disclosure. Those skilled in the art will also recognize that the above - mentioned multiple elements may be physically and / or functionally separated into modules and / or sub - modules or combined together.

[0182] Those skilled in the art should understand that the foregoing examples and embodiments are exemplary and are not intended to limit the scope of the present disclosure. All permutations, enhancements, equivalents, combinations, and improvements that are obvious to those skilled in the art after reading this specification and studying the drawings are included within the essence and scope of the present disclosure. It should also be noted that the elements of any claim may be arranged differently, including having multiple dependencies, configurations, and combinations.

Claims

1. A video generation method, comprising: Given input text that mentions events in an activity, Using a text parsing module to parse the input text to identify the events mentioned in the input text; Using a text-to-speech (TTS) module to convert the input text into TTS-generated audio; Given an input video of at least a part of the activity and the identified events: Using optical character recognition on a set of video frames of the input video to read the time of a clock displayed in the input video; Given the identified time of the clock, generating a set of time anchors including the start time of the activity and any time offsets; Using at least some of the set of time anchors to generate a time map for associating the time of the clock with the running time of the input video; Parsing data from metadata to obtain an approximate time of the event, the metadata being comment information and / or tag information of the event corresponding to the input text and the activity; using the approximate time of the event in the input video and the time map to generate an initial video clip; Initializing a slow-fast neural network model with pre-trained weights, adjusting the slow-fast neural network model to be a classifier using a combination of different loss functions for visual and audio modalities, and using a feature extractor part of the slow-fast neural network model to extract features from the initial video clip, the feature extractor being used to classify the video clip into the following 4 categories: away from the event of interest, before the event of interest, event of interest, after the event of interest; Using the extracted features and the trained neural network model to obtain a final time value of the event in the initial video clip; In response to the running time of the initial video clip being inconsistent with the running time of the TTS-generated audio, generating a final video clip by editing the initial video clip to have a running time consistent with the running time of the TTS-generated audio; In response to the running time of the initial video clip being consistent with the running time of the TTS-generated audio, using the initial video clip as the final video clip; Combining the TTS-generated audio with the final video clip to generate an event highlight video.

2. The video generation method according to claim 1, wherein The steps of extracting features from the initial video clip and using the extracted features and the trained neural network model to obtain a final time value of the event in the initial video clip include: Using a set of two or more models to perform feature extraction on the initial video clip; and Using an integrated neural network model to obtain the final time value of the event in the initial video clip, the integrated neural network model receiving inputs related to the features from the set of two or more models and outputting the final time value.

3. The video generation method according to claim 2, wherein, The set of two or more models includes: A neural network model that extracts video features from the initial video clip; A multi-modal feature neural network model that uses video and audio information in the initial video clip to generate multi-modal-based features; and An audio feature extractor that generates features of the initial video clip based on the audio level in the initial video clip.

4. The video generation method according to claim 1 further includes: Generating a data set, the data set including one or more entities related to the activity; And Using the data set to filter the text parsed from the input text to help identify which text parsed from the input text includes information about the event.

5. The video generation method according to claim 1 further includes: Given supplementary text about the activity or event: Using a text parsing module to parse the supplementary text to identify information related to the event mentioned in the supplementary text; And Using the information extracted from the supplementary text to help identify the occurrence time of the event.

6. The video generation method according to claim 1, wherein the parsing is neural network-based, rule-based, or both neural network-based and rule-based.

7. A video generation system includes: One or more processors; And A non-transitory computer-readable medium including a set or multiple sets of instructions, which when executed by at least one of the one or more processors, cause the following steps to be performed, the steps including: Given an input text mentioning an event in an activity: Using a text parsing module to parse the input text to identify the event mentioned in the input text; and Using a text-to-speech (TTS) module to convert the input text into TTS-generated audio; Given an input video of at least a part of the activity and the identified event: Using optical character recognition on a set of video frames of the input video to read the time of a clock displayed in the input video; given the identified time of the clock, generating a set of time anchors including the start time of the activity and any time offsets; using at least some of the set of time anchors to generate a time map for associating the time of the clock with the running time of the input video; Parsing data from metadata to obtain an approximate time of the event, the metadata being comment information and / or tag information of the event corresponding to the input text and the activity; using the approximate time of the event in the input video and the time map to generate an initial video clip; Initializing a slow-fast neural network model with pre-trained weights, adjusting the slow-fast neural network model to be a classifier using a combination of different loss functions for visual and audio modes, and using a feature extractor part of the slow-fast neural network model to extract features from the initial video clip, the feature extractor being used to classify the video clip into the following 4 categories: away from the event of interest, before the event of interest, event of interest, after the event of interest; Using the extracted features and the trained neural network model to obtain a final time value of the event in the initial video clip; In response to the running time of the initial video clip being inconsistent with the running time of the TTS-generated audio, generating a final video clip by editing the initial video clip to have a running time consistent with the running time of the TTS-generated audio; and In response to the running time of the initial video clip being consistent with the running time of the TTS-generated audio, use the initial video clip as the final video clip; and Combine the TTS-generated audio with the final video clip to generate an event highlight video.

8. The video generation system according to claim 7, wherein The steps of extracting features from the initial video clip and obtaining the final time value of the event in the initial video clip using the extracted features and a trained neural network model include: Performing feature extraction on the initial video clip using a set of two or more models; and Obtaining the final time value of the event in the initial video clip using an integrated neural network model that receives inputs related to the features from the set of two or more models and outputs the final time value.

9. The video generation system according to claim 8, wherein, The set of two or more models includes: A neural network model that extracts video features from the initial video clip; A multi-modal feature neural network model that utilizes video and audio information in the initial video clip to generate multi-modal-based features; and An audio feature extractor that generates features of the initial video clip based on the audio level in the initial video clip.

10. The video generation system according to claim 7, wherein, The one or more non-transitory computer-readable media further include one or more sets of instructions that, when executed by at least one of the one or more processors, cause the following steps to be performed, the steps including: Generating a data set that includes one or more entities related to the activity; and Filtering the text parsed from the input text using the data set to assist in identifying which text parsed from the input text includes information about the event.

11. The video generation system according to claim 7, wherein, The one or more non-transitory computer-readable media further include one or more sets of instructions that, when executed by at least one of the one or more processors, cause the following steps to be performed, the steps including: Given supplementary text about the activity or event: Parsing the supplementary text using a text parsing module to identify information related to the event mentioned in the supplementary text; and Using the information extracted from the supplementary text to assist in identifying the time of occurrence of the event.

12. A non-transitory computer-readable medium including one or more sequences of instructions that, when executed by at least one processor, cause the following steps to be performed, the steps including: Given an input text that mentions an event in an activity, Parsing the input text using a text parsing module to identify the event mentioned in the input text; And Converting the input text to TTS-generated audio using a text-to-speech (TTS) module; Given an input video of at least a portion of the activity and the identified event: Performing optical character recognition on a set of video frames of the input video to read the time of a clock displayed in the input video; Given the identified time of the clock, generating a set of time anchors that includes the start time of the activity and any time offsets. Generate a time map for associating the time of the clock with the running time of the input video using at least some of the set of time anchors; Parse data from the metadata to obtain an approximate time of the event, the metadata being comment information and / or tag information of the event corresponding to the input text and the activity; generate an initial video clip using the approximate time of the event in the input video and the time map; Initialize a slow-fast neural network model with pre-trained weights, adjust the slow-fast neural network model to be a classifier using a combination of different loss functions for visual and audio modalities, and extract features from the initial video clip using the feature extractor part of the slow-fast neural network model, the feature extractor being used to classify the video clip into the following 4 categories: away from the event of interest, before the event of interest, event of interest, after the event of interest; Using the extracted features and the trained neural network model, obtain the final time value of the event in the initial video clip; In response to the running time of the initial video clip being inconsistent with the running time of the audio generated by the TTS, generate a final video clip by editing the initial video clip to have a running time consistent with the running time of the audio generated by the TTS; And In response to the running time of the initial video clip being consistent with the running time of the audio generated by the TTS, use the initial video clip as the final video clip; And Combine the audio generated by the TTS with the final video clip to generate an event highlight video.

13. The non-transitory computer-readable medium according to claim 12, wherein, The steps of extracting features from the initial video clip and obtaining the final time value of the event in the initial video clip using the extracted features and the trained neural network model include: Performing feature extraction on the initial video clip using a set of two or more models; and Obtaining the final time value of the event in the initial video clip using an integrated neural network model, the integrated neural network model receiving inputs related to the features from the set of two or more models and outputting the final time value.

14. The non-transitory computer-readable medium according to claim 13, further comprising one or more instruction sequences that, when executed by at least one processor, cause the following steps to be performed, the steps including: Given supplementary text about the activity or event: Parse the supplementary text using a text parsing module to identify information related to the event mentioned in the supplementary text; And Use the information extracted from the supplementary text to assist in identifying the occurrence time of the event.

Citation Information

Patent Citations

  • Systems and methods for parallel wave generation in end-to-end text-to-speech

    US10872596B2

  • Systems and methods for real-time neural text-to-speech

    US10872598B2

  • Systems and methods for multi-speaker neural text-to-speech

    US10896669B2

  • Parallel neural text-to-speech

    US11017761B2

  • Systems and methods for neural text-to-speech using convolutional sequence learning

    US20190122651A1