Systems and methods for computer learning
Through the combination of time anchoring and neural network model, the event of interest in videos is automatically identified and refined, and the problem of generating highlight videos in the prior art requires a lot of manpower editing, achieving efficient and accurate automatic video editing.
Patent Information
- Application Number
- CN202111345521.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-03
- Filing Date
- 2021-11-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-15
AI Technical Summary
The prior art is difficult to generate highlight videos automatically and accurately, especially when processing large amounts of original videos, requiring a lot of manpower for manual editing.
By performing time anchoring, the video runtime is associated with the time of the captured items in the video, the approximate time of the event of interest is identified using metadata and association time, thereby generating clips, and using neural network models to extract features to accurately locate the final time value of the event of interest.
The automatic, large and precise generation of highlight videos is achieved, reducing the time and cost of manual editing, and accurately capturing key events in the video.
Smart Images

Figure CN114691923B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This patent application claims the benefit of priority of co - pending and co - owned U.S. Patent Application No. 63 / 124,832, filed on December 13, 2020, entitled "AUTOMATICALLY AND PRECISELY GENERATING HIGHLIGHT VIDEOS WITH ARTIFICIAL INTELLIGENCE", naming Zhiyu Cheng, Le Kang, Xin Zhou, Hao Tian, and Xing Li as inventors (Docket No. 28888 - 2450P (BN201118USN1 - Provisional)), and is related to the same. This patent document is hereby incorporated by reference in its entirety for all purposes. Background
[0003] A. Technical Field
[0004] The present disclosure generally relates to systems and methods for machine learning that can provide improved machine performance, features, and uses. More specifically, the present disclosure relates to systems and methods for automatically generating summaries or highlights of content.
[0005] B. Background Art
[0006] With the rapid development of Internet technologies and emerging tools, the amount of video content generated online (such as sports - related or other event videos) is growing at an unprecedented pace. Especially during the COVID - 19 pandemic, the viewership of online videos has surged as fans are not allowed to attend events at venues such as stadiums or arenas. Creating highlight videos or other event - related videos typically requires human effort to manually edit the original untrimmed videos. For example, the most popular sports videos usually consist of short clips of a few seconds, and it is extremely challenging for machines to precisely understand the videos and detect key events. Coupled with the large amount of raw videos available, refining the raw videos into appropriate highlight videos is very time - consuming and expensive. Moreover, given the limited time for viewing content, it is important for viewers to be able to obtain condensed content that appropriately captures the prominent elements or events.
[0007] Therefore, there is a need for systems and methods that can automatically and precisely generate refined or condensed video content, such as highlight videos. Summary of the Invention
[0008] In a first aspect, there is provided a computer - implemented method, comprising:
[0009] For each video from a set of videos:
[0010] Perform time anchoring to associate the video running time with the time of an event captured in the video;
[0011] Generate a clip from the video including the interesting event by using metadata related to the event and the associated time obtained by time anchoring to identify an approximate time of the interesting event at which the event occurs;
[0012] Perform feature extraction on the clip; and
[0013] Use the extracted features and a neural network model to obtain a final time value of the interesting event in the clip;
[0014] For each clip from a set of clips generated from the set of videos, compare the final time value with a corresponding ground truth value to obtain a loss value; and
[0015] Use the loss value to update the neural network model.
[0016] In a second aspect, a system is provided, comprising:
[0017] One or more processors; and
[0018] One or more non-transitory computer-readable media or mediums, the non-transitory computer-readable media or mediums including one or more instruction sets, the instruction sets when executed by at least one of the one or more processors cause steps to be performed including the following:
[0019] For each video from a set of one or more videos:
[0020] Perform time anchoring to associate the video running time with the time of an event captured in the video;
[0021] Generate a clip from the video including the interesting event by using metadata related to the event and the associated time obtained by time anchoring to identify an approximate time of the interesting event at which the event occurs;
[0022] Perform feature extraction on the clip; and
[0023] Use the extracted features and a neural network model to obtain a final time value of the interesting event in the clip; and
[0024] Use the final time value of the interesting event to produce a shorter clip including the interesting event.
[0025] In a third aspect, a computer-implemented method is provided, comprising:
[0026] For each video from a set of one or more videos:
[0027] Perform time anchoring to associate the video running time with the time of an event captured in the video;
[0028] Generate a clip from the video including the interesting event by using metadata related to the event and the associated time obtained by time anchoring to identify an approximate time of the interesting event that occurred during the event;
[0029] Perform feature extraction on the clip; and
[0030] Use the extracted features and a neural network model to obtain a final time value of the interesting event in the clip; and
[0031] Use the final time value of the interesting event to produce a shorter clip including the interesting event.
[0032] In a fourth aspect, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to the first aspect.
[0033] In a fifth aspect, a non-transitory computer-readable medium storing instructions is provided, wherein the instructions, when executed by a processor, cause the execution of the method according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Embodiments of the present disclosure will be referred to, examples of which may be shown in the drawings. These drawings are illustrative and not restrictive. Although the present disclosure is generally described in the context of these embodiments, it should be understood that it is not intended to limit the scope of the present disclosure to these specific embodiments. Items in the drawings may not be drawn to scale.
[0035] Figure 1 Depicts an overview of a highlight generation system according to an embodiment of the present disclosure.
[0036] Figure 2 Depicts an overview method for training a generation model according to an embodiment of the present disclosure.
[0037] Figure 3 Depicts an overall overview of a dataset generation process according to an embodiment of the present disclosure.
[0038] Figure 4 Summarizes some of the cloud-based text data of comments and tags according to an embodiment of the present disclosure.
[0039] Figure 5 Summarizes the untrimmed game videos collected according to embodiments of the present disclosure.
[0040] Figure 6 Depicts a user interface embodiment designed for a person to time video annotation events according to embodiments of the present disclosure.
[0041] Figure 7 Depicts a method for correlating event time and video running time according to embodiments of the present disclosure.
[0042] Figure 8 Shows an example of identifying timer digits in a match video according to embodiments of the present disclosure.
[0043] Figure 9 Depicts a method for generating clips based on an input video according to embodiments of the present disclosure.
[0044] Figure 10 Depicts feature extraction according to embodiments of the present disclosure.
[0045] Figure 11 Shows a pipeline for feature extraction according to embodiments of the present disclosure.
[0046] Figure 12 Graphically depicts a neural network model that can be used for feature extraction according to embodiments of the present disclosure.
[0047] Figure 13 Depicts feature extraction using a Slowfast neural network model according to embodiments of the present disclosure.
[0048] Figure 14 Depicts a method for audio feature extraction and prediction of the time of an event of interest in a video according to embodiments of the present disclosure.
[0049] Figure 15A Shows an example of an original audio waveform according to embodiments of the present disclosure, and FIG. 15B shows its corresponding mean absolute feature.
[0050] Figure 16 Depicts a method for predicting the time of an event of interest in a video according to embodiments of the present disclosure.
[0051] Figure 17 Shows a pipeline for time localization according to embodiments of the present disclosure.
[0052] Figure 18 Depicts a method for predicting the likelihood of an event of interest in a video clip according to embodiments of the present disclosure.
[0053] Figure 19 Shows a pipeline for action discovery prediction according to an embodiment of the present disclosure.
[0054] Figure 20 Describes a method for predicting the likelihood of an event of interest in a video clip according to an embodiment of the present disclosure.
[0055] Figure 21 Describes a method for performing final time prediction using an integrated neural network model according to an embodiment of the present disclosure.
[0056] Figure 22 Describes the target discovery results compared with another method according to an embodiment of the present disclosure.
[0057] Figure 23 Shows the target discovery results of three (3) clips according to an embodiment of the present disclosure. Ensemble learning achieves the best results.
[0058] Figure 24 Describes a simplified block diagram of a computing device / information processing system according to an embodiment of the present invention. Detailed Description
[0059] In the following description, for purposes of explanation, specific details are set forth in order to provide an understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these details. In addition, those skilled in the art will recognize that the embodiments of the present disclosure described below may be implemented in a variety of ways, such as a process, apparatus, system, device, or method on a tangible computer-readable medium.
[0060] The components or modules shown in the drawings are examples of exemplary embodiments of the present disclosure and are intended to avoid obscuring the present disclosure. It should also be understood that throughout the discussion, a component may be described as a separate functional unit that may include sub-units, but those skilled in the art will recognize that various components or portions thereof may be divided into separate components or may be integrated together, including, for example, within a single system or component. It should be noted that the functions or operations discussed herein may be implemented as components. The components may be implemented in software, hardware, or a combination thereof.
[0061] In addition, the connections between components or systems in the accompanying drawings are not limited to direct connections. Instead, the data between these components can be modified, reformatted, or otherwise changed by intermediate components. In addition, additional or fewer connections can be used. It should also be noted that the terms "coupled", "connected", "communicatively coupled", "interface connection", "interface", or any derivatives thereof should be understood to include direct connections, indirect connections through one or more intermediate devices, and wireless connections. It should also be noted that any communication such as a signal, response, reply, confirmation, message, query, etc. can include one or more information exchanges.
[0062] References in the specification to "one or more embodiments", "preferred embodiment", "embodiment", "multiple embodiments", etc. mean that the particular features, structures, characteristics, or functions described in connection with that embodiment are included in at least one embodiment of the present disclosure and may be in more than one embodiment. Further, the above phrases that appear in various places in the specification do not necessarily all refer to the same one or more embodiments.
[0063] The use of certain terms in various places in this specification is for illustrative purposes and should not be construed as limiting. A service, function, or resource is not limited to a single service, function, or resource; the use of these terms can refer to a combination of related services, functions, or resources, which can be distributed or aggregated. The terms "comprises", "comprising", "includes", and "including" should be understood as open-ended terms, and any list below is an example and is not meant to be limited to the items listed. A "layer" can include one or more operations. Words such as "optimal", "optimize", "optimization", etc. refer to an improvement in a result or process and do not require that the specified result or process reach an "optimal" or peak state. The use of memory, database, information repository, data storage, table, hardware, cache, etc. in this document can be used to refer to system components or components to which information can be input or otherwise recorded.
[0064] In one or more embodiments, the stop conditions can include: (1) a set number of iterations have been performed; (2) a certain amount of processing time has been reached; (3) convergence (e.g., the difference between consecutive iterations is less than a first threshold); (4) divergence (e.g., performance degradation); and (5) an acceptable result has been reached.
[0065] Those skilled in the art should recognize that: (1) certain steps can be optionally performed; (2) the steps are not limited to the specific order described herein; (3) certain steps can be performed in a different order; and (4) certain steps can be performed simultaneously.
[0066] Any headings used in this document are for organizational purposes only and should not be used to limit the scope of the specification or claims. Each reference / document mentioned in this patent document is incorporated herein by reference in its entirety.
[0067] It should be noted that any experiments and results provided herein are provided by way of example and are conducted under specific conditions using one or more specific embodiments; accordingly, neither these experiments nor their results should be used to limit the scope of the disclosure of this patent document.
[0068] It should also be noted that although the embodiments described herein may be in the context of a sports event (such as football), aspects of the present disclosure are not limited thereto. Accordingly, aspects of the present disclosure may be applied or adapted to other contexts.
[0069] A. Overview
[0070] 1. General Overview
[0071] Embodiments for automatically, massively, and precisely generating highlight videos are presented herein. For illustrative purposes, a football game will be used. However, it should be noted that the embodiments herein can be used or adapted for other sports events and non-sports events, such as concerts, shows, speeches, presentations, news, shows, video games, games, sports events, animations, social media posts, movies, etc. Each of these activities may be referred to as a matter or event, and the highlights of the matter may be referred to as events of interest, occurrences, or highlights.
[0072] Using a large-scale multimodal dataset, an existing deep learning model is created and trained to detect one or more events in the game, such as goals, but can also use a plethora of events of interest (e.g., penalties, injuries, fights, red cards, corner kicks, free kicks, etc.). Embodiments of an ensemble learning module for improving the performance of event-of-interest discovery are also presented herein.
[0073] Figure 1 An overview of a highlight generation system according to embodiments of the present disclosure is depicted. In one or more embodiments, large-scale cloud-derived text data and untrimmed football game videos are collected and fed into a series of data processing tools to generate candidate long clips (e.g., 70 seconds, but other time lengths can be used) that contain the main game events of interest (e.g., goal events). In one or more embodiments, a novel event-of-interest discovery pipeline precisely locates the instants of the events in the clips. Finally, embodiments can build one or more custom highlight videos / stories around the detected highlights.
[0074] Figure 2Depicts an overview method for training a generative model according to an embodiment of the present disclosure. To train a generative system, a large-scale multimodal dataset of event-related data must be obtained or generated (205) such that it can be used as training data. Since the video running time may not correspond to the time in the event, in one or more embodiments, for each video in a set of training videos, time anchoring is performed (210) to associate the video running time with the event time. The metadata (e.g., comments and / or tags) and the associated time obtained through time anchoring can then be used to identify (215) the approximate time of the event of interest to generate clips from the video including the event of interest. By using clips instead of the entire video, the processing requirements can be greatly reduced. For each clip, features are extracted (220). In one or more embodiments, a set of pre-trained models can be used to obtain the extracted features, which can be multimodal.
[0075] In one or more embodiments, for each clip, a neural network model is used to obtain (225) the final time value of the event of interest. In an embodiment, the neural network model can be an ensemble model that receives features from the set of models and outputs the final time value. Given the predicted final time value for each clip, the predicted final time value is compared with its corresponding ground truth value (230) to obtain a loss value; and the loss value can be used to update (235) the model.
[0076] Once trained, the generative system can output and be used to generate a highlight video in view of the input event video.
[0077] 2. Related Work
[0078] In recent years, artificial intelligence has been applied to analyze video content and generate videos. In sports analysis, many computer vision techniques have been developed to understand sports broadcasts. Specifically, in football, researchers have proposed the following algorithms: identifying key match events and player actions, using the body orientation of players to analyze passing feasibility, combining both audio and video streams to detect events, using broadcast streams and trajectory data to identify group activities on the field, aggregating depth frame features to discover major match events, and using the event context information around an action to process the inherent temporal patterns representing these actions.
[0079] Deep neural networks are trained with large-scale datasets for various video understanding tasks. Recent challenges include finding the temporal boundaries of activities or localizing events in the time domain. In football video understanding, some people define the goal event as the moment the ball crosses the goal line.
[0080] In one or more embodiments, this definition of a goal is adopted, and existing deep learning models and methods as well as audio stream processing techniques are utilized. Additionally, an ensemble learning module is adopted in the embodiments to accurately discover events in football video clips.
[0081] 3. Some contributions of the embodiments
[0082] In this patent document, embodiments of an automatic highlight generation video that can accurately identify the occurrence of events in a video are presented. In one or more embodiments, the system can be used to generate highlight videos in large quantities without traditional human editing. Some contributions provided by one or more embodiments include, but are not limited to, the following items:
[0083] – A large-scale multimodal football dataset is created, which includes cloud-derived text data and high-definition videos. Moreover, in one or more embodiments, various data processors are applied to parse, clean, and annotate the collected data.
[0084] – Align multimodal data from multiple sources and generate candidate long video clips by cutting the original video into 70-second clips using parsing tags from cloud-derived comment data.
[0085] – Embodiments of an event discovery pipeline are presented herein. The embodiments extract high-level feature representations from multiple perspectives and apply a temporal localization method to assist in discovering events in the clips. Additionally, the embodiments are further designed with an ensemble learning model to improve the performance of event discovery. It should be noted that although the matter can be a football game and the event of interest can be a goal, the embodiments can be used for or applied to other matters and other events of interest.
[0086] – Experimental results show that the tested embodiments achieve an accuracy close to 1 (0.984) with a 5-second tolerance in discovering goal events in the clips, which is better than existing work and establishes a new state of the art. This result helps to capture accurate goal moments and accurately generate highlight videos.
[0087] 4. Patent document layout
[0088] This patent document is organized as follows: Section B introduces the created dataset and how the data is collected and annotated. Section C presents the method for constructing an embodiment of the highlight generation system and how to accurately discover goal events in football clip videos using the proposed method. Section D summarizes and discusses the experimental results. It should be reiterated that only by way of illustration is a football game used as the overall content and a goal is used as the event within that content, and those skilled in the art should recognize that the methods herein can be applied to other content areas, including beyond the realm of sports, and applied to other events.
[0089] B. Data Processing Example
[0090] To train and develop system embodiments, a large-scale multi-modal dataset was created. Figure 3 depicts an overall overview of the dataset generation process according to an embodiment of the present disclosure. In one or more embodiments, one or more comments and / or tags associated with a video of an event are collected (305). For example, football game comments and tags (such as corner kicks, goals, blocks, throws, etc.) from websites and other sources can be crawled (see, for example Figure 1 in tags and comments 105) to obtain data. Also, videos associated with metadata (i.e., comments and / or tags) are collected (305). For the embodiments herein, high-definition (HD) uncropped football game videos from various sources are collected. Amazon Mechanical Turk (AMT) is used to annotate (315) the start time of the game in the uncropped original videos. In one or more embodiments, metadata (such as comment and / or tag information) can be used (320) to help identify the approximate time of an event of interest to generate clips (such as clips of goals) based on the video including the event of interest. Finally, Amazon Mechanical Turk (AMT) is used to identify the exact time of an event of interest (such as a goal) in the processed video clips. During the training of an embodiment of the goal discovery model, the annotated goal time can be used as the ground truth.
[0091] 1. Data Collection Example
[0092] In one or more embodiments, sports websites are crawled to obtain over 1,000,000 averages and tags, which cover over 10,000 football games from different leagues from the 2015 to 2020 seasons. Figure 4 summarizes some of the cloud-derived text data of comments and tags according to an embodiment of the present disclosure.
[0093] Comments and tags provide a large amount of information for each game. For example, they include the game date, team names, leagues, game event times (in minutes), event tags (such as goals, shots, corner kicks, substitutions, throw-ins, etc.), and the associated player names. These comments and tags from cloud-derived data can be converted or can be considered rich metadata for the original video processing embodiments as well as the highlight video generation embodiments.
[0094] In addition, over 2600 high-definition (720P or above) uncropped football game videos are collected from various online sources. The games are from various leagues from 2014 to 2020. Figure 5Summarizes untrimmed game videos collected according to embodiments of the present disclosure.
[0095] 2. Data Annotation Embodiments
[0096] In one or more embodiments, the untrimmed original video is first sent to Amazon Mechanical Turk (AMT) workers to annotate the start time of the game (defined as the time when the referee blows the whistle to start the game), and then the game commentary and tags sourced from the cloud are parsed to obtain the goal times (in minutes) for each game. By combining the goal minute tags with the start time of the game in the video, candidate 70 - second clips containing goal events are generated. Next, in one or more embodiments, these candidate clips are sent to AMT to annotate the goal times in seconds. Figure 6 Depicts a user interface embodiment designed for goal time annotation by AMT according to embodiments of the present disclosure.
[0097] For goal time annotation on AMT, each HIT (Artificial Intelligence Task, single worker assignment) contains one (1) candidate clip. Each HIT is assigned to five (5) AMT workers, and the median timestamp value is collected as the ground truth label.
[0098] C. Method Embodiments
[0099] In this section, details of embodiments of each of the five modules of the highlight generation system are presented. By way of a brief overview, the first module embodiment in section C.1 is a game time anchoring embodiment that checks the temporal integrity of the video and maps any time in the game to a time in the video.
[0100] The second module embodiment in section C.2 is a coarse interval extraction embodiment. This module is the main difference from the commonly studied event discovery pipeline. In the embodiments of this module, 70 - second intervals (although intervals of other sizes can be used) are extracted, where specific events are located by leveraging text metadata. There are at least three reasons for preferring this method over the common end - to - end visual event discovery pipeline. First, the clips extracted with metadata contain more context information and can be used across different domains. With metadata, the clips can be used for temporal cuts (such as game highlight videos) or can be used with other clips of the same team or player to generate team, player, and / or season highlight videos. The second reason is the robustness of low event ambiguity stemming from text data. And third, by analyzing shorter clips of the events of interest instead of the entire video, many resources (processing, processing time, memory, energy consumption, etc.) are conserved.
[0101] An embodiment of the third module in the system embodiment is multi-modal feature extraction. Video features are extracted from multiple perspectives.
[0102] An embodiment of the fourth module is precise time localization. Extensive studies of the techniques for designing and implementing embodiments of feature extraction and time localization are provided in Sections C.3 and C.4 respectively.
[0103] Finally, an embodiment of the integrated learning module is described in Section C.5.
[0104] 1. Embodiment of game time anchoring
[0105] It has been found that the event clock in event videos is sometimes irregular. The main reason seems to be that at least some of the time video files collected from the Internet contain corrupted timestamps or frames. It has been observed that in video collection, about 10% of the video files contain time corruption that moves the video in time (sometimes by more than 10 seconds). Some of the observed severe corruptions include the loss of more than 100 seconds of frames. In addition to errors in the video files, some unexpectedly rare events may have occurred during the matter / event, and the event clock must stop for a few minutes and then continue. If the video content is corrupted or the game is interrupted, the time violation can be regarded as a time jump forward or backward. To accurately locate the clip of the event specified by the metadata, in one or more embodiments, time jumps are detected and calibrated accordingly. Therefore, in one or more embodiments, an anchoring mechanism is designed and used.
[0106] Figure 7 A method for correlating event time and video running time according to an embodiment of the present disclosure is depicted. In one or more embodiments, OCR (Optical Character Recognition) is performed (705) on video frames at 5-second intervals (although other intervals can be used) to read the game clock displayed in the video. The start time of the game in the video can be inferred (710) from the recognized game clock. Whenever a time jump occurs, in one or more embodiments, a record of the game time after the time jump is retained (710), and this is referred to as or called a time anchor. Using the time anchor, in one or more embodiments, any time in the game can be mapped (715) to the time in the video (i.e., the video running time), and any clip specified by the metadata can be accurately extracted. Figure 8 An example of identifying timer digits in a game video according to an embodiment of the present disclosure is shown.
[0107] As Figure 8As shown, the timer digits 805 to 820 can be recognized and associated with the video running time. Embodiments can collect multiple recognition results over time and can perform self-correction based on spatial smoothness and temporal continuity.
[0108] 2. Embodiment of Coarse Interval Extraction
[0109] Figure 9 A method for generating clips according to an input video according to an embodiment of the present disclosure is depicted. In one or more embodiments, metadata from cloud-derived game commentary and tags is parsed (905), the metadata including minute-by-minute timestamps of goal events. Combining with the game start time detected by an embodiment of the OCR tool (discussed above), the original video can be edited to generate x-second (e.g., 70-second) candidate clips containing events of interest. In one or more embodiments, the extraction rule can be described by the following equations:
[0110] t {clipStart} = t {gameStart} + 60 * t {goalMinute} - tolerance 1)
[0111] t {clipEnd} = t {clipStart} +(base clip length + 2 * tolerance) 2)
[0112] In one or more embodiments, given the goal minute t {goalMinute} and the game start time t {gameStart} , a clip starting from t {clipStart} seconds in the video is extracted. In one or more embodiments, the duration of the candidate clip can be set to 70 seconds (where the base clip length is 60 seconds and the tolerance is 5 seconds, but it should be noted that different values and different schemes can be used), because this covers the corner kick situation where the event of interest occurs very close to the goal minute and also tolerates small deviations from the game start time detected by OCR. In the next section, method embodiments for finding the goal second (the moment the ball crosses the goal line) in the candidate clip are presented.
[0113] 3. Embodiments of Multi-Mode Feature Extraction
[0114] In this section, three embodiments for obtaining a high-level feature representation from candidate clips are disclosed.
[0115] a) Embodiment of Feature Extraction Using a Pre-Trained Model
[0116] Figure 10Depicts feature extraction according to an embodiment of the present disclosure. Given video data, in one or more embodiments, temporal frames are extracted (1005), and if necessary to match the input size, resized (1010) in the spatial domain to feed a deep neural network model to obtain a high-level feature representation. In one or more embodiments, a ResNet-152 model pre-trained on an image dataset is used, but other networks may be used. In one or more embodiments, temporal frames are extracted at the frames per second (fps) inherent in the original video and then downsampled to 2 fps, i.e., a ResNet-152 feature representation of 2 frames per second in the original video is obtained. ResNet is an extremely deep neural network that outputs a feature representation of 2048 dimensions per frame at the fully connected 1000-layer. In one or more embodiments, the output of the layer before the softmax layer can be used as the extracted high-level feature. It should be noted that ResNet-152 can be used to extract high-level features from a single image; this is not inherently embedded temporal context information. Figure 11 Illustrates a pipeline 1100 for extracting high-level features according to an embodiment of the present disclosure.
[0117] b) Slowfast feature extractor embodiment
[0118] As part of a video feature extractor, in one or more embodiments, a Slowfast network architecture can be used, such as the one proposed by Feichtenhofer et al. (Feichtenhofer, C., Fan, H., Malik, J., & He, K., Slowfast Networks for Video Recognition, Proceedings of the IEEE International Conference on Computer Vision (pp. 6202 - 6211) (2019), which is incorporated herein by reference in its entirety) or by Xiao et al. (Xiao et al., Audiovisual SlowFast Networks for Video Recognition, available at arxiv.org / abs / 2001.08740v1 (2020), which is incorporated herein by reference in its entirety); however, it should be noted that other network architectures may be used. Figure 12 Graphically depicts a neural network model that can be used to extract features according to an embodiment of the present disclosure.
[0119] Figure 13Depicts feature extraction using the Slowfast neural network model according to an embodiment of the present disclosure. In one or more embodiments, the Slowfast network is initialized (1305) with pre-trained weights using a training dataset. The network can be fine-tuned (1310) into a classifier. The second column in Table 1 below shows the event classification results using the baseline network in the case of the test dataset. In one or more embodiments, a feature extractor is used to divide 4-second clips into 4 categories: 1) far from the event of interest (e.g., a goal), 2) just before the event of interest, 3) the event of interest, and 4) just after the event of interest.
[0120] Several techniques can be implemented to find the best classifier evaluated by the top-1 error percentage. First, apply a network constructed such as Figure 12 which adds audio as an additional path to the Slowfast network (AVSlowfast). The visual part of the network can be initialized with the same weights. It can be seen that directly combining the training of visual and audio features actually degrades the performance. This is a common problem found when training multi-modal networks. In one or more embodiments, a technique of adding different loss functions to the visual and audio modalities respectively and training the entire network with a multi-task loss is applied. In one or more embodiments, a linear combination of the audio-visual results and the cross-entropy losses of each of the audio and visual branches can be used. The linear combination can be a weighted combination, where the weights can be learned or selected as hyperparameters. The best top-1 error result shown in the bottom row of Table 1 is obtained.
[0121] Table 1. Results of event classification.
[0122] Algorithm Error % of the 1st place Slowfast 33.27 Audio only 60.01 AVSlowfast 40.84 AVSlowfast multitasking 31.82
[0123] In one or more embodiments of the goal discovery pipeline, the feature extractor part of the network (AVSlowfast with multi-task loss) can be utilized. Thus, the aim is to reduce the top-1 error corresponding to stronger features.
[0124] c) Average absolute audio feature embodiment
[0125] By listening to the vocal cords of the event (e.g., a game without live commentary), one can often simply decide when the event of interest occurs based on the volume of the audience. Inspired by this observation, a simple method of directly extracting key information about the event of interest from audio has been developed.
[0126] Figure 14Describes a method for audio feature extraction and event time prediction of interest according to an embodiment of the present disclosure. In one or more embodiments, the absolute value of the audio wave is obtained and downsampled (1405) to 1 Hertz (Hz). This feature representation can be referred to as the average absolute feature because it represents the average sound amplitude per second. Figure 15A and Figure 15B Respectively show examples of the original audio wave and its average absolute feature of a clip according to an embodiment of the present disclosure.
[0127] For each clip, the maximum value 1505 of this average absolute audio feature 1500B can be located (1410). By locating the maximum value (e.g., maximum value 1505) of this average absolute audio feature and its corresponding time (e.g., time 1510) for the clips in the test dataset, an accuracy of 79% for event time localization is achieved (with a tolerance of 5 seconds).
[0128] In one or more embodiments, for the time in a clip, the average absolute audio feature (e.g., Figure 15B 1500B in ) can be regarded as a likelihood prediction of an event of interest. As will be discussed below, this average absolute audio feature can be a feature input to an ensemble model for predicting the final time within a clip where an event of interest occurs.
[0129] 4. Action discovery embodiments
[0130] To accurately discover the goal-scoring moment in a sufficient number of videos, in one or more embodiments, the temporal context information surrounding that moment is combined to learn what is happening in the video. For example, before the goal event occurs, the goal-scoring player will shoot (or head the ball), and the ball will move towards the goal. In some cases, both offensive and defensive players gather in the penalty area and are not far from the goal. After the goal event, typically, the goal-scoring player will run to the sidelines, hug their teammates, and the viewers and coaches will also celebrate. Intuitively, these patterns in the video can help the model learn what is happening and discover the moment of the goal event.
[0131] Figure 16A method for predicting the likelihood of an event of interest in a video clip according to an embodiment of the present disclosure is depicted. In one or more embodiments, to construct a temporal localization model, a temporal convolutional neural network is used (1605) that takes the extracted visual features as input. In one or more embodiments, the input features can be extracted features from one or more of the prior models discussed above. For each frame, a set of intermediate features that mix temporal information over the frame is output. Then, in one or more embodiments, the intermediate features are input (1610) into a segmentation model that generates a segmentation score that is evaluated by a segmentation loss function. A cross entropy loss function can be used for the segmentation loss function:
[0132]
[0133] where t i is the ground truth label, and p i is the softmax probability of the i-th class.
[0134] In one or more embodiments, the segmentation scores and intermediate features are concatenated and fed (1615) into an action discovery module, which generates (1620) discovery predictions (e.g., likelihood predictions over the span of the clip that an event of interest occurred at each time point) that can be evaluated by an action discovery loss function similar to YOLO. An L2 loss function can be used for the action discovery loss function:
[0135]
[0136] Figure 17 A pipeline for temporal localization according to an embodiment of the present disclosure is shown. In one or more embodiments, the temporal CNN may include a convolutional layer, the segmentation model may include a convolutional layer and a batch normalization layer, and the action discovery model may include a pooling layer and a convolutional layer.
[0137] In one or more embodiments, a model embodiment is trained with a segmentation and action discovery loss function that takes into account temporal context information, as described by Cioppa et al. (A. Deliège, A. Giancola, S. Ghanem, B. Droogenbroeck, M. V. Gade R., and Moeslund T., “A Context-Aware Loss Function for Action Spotting in Soccer Videos,” 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), which is incorporated herein by reference in its entirety). In one or more embodiments, a segmentation loss is used to train a segmentation model, where each frame is associated with a score representing how likely the frame belongs to an action class, and an action discovery loss is used to train an action discovery module, where temporal localizations are predicted for action classes.
[0138] At least one of the main differences between the embodiments herein and the method of Cioppa et al. is that the embodiments herein process short clips, while Cioppa et al. take the entire match video as input and thus require much more time to process the video and extract features when implemented in real time.
[0139] In one or more embodiments, the input of the extracted features can be the extracted features from the ResNet model discussed above or the AVSlowFast multi-task model discussed above. Alternatively, for the AVSlowFast multi-task model, the segmentation part of the action discovery model can be removed. Figure 18 A method for predicting the likelihood of an event of interest in a video clip according to an embodiment of the present disclosure is depicted. In one or more embodiments, a temporal convolutional neural network receives (1805) the extracted features from the AVSlowfast multi-task model as input. For each frame, a set of intermediate features that incorporate the temporal information on the frame is output. Then, in one or more embodiments, the intermediate features are input (1810) into an action discovery module that generates (1815) a discovery prediction (e.g., a likelihood prediction over the span of a clip of an event of interest occurring at each time point) that can be evaluated by an action discovery loss function. Figure 19 A pipeline for action discovery prediction according to an embodiment of the present disclosure is shown.
[0140] 5. Ensemble learning embodiments
[0141] In one or more embodiments, a single predicted time (e.g., pick the maximum value) of the event of interest in the clip can be obtained from each of the three models discussed above. One of the predictions can be used or the predictions can be combined (e.g., averaged). Alternatively, an ensemble model can be used to combine the information from each model in order to obtain a final prediction of the event of interest in the clip.
[0142] Figure 20 A method for predicting the likelihood of an event of interest in a video clip according to an embodiment of the present disclosure is depicted, and Figure 21 A pipeline for final time prediction according to an embodiment of the present disclosure is depicted. In one or more embodiments, the outputs of the three models / features described in the above subsections can be aggregated in an ensemble manner to enhance the final accuracy. In one or more embodiments, the outputs of all three of the previous models, along with the position encoding vector, can be combined as the input (2005) to the ensemble module. Cascade can be used to accomplish the combination. For example, four d-dimensional vectors become a 4×d matrix. For the ResNet and AVSlowfast multi-task models, the input can be the likelihood prediction output of their action discovery models from Section 5 above. Also, for audio, the input can be the average absolute audio features of the clip (e.g., Figure 15B ). In one or more embodiments, the position encoding vector is a 1-D vector (i.e., index) representing the temporal length of the clip.
[0143] In one or more embodiments, the core of the ensemble module is an 18-layer 1-D ResNet with a recurrent head. Essentially, the ensemble module learns a mapping from multi-dimensional input features including multiple patterns to the final temporal position of the event of interest in the clip. In one or more embodiments, the output (2010) is the final time value prediction from the ensemble model, and it can be compared with the ground truth time to calculate the loss. The losses for individual clips can be used to update the parameters of the ensemble model.
[0144] 6. Inference Embodiments
[0145] Once trained, the overall highlight generation system can be deployed, such as Figure 1Depicted. In one or more embodiments, the system may additionally include an input that allows a user to select one or more parameters of the generated clip. For example, the user may use a specific layer, the span of the game, one or more events of interest (e.g., goals and penalties), and the number of clips for creating the highlight video (or the time length of each clip and / or the overall highlight compilation video). The highlight generation system may then access the video and metadata and generate a highlight compilation video by concatenating the clips. For example, the user may want 10-second clips for each event of interest. Thus, in one or more embodiments, the customized highlight video generation module may obtain the final predicted time of the clip and select 8 seconds before and 2 seconds after the event of interest. Alternatively, as Figure 1 shown, key events in a player's career may be events of interest, and these events may be automatically identified and compiled into a "story" of the player's career. Audio and other multimedia features may be added to the video by the customized highlight video generation module, and the audio and features may be selected by the user. Those skilled in the art should recognize other applications of the highlight generation system.
[0146] D. Experimental Results
[0147] Note that these experiments and constructs are provided by way of illustration only and are performed using one or more specific embodiments under specific conditions; thus, neither these experiments nor their results should be used to limit the scope of the disclosure of this patent document.
[0148] 1. Goal Discovery
[0149] For a fair comparison with existing work, the model embodiments of the test are trained with candidate clips containing goals extracted from the games in the training set of the dataset, and validated / tested with candidate clips containing targets extracted from the games in the validation / test set of the dataset.
[0150] Figure 22 Shows the main results: To discover goals in 70-second clips, the tested embodiment 2205 is significantly better than the current state-of-the-art method 2210 in discovering goals in soccer, called the context-aware method.
[0151] Also shows the intermediate prediction results obtained by using the three different features described in Section C.3 or C.4, and the final results predicted by the ensemble learning module described in Section C.5. The goal discovery results for 3 clips are stacked in Figure 23 As Figure 23 shown, in terms of its proximity to the ground truth label (shown by the dashed ellipse), the final prediction output of the ensemble learning module embodiment is the best.
[0152] 2. Some Discussion Notes
[0153] As Figure 9 shown, an embodiment can achieve an accuracy close to 1 (0.984) with a tolerance of 5 seconds. This result is remarkable as it can be used to correct mislabeling of text and synchronize with custom audio commentary. It also helps to precisely generate highlights and thus gives the user / editor the option to customize their video around the exact goal moment. The pipeline embodiment can be naturally extended to capture moments of other events such as corner kicks, free kicks, and penalties.
[0154] It should be reiterated again that the soccer game is used as the overall content and the goal as an event within that content by way of illustration only, and those skilled in the art should recognize that the methods herein can be applied to other content areas, including beyond the realm of sports, and to other events.
[0155] E. Computing System Embodiments
[0156] In one or more embodiments, aspects of this patent document can be directed to one or more information processing systems (computing systems), can include one or more information processing systems (computing systems), or can be implemented on one or more information processing systems (computing systems). An information processing system / computing system can include any tool or collection of tools operable to compute, calculate, determine, classify, process, send, receive, retrieve, originate, route, transform, store, display, communicate, manifest, detect, record, reproduce, handle, or utilize any form of information, knowledge, or data. For example, a computing system can be or can include a personal computer (e.g., a laptop computer), a tablet computer, a mobile device (e.g., a personal digital assistant (PDA), a smart phone, a phablet, a tablet computer, etc.), a smart watch, a server (e.g., a blade server or a rack server), a network storage device, a camera, or any other suitable device, and can vary in size, shape, performance, function, and price. A computing system can include random access memory (RAM), one or more processing resources (such as a central processing unit (CPU) or hardware or software control logic), read-only memory (ROM), and / or other types of memory. Additional components of a computing system can include one or more drives (e.g., a hard disk drive, a solid state drive, or both), one or more network ports for communicating with external devices, and various input and output (I / O) devices (e.g., a keyboard, a mouse, a touch screen, and / or a video display). A computing system can also include one or more buses for transmitting communications between the various hardware components.
[0157] Figure 24FIG. depicts a simplified block diagram of an information processing system (or computing system) in accordance with an embodiment of the present disclosure. It should be understood that the functionality shown in system 2400 can be used to support various embodiments of a computing system, although it should be understood that the computing system can be configured differently and include different components, including having fewer or more components as shown in Figure 24 shown.
[0158] As Figure 24 shown, computing system 2400 includes one or more central processing units (CPUs) 2401, which provide computing resources and control the computer. The CPU 2401 can be implemented with a microprocessor or the like, and can also include one or more graphics processing units (GPUs) 2402 and / or floating point coprocessors for mathematical calculations. In one or more embodiments, one or more CPUs 2402 can be incorporated within a display controller 2409, such as part of one or more graphics cards. System 2400 can also include system memory 2419, which can include RAM, ROM, or both.
[0159] Multiple controllers and peripherals can also be provided, such as Figure 24As shown. The input controller 2403 represents an interface to various input devices 2404 such as a keyboard, mouse, touch screen, and / or stylus. The computing system 2400 may also include a storage controller 2407 for interfacing with one or more storage devices 2408, each of the one or more storage devices 2408 including a storage medium such as a magnetic tape or disk or an optical medium that may be used to record programs of instructions for an operating system, utilities, and applications, the operating system, utilities, and applications may include embodiments of programs that implement various aspects of the present disclosure. The storage device 2408 may also be used to store data processed according to the present disclosure or data to be processed. The system 2400 may also include a display controller 2409 for providing an interface to a display device 2411, which may be a cathode ray tube (CRT) display, a thin film transistor (TFT) display, an organic light emitting diode, an electroluminescent panel, a plasma panel, or any other type of display. The computing system 2400 may also include one or more peripheral device controllers or interfaces 2405 for one or more peripheral devices 2406. Examples of peripheral devices may include one or more printers, scanners, input devices, output devices, sensors, etc. The communication controller 2414 may interface with one or more communication devices 2415, which enables the system 2400 to connect to remote devices through any one of various networks including the Internet, cloud resources (e.g., Ethernet cloud, Fibre Channel over Ethernet (FCoE) / Data Center Bridging (DCB) cloud over Ethernet, etc.), local area network (LAN), wide area network (WAN), storage area network (SAN)), or through any suitable electromagnetic carrier signal including infrared signals. As shown in the depicted embodiment, the computing system 2400 includes one or more fans or fan trays 2418 and one or more cooling subsystem controllers 2417, and the cooling subsystem controller 2417 monitors the thermal temperature of the system 2400 (or its components) and operates the fan / fan tray 2418 to help regulate the temperature.
[0160] In the system shown, all of the major system components can be connected to a bus 2416, which can represent more than one physical bus. However, the various system components can be physically close to or not physically close to each other. For example, input data and / or output data can be transmitted remotely from one physical location to another physical location. In addition, programs implementing aspects of the present disclosure can be accessed over a network from a remote location (e.g., a server). Such data and / or programs can be conveyed via any of a variety of machine-readable media, including, for example: magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as compact discs (CDs) and holographic devices; magneto-optical media; and hardware devices specifically configured to store or execute program code, such as application specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices.
[0161] Aspects of the present disclosure can be encoded on one or more non-transitory computer-readable media having instructions for one or more processors or processing units to cause steps to be performed. It should be noted that one or more non-transitory computer-readable media can include volatile and / or non-volatile memory. It should be noted that alternative implementations are possible, including hardware implementations or software / hardware implementations. Hardware-implemented functions can be implemented using ASICs, programmable arrays, digital signal processing circuitry, etc. Accordingly, the term "means" in any claim is intended to cover both software and hardware implementations. Similarly, the term "computer-readable media" as used herein includes software and / or hardware, or combinations thereof, having an instruction program contained thereon. Given these implementation alternatives, it should be understood that the drawings and the accompanying description provide the functional information necessary for a person of ordinary skill in the art to write program code (i.e., software) and / or fabricate circuitry (i.e., hardware) to perform the required processing.
[0162] It should be noted that embodiments of the present disclosure may also relate to computer products having a non - transitory, tangible computer - readable medium, having computer code thereon for performing various computer - implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the type known or available to those skilled in the relevant art. Examples of tangible computer - readable media include, for example: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CDs and holographic devices; magneto - optical media; and hardware devices specifically configured to store or execute program code, such as application - specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non - volatile memory (NVM) devices (such as 3D XPoint - based devices), and ROM and RAM devices. Examples of computer code include, for example, machine code generated by a compiler, and files containing higher - level code that is executed by a computer using an interpreter. Embodiments of the present disclosure may be implemented, in whole or in part, as machine - executable instructions that may be in program modules executed by a processing device. Examples of program modules include libraries, programs, routines, objects, components, and data structures. In a distributed computing environment, program modules may be physically located locally, remotely, or in a combination of both settings.
[0163] Those skilled in the art will recognize that no computing system or programming language is critical for the practice of the present disclosure. Those skilled in the art will also recognize that the above - mentioned multiple elements may be physically and / or functionally separated into modules and / or sub - modules or combined together.
[0164] Those skilled in the art should understand that the foregoing embodiments and implementations are exemplary and do not limit the scope of the present disclosure. All permutations, enhancements, equivalents, combinations, and improvements thereof that are obvious to those skilled in the art after reading this specification and studying the drawings are included within the true spirit and scope of the present disclosure. It should also be noted that the elements of any claim may be arranged differently, including having multiple dependencies, configurations, and combinations.
Claims
1. A computer-implemented method, comprising: For each video from a set of videos: Perform time anchoring to associate the video running time with any time of the matter captured in the video; Generate a clip from the video including the interesting event by using the metadata related to the matter and the associated time obtained by time anchoring to identify the approximate time of the interesting event when the matter occurs, the metadata including comments and / or tags; Perform feature extraction on the clip using a set of two or more models, including extracting video features, audio features, and multimodal features from the clip; and Use the extracted features and a neural network model to obtain the final time value of the interesting event in the clip, the neural network model being an integrated neural network model, the integrated neural network model receiving inputs related to the features from the set of two or more models and outputting the final time value; For each clip from a set of clips generated from the set of videos, compare the final time value with the corresponding ground truth value to obtain a loss value; and Use the loss value to update the neural network model.
2. The computer-implemented method according to claim 1, wherein the step of performing time anchoring to associate the video running time with the time of the matter captured in the video further comprises: Use optical character recognition on a set of video frames of the video to read the time of the clock displayed in the video; Generate a set of time anchors including the start time of the matter and any time offset in view of the recognized time of the clock; and Use at least some of the time anchors in the set of time anchors to generate a time mapping that maps the time of the clock to the video running time.
3. The computer-implemented method according to claim 1, wherein the step of generating a clip from the video including the interesting event by using the metadata related to the matter and the associated time obtained by time anchoring to identify the approximate time of the interesting event when the matter occurs comprises: For the video, parse the data from the metadata to obtain the approximate timestamp of the interesting event; and Use the approximate timestamp of the interesting event in the video and the time mapping to generate one or more candidate clips, where the candidate clips include the interesting event.
4. The computer-implemented method according to claim 1, wherein the set of two or more models comprises: A neural network model that extracts video features from the clip; A multimodal feature neural network model that uses video and audio information in the clip to generate multimodal-based features; and An audio feature extractor that generates features of the clip based on the audio volume in the clip.
5. The computer-implemented method according to claim 1, wherein the integrated neural network model receives the following as inputs: The first likelihood prediction feature of the clip obtained as input from an action discovery model system, the action discovery model system receiving the extracted features from a neural network model that extracts video features from the clip; The second likelihood prediction feature of the clip obtained as input from an action discovery and segmentation model system, the action discovery and segmentation model system receiving the extracted features from a multimodal feature neural network model that uses video and audio information in the clip to generate multimodal-based features; And The mean absolute audio feature obtained from an audio feature extractor, the audio feature extractor generating the mean absolute audio feature of the clip based on the audio volume in the clip.
6. The computer-implemented method according to claim 5, further comprising: Combining the positional encoding representation with the first likelihood prediction feature, the second likelihood prediction feature, and the mean absolute audio feature as input to the integrated neural network model.
7. A system, comprising: One or more processors; And One or more non-transitory computer-readable media or media, the non-transitory computer-readable media or media comprising one or more instruction sets that, when executed by at least one of the one or more processors, cause steps to be performed including the following: For each video from a set of one or more videos: Perform time anchoring to associate the video runtime with the time of any matter captured in the video; Generate a clip from the video including the interesting event by using metadata related to the matter and the associated time obtained by time anchoring to identify the approximate time of the interesting event in which the matter occurs, the metadata including comments and / or tags; Perform feature extraction on the clip using a set of two or more models, including extracting video features, audio features, and multimodal features from the clip; And Using the extracted features and a neural network model to obtain a final time value of the interesting event in the clip, the neural network model being an integrated neural network model that receives inputs related to the features from the set of two or more models and outputs the final time value; And Using the final time value of the interesting event to produce a shorter clip including the interesting event.
8. The system according to claim 7, wherein the step of performing time anchoring to associate the video runtime with the time of the matter captured in the video further comprises: Using optical character recognition on a set of video frames of the video to read the time of a clock displayed in the video; Generating a set of time anchors including the start time of the matter and any time offset in view of the recognized time of the clock; And Using at least some of the time anchors in the set of time anchors to generate a time mapping that maps the time of the clock to the video runtime.
9. The system according to claim 8, wherein the step of generating a clip from the video including the interesting event by using metadata related to the matter and the associated time obtained by time anchoring to identify an approximate time of the interesting event occurring at the matter comprises: For a video, parsing data from the metadata to obtain an approximate timestamp of the interesting event; and using the approximate timestamp of the interesting event in the video and the time mapping to generate one or more candidate clips, wherein the candidate clips include the interesting event.
10. The system according to claim 7, wherein the set of two or more models comprises: A neural network model that extracts video features from the clip; A multi-modal feature neural network model that utilizes video and audio information in the clip to generate multi-modal based features; and An audio feature extractor that generates features of the clip based on the audio volume in the clip.
11. The system according to claim 7, wherein the integrated neural network model receives the following as inputs: The first likelihood prediction feature of the clip obtained as an input from an action discovery model system, the action discovery model system receiving the extracted features from a neural network model that extracts video features from the clip; The second likelihood prediction feature of the clip obtained as an input from an action discovery and segmentation model system, the action discovery and segmentation model system receiving the extracted features from a multi-modal feature neural network model that utilizes video and audio information in the clip to generate multi-modal based features; and The average absolute audio feature obtained from an audio feature extractor, the audio feature extractor generating the average absolute audio feature of the clip based on the audio volume in the clip.
12. The system according to claim 11, wherein the one or more non-transitory computer-readable media or media further comprise one or more instruction sets, the one or more instruction sets further causing, when executed by at least one of the one or more processors, the execution of steps including the following: Combining a positional encoding representation with the first likelihood prediction feature, the second likelihood prediction feature, and the average absolute audio feature as an input to the integrated neural network model.
13. The system according to claim 7, wherein the one or more non-transitory computer-readable media or media further comprise one or more instruction sets, the one or more instruction sets further causing, when executed by at least one of the one or more processors, the execution of steps including the following: Combining a set of short clips together to produce a compilation highlight video.
14. A computer-implemented method, comprising: For each video from a set of one or more videos: Performing time anchoring to associate the video running time with any time of the matter captured in the video; Generating a clip from a video including the event of interest by using metadata related to the matter and an associated time obtained by time anchoring to identify an approximate time of the event of interest occurring during the matter, the metadata including comments and / or tags; Performing feature extraction on the clip using a set of two or more models, including extracting video features, audio features, and multimodal features from the clip; and using the extracted features and a neural network model to obtain a final time value of the event of interest in the clip, the neural network model being an integrated neural network model that receives inputs related to the features from the set of two or more models and outputs the final time value; and using the final time value of the event of interest to produce a shorter clip including the event of interest.
15. The computer-implemented method according to claim 14, wherein the step of performing time anchoring to associate the video running time with the time of a matter captured in the video further includes: using optical character recognition on a set of video frames of the video to read the time of a clock displayed in the video; generating a set of time anchors including a start time of the matter and any time offsets in view of the recognized time of the clock; and using at least some of the time anchors in the set of time anchors to generate a time map that maps the time of the clock to the video running time.
16. The computer-implemented method according to claim 15, wherein the step of generating a clip from a video including the event of interest by using metadata related to the matter and the associated time obtained by time anchoring to identify an approximate time of the event of interest occurring during the matter includes: for the video, parsing data from the metadata to obtain an approximate timestamp of the event of interest; and using the approximate timestamp of the event of interest in the video and the time map to generate one or more candidate clips, wherein a candidate clip includes the event of interest.
17. The computer-implemented method according to claim 14, wherein the integrated neural network model receives as inputs: a first likelihood prediction feature of the clip obtained as an input from an action discovery model system that receives the extracted features from a neural network model that extracts video features from the clip; a second likelihood prediction feature of the clip obtained as an input from an action discovery and segmentation model system that receives the extracted features from a multimodal feature neural network model that uses video and audio information in the clip to generate multimodal-based features; an average absolute audio feature obtained from an audio feature extractor that generates the average absolute audio feature of the clip based on the audio volume in the clip; and a positional encoding representation.
18. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.
19. A non-transitory computer-readable medium storing instructions, wherein, the instructions, when executed by a processor, cause the method according to any one of claims 1-6 to be executed.
Citation Information
Patent Citations
Video event recognition method and system, electronic equipment and medium
CN110942011A
System and method for audio event detection in surveillance systems
CN111742365A
Competition video structuring method, device and system
CN111757147A