Technologies for generating media content

By automatically detecting and analyzing triggers in the media system, the highest quality media segments are selected for aggregation, solving the problem of generating and aggregating media content in a multi-device environment, improving efficiency and reducing costs.

CN117041694BActive Publication Date: 2026-05-26OPEN TV INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OPEN TV INC
Filing Date
2019-09-26
Publication Date
2026-05-26

Smart Images

  • Figure CN117041694B_ABST
    Figure CN117041694B_ABST
Patent Text Reader

Abstract

This disclosure relates to techniques for generating media content. Techniques and systems for generating media content are provided. (E.g., via a server computer or other device or system) Triggers associated with events at the scene can be detected from devices located at the scene. Media segments captured by multiple media capture devices located at the scene can be obtained. At least one of the media segments corresponds to a detected trigger. One or more quality metrics of the media segment can be determined based on a first motion of an object captured in the media segment and / or a second motion of the media capture device used to capture the media segment. A subset of media segments can be selected from the media segments based on the quality metrics determined for the obtained media segments. An aggregate of media segments including the subset of media segments can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with application number 201980063064.7, application date September 26, 2019, entitled "Technology for Generating Media Content".

[0002] Cross-reference to related applications

[0003] This application claims the benefit of U.S. Patent Application No. 16 / 145,774, filed September 28, 2018, which is incorporated herein by reference in its entirety for all purposes. Technical Field

[0004] This disclosure generally relates to technologies and systems for generating media content, and more specifically to media content for generating events. Background Technology

[0005] Media capture devices can capture various types of media content, including images, video, and / or audio. For example, a camera can capture image or video data of a scene. Media data from a media capture device can be captured and output for processing and / or consumption. For example, video of a scene can be captured and processed for display on one or more viewing devices.

[0006] Media content from multiple media capture devices can be used to generate media clips that include different qualities, perspectives, and events within a scene. However, it can be difficult to capture the highest quality and most relevant media clips from different media capture devices at different times. This problem becomes even more challenging when a large amount of media content is available for processing. Summary of the Invention

[0007] In some examples, this document discloses techniques and systems for generating media content. For instance, a media system may acquire different media content items from one or more media capture devices located at a site. Media content items may include video, audio, images, any combination thereof, and / or any other type of media. The media system may curate certain media segments from the acquired media content items into a collection of media content.

[0008] In some examples, a media content aggregation may include a set of selected media clips organized in a way that results in multiple shortened clips of the captured media content having different characteristics. For example, the media clips in a media content aggregation may come from different points in time, from different viewpoints or perspectives within the location, may have different zoom levels, may have different display and / or audio characteristics, may have combinations of these, and / or may have any other suitable variations between the media clips. In some cases, the acquired media content may be captured during an event occurring at the location. For example, a media content aggregation may provide a highlight reel of an event by including a set of media clips capturing an event at different points in time, either from a single viewpoint or multiple viewpoints (e.g., different viewpoints from different perspectives).

[0009] Candidate media segments that can be included in a media content aggregation can be selected based on triggers associated with the media segments. Triggers can be used to indicate that a moment of interest has occurred within the event during which the media content is captured. In some cases, triggers can be associated with time within the event. Other data associated with the event can also be obtained, and in some cases, associated with time within the event. In some examples, triggers can be generated by a device located at the event location where media content is being captured and / or by a device of a user remotely observing the event (e.g., a user watching the event on a television, mobile device, personal computer, or other device). In some examples, triggers can be automatically generated based on indications that a moment of interest may have occurred during the event. For example, triggers can be generated by a device based on the detection of the event's occurrence (also referred to as a moment) during the event, based on characteristics of users of devices located at the event location, based on characteristics of users remotely observing the event, any combination thereof, and / or other indications that a significant event occurred during the event.

[0010] Media systems can analyze the quality of selected candidate media segments to determine which segments will be included in the media content aggregate. For example, the quality of a media segment can be determined based on factors that indicate the level of interest those segments might have for potential viewers. In some examples, the quality of a media segment can be determined based on the motion of one or more objects captured in the segment, the motion of the media capture device used to capture the segment, the number of triggers associated with the segment, the presence of specific objects in the segment, their combination, and / or any other suitable characteristics of the segment. The highest quality candidate media segments can then be selected from the candidate media segments to be included in the media content aggregate.

[0011] According to at least one example, a method for generating media content is provided. The method includes a server computer detecting a trigger from a device associated with an event at a location. The device does not capture media segments for the server computer. The method also includes the server computer acquiring media segments captured by a plurality of media capture devices located at the location. At least one of the acquired media segments corresponds to the detected trigger. The server computer does not acquire media segments that do not correspond to the trigger. The method further includes the server computer determining one or more quality metrics for each media segment among the acquired media segments. The one or more quality metrics for the media segments are determined based on at least one of a first motion of an object captured in the media segment and a second motion of the media capture device used to capture the media segment. The method also includes selecting a subset of media segments from the acquired media segments, the subset being selected based on the one or more quality metrics determined for each media segment among the acquired media segments. The method further includes generating an aggregate of media segments including the subset of said media segments.

[0012] In another example, a system is provided that includes one or more processors and a memory configured to store media data. The one or more processors are configured to detect triggers from a device associated with an event at a location. The device does not capture media segments for a server computer. The one or more processors are also configured to acquire media segments captured by a plurality of media capture devices located at the location. At least one of the acquired media segments corresponds to a detected trigger. The server computer does not acquire media segments that do not correspond to a trigger. The one or more processors are also configured to determine one or more quality metrics for each of the acquired media segments. The one or more quality metrics for the media segment are determined based on at least one of a first motion of an object captured in the media segment and a second motion of a media capture device used to capture the media segment. The one or more processors are also configured to select a subset of media segments from the acquired media segments, the subset being selected based on the one or more quality metrics determined for each of the acquired media segments. The one or more processors are also configured to generate an aggregate of media segments including the subset of media segments.

[0013] In another example, a non-transitory computer-readable medium for a server computer is provided, the non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: detect a trigger from a device associated with an event at a location, wherein the device does not capture media segments for the system; obtain media segments captured by a plurality of media capture devices located at the location, at least one of the obtained media segments corresponding to the detected trigger, wherein the system does not obtain media segments that do not correspond to the trigger; determine one or more quality metrics for each of the obtained media segments, wherein the one or more quality metrics for the media segments are determined based on at least one of a first motion of an object captured in the media segment and a second motion of a media capture device used to capture the media segment; select a subset of media segments from the obtained media segments, the subset of media segments being selected based on the one or more quality metrics determined for each of the obtained media segments; and generate a collection of media segments including the subset of media segments.

[0014] In some aspects, the trigger is generated in response to user input obtained by the device. In other aspects, the trigger is automatically generated by the device based on at least one or more of the detection of a time during the event and characteristics of the user during the event.

[0015] In some aspects, obtaining a media segment includes: receiving a first media stream captured by a first media capture device among the plurality of media capture devices; receiving a second media stream captured by a second media capture device among the plurality of media capture devices; and extracting a first media segment from the first media stream and extracting a second media segment from the second media stream.

[0016] In some aspects, the trigger is a first trigger. In these aspects, the methods, apparatus, and computer-readable media described above may further include detecting a second trigger associated with an event at the location. In these aspects, a first media segment corresponds to the first trigger, and a second media segment corresponds to the second trigger. In some aspects, obtaining a media segment includes: sending a first trigger to a first media capture device among the plurality of media capture devices, the first trigger causing the extraction of a first media segment from media captured by the first media capture device; sending a second trigger to at least one of the first media capture device and a second media capture device among the plurality of media capture devices, the second trigger causing the extraction of a second media segment from media captured by at least one of the first media capture device and the second media capture device; and receiving the extracted first media segment and the extracted second media segment from at least one of the first media capture device and the second media capture device.

[0017] In some aspects, the methods, apparatus, and computer-readable media described above further include: a server computer (or system) associating the first trigger and the second trigger with the time of the event at that location. In these aspects, the obtained media segment corresponds to the time associated with the first trigger and the second trigger.

[0018] In some respects, the subset of media segments is further selected based on the number of triggers associated with each media segment in the subset of media segments.

[0019] In some respects, the one or more quality metrics of the media segment are further based on the presence of objects in the media segment.

[0020] In some aspects, selecting a subset of media segments from the acquired media segments includes: acquiring a first media segment of an event captured by a first media capture device; acquiring a second media segment of an event captured by a second media capture device, wherein the first and second media segments are captured simultaneously from different perspectives of the event; determining that the motion of one or more objects captured in the first media segment is greater than the motion of one or more objects captured in the second media segment; and selecting the first media segment based on the fact that the motion of the one or more objects captured in the first media segment is greater than the motion of the one or more objects captured in the second media segment.

[0021] In some aspects, the methods, apparatus, and computer-readable media described above also include providing an assembly of media segments to one or more devices.

[0022] In some respects, the device is located at the stated location. In other respects, the device is located at a location remote from the stated location.

[0023] The examples disclosed herein with respect to exemplary methods, apparatuses, and computer-readable media can be implemented individually or in any combination.

[0024] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood through reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0025] The foregoing and other features and embodiments will become more apparent from the following description, claims and drawings. Attached Figure Description

[0026] The illustrative embodiments of this application are described in detail below with reference to the accompanying drawings:

[0027] Figure 1This is a block diagram illustrating an example of a network environment based on some examples;

[0028] Figure 2 This is a diagram illustrating, based on some examples, the locations where events occurred;

[0029] Figure 3 This is a diagram illustrating examples of media systems based on some examples;

[0030] Figure 4 It is a diagram of a scoreboard that provides information for media systems based on some examples;

[0031] Figure 5 This is a diagram showing another example of the location where an event occurred, based on some examples;

[0032] Figure 6 This is a diagram illustrating example use cases of media systems for generating media content aggregations, based on several examples.

[0033] Figure 7 This is a flowchart illustrating an example of the process of generating media content based on some examples; and

[0034] Figure 8 This is a block diagram illustrating an example computing system architecture based on some examples. Detailed Implementation

[0035] Certain aspects and embodiments of this disclosure are provided below. It will be apparent to those skilled in the art that some of these aspects and embodiments can be applied independently, and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The accompanying drawings and description are not intended to be limiting.

[0036] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of exemplary embodiments is intended to provide those skilled in the art with enabling descriptions for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0037] Specific details are set forth in the following description to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0038] Furthermore, please note that the various embodiments can be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart can describe operations as a sequential process, many operations can be performed in parallel or simultaneously. Additionally, the order of operations can be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process can correspond to a method, function, program, subroutine, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.

[0039] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, including, or carrying one or more instructions and / or data. Computer-readable media can include non-transitory media capable of storing data but excluding carrier waves and / or transient electronic signals propagated wirelessly or via wired connections. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions that can represent any combination of procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted by any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0040] Furthermore, embodiments can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., a computer program product) that perform the necessary tasks can be stored on a computer-readable or machine-readable medium. One or more processors can perform the necessary tasks.

[0041] Media capture devices can be used to capture media content (e.g., images, videos, audio, and / or other types of media content). The captured media content can be processed and output for consumption. In some cases, media content from multiple media capture devices can be combined and used to generate one or more media content items with various perspectives of an event or scene (e.g., different field of view, sound, zoom levels, etc.).

[0042] However, problems arise when using media content from multiple media capture devices. For example, capturing and selecting the highest quality and most relevant media clips at different points in time (e.g., during an event) can be challenging. It may be difficult to collect high-quality media content (e.g., video and / or audio, images, or other content) from the best viewpoint at the right time. It may also be difficult to collect and organize video captured from different user devices at different times and angles. This difficulty is exacerbated when a large amount of media content is available for processing. For example, finding the best quality and / or most relevant content among a large amount of video can be challenging, and creating media content compilations (e.g., highlights) using the best quality and / or most relevant content can also be challenging.

[0043] This document discloses techniques for generating media content (e.g., methods or processes implemented by one or more systems, devices, and / or computer-readable media). For example, the techniques disclosed herein can automate the aggregation and curation of media content. In some cases, the techniques disclosed herein can automatically aggregate an aggregation of curated media content into a single media content item. The techniques disclosed herein can reduce the amount of human activity required to generate such media content, thereby reducing the required time and cost. The techniques disclosed herein can reduce the amount of computer resources, such as equipment, storage devices, and processor usage, required to generate an aggregation of media content by minimizing the amount of media content captured and / or stored, by minimizing the amount of analysis required to determine the quality and relevance of the content, etc. The amount of network bandwidth required to upload media content to one or more servers (e.g., one or more cloud servers, one or more enterprise servers, and / or other servers) can also be reduced by aggregating capture requests triggered by multiple users into a single media file delivered to multiple users.

[0044] Figure 1 This is a block diagram illustrating an example of a network environment 100. Network environment 100 includes a media system 102, one or more media capture devices 104, one or more user devices 106, and one or more triggering devices 108. Media system 102 may include one or more server computers capable of processing media data, triggers, and / or other data. References below... Figure 3 Further details of the example media system 302 are described.

[0045] In some embodiments, media system 102 may include a cloud infrastructure system (also referred to as a cloud network) that provides cloud services to one or more media capture devices 104 and one or more user devices 106 (and in some cases, one or more triggering devices 108). For example, one or more media capture devices 104 and / or one or more user devices 106 may install and / or execute applications (e.g., mobile applications or other suitable device applications) and / or websites associated with a provider of media system 102. In this case, the applications and / or websites can access the cloud services provided by media system 102 (via the network). In another example, the cloud network of media system 102 may host applications, and users may subscribe to and use the applications on demand via a communication network (e.g., the Internet, WiFi networks, cellular networks, and / or using another suitable communication network). In some embodiments, the cloud services provided by media system 102 may include hosts of services available on demand to users of the cloud infrastructure system. The services provided by the cloud infrastructure system may be dynamically scaled to meet the needs of its users. The cloud network of media system 102 may include one or more computers, servers, and / or systems. In some cases, the computers, servers, and / or systems that make up a cloud network are different from local computers, servers, and / or systems that may be located at a location (e.g., the location where an event is hosted).

[0046] One or more server computers can communicate with one or more media capture devices 104, one or more user devices 106, and one or more triggering devices 108 using a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network can include any wireless interface or combination of wireless interfaces (e.g., the Internet, cellular networks such as 3G, LTE, or 5G, combinations thereof, and / or other suitable wireless networks). The wired network can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, Ethernet over coaxial cable, digital signal line (DSL), or other suitable wired networks). Various routers, access points, bridges, gateways, etc., that can connect the media system 102, media capture device 104, user device 106, and triggering device 108 to the network can be used to implement wired and / or wireless networks.

[0047] One or more media capture devices 104 may be configured to capture media data. Media data may include video, audio, images, any combination thereof, and / or any other type of media. As used herein, media content items may include video files, audio files, image files, any suitable combination thereof, and / or other media items. The one or more media capture devices 104 may include any suitable type of media capture device, such as a personal or commercial camera (e.g., a digital camera, IP camera, video streaming device, or other suitable type of camera), a mobile or landline telephone handset (e.g., a smartphone, cellular phone, etc.), an audio capture device (e.g., a recorder, microphone, or other suitable audio capture device), a camera for capturing still images, any combination thereof, and / or any other type of media capture device. In one illustrative example, the media capture device in the one or more media capture devices 104 may include a camera (or a device with a built-in camera) that captures video content. In such an example, the camera may also capture audio content synchronized with the video content using known techniques.

[0048] The one or more triggering devices 108 are configured to generate triggers that can be used to determine candidate media segments from the captured media to be considered for inclusion in a media content aggregation. Triggering devices 108 may include an interface that a user can interact with to generate a trigger (e.g., a button, keypad, touchscreen display, microphone, or other suitable interface). For example, a user can press a button on triggering device 108, which can cause triggering device 108 to send a trigger to media system 102 at some point during a media content item.

[0049] The one or more user devices 106 (also referred to as client devices) may include personal electronic devices such as mobile or landline telephone handsets (e.g., smartphones, cellular phones, etc.), cameras (e.g., digital cameras, IP cameras, camcorders, camera phones, video phones, or other suitable capture devices), laptops or notebook computers, tablet computers, digital media players, wearable devices (e.g., smartwatches, fitness trackers, virtual reality headsets, augmented reality headsets, or other wearable devices), video game consoles, video streaming devices, desktop computers, or any other suitable type of electronic device. The one or more user devices 106 may also generate triggers that can be used to determine candidate media segments from the captured media to be considered for inclusion in the media content aggregation. In some cases, the one or more user devices 106 may not provide media content to the media system 102. For example, a user of the one or more user devices 106 may choose not to provide media content to the media system 102 and / or may not be authenticated to provide media content.

[0050] Media system 102 is configured to obtain media from one or more media capture devices 104. In some examples, media system 102 may obtain media from a single source, such as a single storage location (e.g., a file location local to or remote from media system 102, or other storage location), a single media capture device, and / or other sources. In one illustrative example, media may be captured by multiple media capture devices and may be stored in a computer file at media system 102 (or, in some cases, in a remote storage location accessible via a network, for example). Media system 102 may obtain the media as a single media stream by accessing (e.g., opening and executing) the computer file. In another illustrative example, media may be captured by a single media capture device. Media captured by a single media capture device may be stored in a computer file at media system 102 (or, in some cases, in a remote storage location accessible via a network, for example), and media system 102 may access (e.g., open and execute) the computer file to obtain the media. In another illustrative example, a single media capture device can send media content directly (e.g., as streaming media) to media system 102. In some examples, media system 102 can obtain media from multiple sources (e.g., from multiple storage locations (e.g., file locations local to or remote from media system 102, or other storage locations), from media capture devices, and / or other sources). In one illustrative example, media content from one or more media capture devices can be stored in multiple computer files at media system 102 (or in some cases, in a remote storage location accessible via a network, for example), and media system 102 can obtain the media by accessing (e.g., opening and executing) multiple computer files. In another illustrative example, multiple media capture devices can send media content directly (e.g., as streaming media) to media system 102.

[0051] Candidate media segments from captured media can be selected based on one or more triggers. A media segment can include a portion of an entire media content item (e.g., a movie, song, or other media content item). In an illustrative example, a media content item can include video captured by a video capture device, and a media segment can include a portion of the video (e.g., the first ten seconds of the entire video, or other portions of the video). In some cases, media system 102 can receive media captured by one or more media capture devices 104 and can use one or more triggers to select candidate media segments from the received media. In some cases, one or more media capture devices 104 can use one or more triggers to select candidate media segments and can send the selected media segments to media system 102. One or more triggers can include triggers generated by user device 106 and / or triggering device 108 as described above, and / or can include automatically generated triggers. In some cases, one or more media capture devices 104 can generate one or more triggers.

[0052] As described in more detail below, triggers can be automatically generated based on indications of moments of interest that may have occurred during the event captured on the media. In some examples, triggers can be generated based on the detection of an event (or moment) occurring during the event at a location, based on the characteristics of users on devices present at the event, based on the characteristics of users remotely observing the event, any combination thereof, and / or other indications of significant events occurring during the event.

[0053] Media system 102 can evaluate candidate media segments to determine which media segments should be included in the media content collection. As described in more detail below, one or more quality metrics can be determined for a media segment and used to determine whether that media segment will be included in the media content collection. The quality of a media segment can be determined based on factors that indicate the degree of interest those media segments might have for potential viewers. In some cases, the quality metrics determined for media segments can be compared to determine which media segments will be included in the media content collection. For example, the quality metrics determined for a first media segment can be compared with the quality metrics determined for a second media segment, and the first or second media segment with higher quality based on that comparison can be selected for inclusion in the media content collection.

[0054] In some examples, a media content aggregation may include a selected set of media segments that are combined, resulting in multiple shortened clips of media content captured from different media capture devices. The group of selected media segments may have different characteristics, which may be based on the characteristics or settings of the media devices when capturing different media segments. For example, taking video and audio as examples, the media segments in the media content aggregation may come from different points in time, from different viewpoints of a location, have different zoom levels, have different display and / or audio characteristics, may have combinations of these, and / or may have any other suitable variations between the media segments in the aggregation. In some examples, a media content aggregation may include one or more separate, individual media clips. For example, one or more video segments captured by a video capture device may be selected as video clips to be provided for viewing on one or more devices.

[0055] Some of the examples provided herein use video as illustrative examples of media content. However, those skilled in the art will recognize that the systems and techniques disclosed herein are applicable to any type of media content, including audio, images, metadata, and / or any other type of media. Furthermore, in some cases, video content may also include audio content synchronized with the video content using known techniques. The examples disclosed herein may be used alone or in any combination.

[0056] In some cases, it is possible to capture media content during an event at a location. Figure 2 This is a diagram showing an example of location 200 where an event occurred. Events can include sporting events (such as...). Figure 2 This includes events such as performance arts events (e.g., music events, theatrical events, dance events, or other performing arts events), art exhibitions, or any other type of event. Multiple media capture devices can be arranged at different locations within location 200. The media capture devices can include any suitable type of media capture device. For example, as shown, three video capture devices 204a, 204b, and 204c are located at three different locations within location 200. Although... Figure 2 The image shows multiple media capture devices, but in some cases, a single media capture device can provide media capture of events occurring at location 200.

[0057] Video capture devices 204a, 204b, and 204c can capture video of an event from different viewpoints (also referred to as fields of view or angles of view) within location 200. In some embodiments, video capture devices 204a, 204b, and 204c can capture video of an event continuously (e.g., from a point in time before the event begins to a point in time after the event ends). In some embodiments, video capture devices 204a, 204b, and 204c can capture video of an event on demand. For example, one or more of video capture devices 204a, 204b, and 204c may only capture video of an event when receiving data from a media system (e.g., Figure 2 Video is captured only when a command to capture video is received from a user device (e.g., user device 206a, 206b, or 206c), a triggering device, any other suitable device, an interface of the video capture devices 204a, 204b, and 204c, or when input is provided by a user using any other command. In an illustrative example, the video capture devices may capture video upon receiving a trigger, input command, and / or other input, as described in more detail herein. In some cases, one or more of the video capture devices 204a, 204b, and 204c may continuously capture video of an event while simultaneously serving on-demand video clip requests to the media system. For example, captured video footage may be stored locally on the video capture devices 204a, 204b, and 204c, while on-demand video clips (e.g., clips determined based on one or more triggers) may be transmitted to a server and, in some cases, made available to a user registered and / or authenticated to a service provided by the media system and associated with the event. For example, as described in more detail below, video clips (e.g., as standalone video clips, in a compilation of media content such as highlights, in a set of selected video clips, etc.) are available to the user who captured the video clips, and also to other registered users associated with the event.

[0058] User equipment 206a, 206b, and 206c may also be located at location 200. User equipment 206a, 206b, and 206c may include devices used by people attending the event, devices used by people remotely observing the event (e.g., watching on a television, mobile device, personal computer, or other device), or other suitable devices communicating with a media system (not shown). Each of user equipment 206a, 206b, and 206c may include any suitable electronic device.

[0059] Despite Figure 2The diagram shows three video capture devices 204a, 204b, and 204c, and three user devices 206a, 206b, and 206c; however, those skilled in the art will recognize that more or fewer video capture devices and / or user devices may be located within location 200. In some embodiments, one or more triggering devices (not shown) may also be used by a user at the event location and / or by a user remotely observing the event.

[0060] Using information from video capture devices 204a, 204b, and 204c and user devices 206a, 206b, and 206c, the media system can generate a media content aggregation. The media content aggregation may include a set of video clips (videos captured by video capture devices 204a, 204b, and 204c) providing different views of an event at different points in time. For example, the media content aggregation may provide a highlight reel of an event. Any given point in time during the event may also have different views of that point in time. In some cases, multiple views of a given point in time may be included in the media content aggregation (e.g., based on the view with the highest quality, as described below). In some cases, a single view of a given point in time, determined to have the highest quality, may be included in the media content aggregation.

[0061] Figure 3 This is a diagram illustrating an example of media system 302. Media system 302 can be used with events hosted (e.g., in...). Figure 2 The media system 302 communicates with one or more media devices 304 and one or more user devices 306 (and in some cases, one or more triggering devices 308) at the location of the event hosted at location 200 shown. The media system 302 can also communicate with devices (not shown) of users remotely observing the event. The media system 302 includes various components, including a trigger detection engine 312, a media segment analysis engine 314, a quality measurement engine 316, and a media aggregation generation engine 318. Components of the media system 302 may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits); and / or components of the media system 302 may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations disclosed herein. Although the media system 302 is shown as including certain components, those skilled in the art will recognize that the media system 302 may include more than […]. Figure 3The components shown may include more or fewer components. For example, in some cases, media system 302 may also include one or more memory devices (e.g., RAM, ROM, cache, buffer, etc.), processing devices, one or more buses, and / or components not shown in the diagram. Figure 3 Other devices shown in the image.

[0062] In some cases, media system 302 may also include an authentication engine (not shown). The authentication engine can authenticate users of one or more user devices 306 (and / or remotely observing the event), which are located at the site hosting the event and will not provide video (and / or other media) to media system 302. Authentication can be performed on users of the one or more user devices 306 (which will not provide video of the event to media system 302). The authentication engine can also authenticate one or more media capture devices 304 (and / or users of one or more media capture devices 304) at the event site and which will provide video (and / or other media) of the event to media system 302.

[0063] Users of one or more user devices 306 and users of one or more media capture devices 304 can be authenticated through an application associated with the provider of media system 302 (installed on devices 304 and / or 306), a website associated with the provider of media system 302 (e.g., the website may be executed by a browser on one or more user devices 306), and / or using any other suitable authentication platform. For example, a user can register for services offered by media system 302 (e.g., through an application, website, or other platform). A user can register for services offered by media system 302 by creating an account specific to the service (e.g., using an application, website, or other platform to select a username and password, provide answers to security questions, enter contact information, and / or provide other information). Upon registration, a user can be associated with login credentials (e.g., username and password) that allow the user to authenticate through media system 302. Authentication is performed when the user provides credentials, so that the user becomes identified as the actual owner of the registered account and is logged into the application, website, or other platform. Once authenticated, activities associated with the user (e.g., content played, captured, created, joined, invited, and other activities) are linked to their account. The user can then retrieve these activities at a later time by re-authenticating with the service (e.g., logging back in). As described below, in some cases, users can also use the service as "anonymous" users.

[0064] Once registered with the services provided by media system 302, users can join and / or be invited to certain events. In one illustrative example, a user can purchase an event ticket, and the ticket information can be used to authenticate the user for that event. In another illustrative example, a user can join an event through an application, website, or other platform associated with the provider of media system 302. When joining or being invited to a specific event, the user can be authenticated for that event. Authenticated users for an event can use their devices (e.g., one or more media devices 304, one or more user devices 306, and / or one or more triggering devices 308) to generate a trigger for the event, access media content provided by media system 302 for that event, and / or provide media content to media system 302 to generate a media content aggregation. In some cases, one or more media capture devices 304 may be personal devices owned by a specific user. In some cases, one or more media capture devices 304 may not be personal devices and therefore may not be owned by a specific user. In some cases, media capture devices that are not personal devices may be pre-registered by media system 302 as authorized media sources.

[0065] In some examples, different types of users can join an event and / or capture its media content. For instance, in addition to authenticated users (as described above), “anonymous” users can also join an event and / or capture its media content. Anonymous users can include users who have not registered with the service or users who have registered with the service but are not logged into the application (and therefore are not authenticated by the service). In some cases, only a subset of characteristics may be available to anonymous users. In some cases, anonymous users may not be able to retrieve their own activity and / or content created for them during subsequent connections to the application or website associated with the provider of media system 302, because they may not know which account such activity and / or content should be paired with.

[0066] Trigger detection engine 312 can detect device-generated or automatically generated triggers. A trigger indicates that a moment of interest occurred during an event. For example, a trigger from a user (via one of one or more user devices 306 or one of one or more media capture devices 304) can indicate a user's request to time-stamp a portion of video (or other media content) related to a moment of interest in the event. Trigger detection engine 312 can receive triggers generated by one or more user devices 306, and in some cases, triggers generated by one or more media capture devices 304 (e.g., when the media capture device is a personal device of a user attending the event). In some cases, triggers are received only from authenticated users who have been authenticated by the authentication engine. In some examples, triggers can be generated by a user device in one or more user devices 306 or by a media capture device in one or more media capture devices 304 in response to user input. User input can be provided via an application associated with the provider of media system 302, installed on the device. User input can also be provided via a website associated with the provider of media system 302. For example, once a user logs into an application or website (and is thus 302 authenticated by the media system), the user can click or otherwise select a virtual button on the application or website's graphical interface. The user device or media capture device can generate a trigger in response to the selection of the virtual button. Although the virtual button is used herein as an example of an optional option for generating a trigger, those skilled in the art will recognize that any other selection mechanism can be used to generate a trigger, such as using a physical button, using one or more gesture inputs, using one or more voice inputs, combinations thereof, and / or other inputs.

[0067] In some cases, the trigger detection engine 312 may also receive triggers from one or more trigger devices 308. As described above, the trigger devices in one or more trigger devices 308 may include an interface that a user can interact with to generate a trigger (e.g., a button, keypad, touchscreen display, microphone, or other suitable interface). For example, a user may press a button on the trigger device (e.g., a physical button, virtual button, etc.), which may cause the trigger device to send a trigger to the media system 302. In another example, a user may open an application or website associated with the provider of the media system 302 that is installed or executed on a user device in one or more user devices 306 and / or a media capture device in one or more media capture devices 304. The trigger device may be associated with or synchronized with an event via the application or website, after which the user can interact with the interface of the trigger device 308 to generate a trigger. For example, the trigger device may (e.g., using a wireless or wired link) connect to a user device in one or more user devices 306 and / or a media capture device in one or more media capture devices 304 running the application or website, and the application or website may facilitate the association of the trigger device with the application or website. The user can then use a triggering device to generate a trigger, which is sent to an application or website (e.g., via a wireless or wired link).

[0068] In some cases, one or more media capture devices 304, one or more user devices 306, and / or one or more triggering devices 308 can be directly or indirectly paired with an event. For example, in an example where the interface for generating the trigger is a software button, the software button can be directly paired with an event by presenting it within the event-view graphical user interface of an application or website (associated with media system 302). In some cases, if the interface for generating the trigger is a hardware button, the hardware button can be indirectly paired with an event via an application or website. For example, the hardware button can be paired with a mobile application running in an event scenario associated with a specific event. In such an example, the hardware button will only work when the application is running the event scenario, in which case the button can work with an application running in the foreground (e.g., when the user opens the application) and / or an application running in the background (e.g., when the device does not currently display the application and / or when the application is closed). Triggers can also be generated by a third-party system (e.g., a scoreboard or other system) that can be indirectly paired with an event. For example, a third-party system can be paired with a "location" object in the system, and the location will be paired with an event for a given date and time. A "location" object is a representation of a place, location, facility, or venue (e.g., a sports field, park, etc.) on media system 302. In some cases, a "location" object can be defined for a permanent installation. In an illustrative example, when the services provided by media system 302 are installed as permanent installations on an ice hockey rink, a "location" object corresponding to each ice rink can be created. Events can then be scheduled to occur on each defined "location" object. In some examples, when a third-party system (e.g., a scoreboard or other system) is paired with a defined "location," the third-party system can be indirectly paired with events because media system 302 starts and stops scheduled events occurring at that "location." As a result, when a scoreboard or other system paired with a location generates a trigger, media system 302 can identify the event that is currently running at that "location" and associate the trigger with that event.

[0069] Media system 302 can detect triggers by interacting with an application or website associated with the provider of media system 302 that is installed and / or running on a device (e.g., a user device in one or more user devices 306 and / or a media capture device in one or more media capture devices 304). For example, an application or website may generate a trigger in response to user input (e.g., on a user device, on a media capture device, and / or on a triggering device) and may send the trigger to media system 302 (e.g., via the Internet, a wireless network, a wired network, and / or other suitable communication networks). In some cases, a trigger generated by a triggering device may be sent to media system 302 without using an application or website installed on or running on the user device and / or media capture device.

[0070] In some implementations, different click patterns can be defined (received through the interfaces of one or more user devices 306, one or more media capture devices 304, and / or one or more trigger devices 308) and associated with different moments during the event. In one illustrative example, a double-click can be associated with a goal scored during a sporting event, and a long click (e.g., longer than 1 second) can be associated with a penalty during a sporting event.

[0071] As described above, the trigger detection engine 312 can also detect automatically generated triggers. Triggers can be automatically generated based on information indicating a moment of interest that may have occurred during the event. In some examples, applications, websites, or other platforms associated with the provider of media system 302 (e.g., user devices installed in one or more user devices 306 or media capture devices in one or more media capture devices 304) can be configured to automatically generate triggers based on the detection of events occurring during the event, based on one or more characteristics of users on devices present during the event, based on one or more characteristics of users remotely observing the event, any combination of these, and / or other indications of significant events occurring during the event. For example, events that may lead to the automatic generation of triggers during an event could include score changes during a sporting event (e.g., detected using scoreboard data or other data as described below), athlete substitutions during a sporting event, significant time changes (e.g., the end of a quarter or half of the event, the end of the event itself, an intermission in a theatrical or musical event, or other significant time changes), penalties during a sporting event, and audio levels during the event (e.g., loud noise such as cheers, buzzers, bells, etc.).

[0072] It can detect in various ways different events that may lead to automatic generation and triggering during an event. For example, as shown below... Figure 4Further described, the scoreboard can be integrated into the provider platform of media system 302 via software integration, hardware integration via a scoreboard interface device (which may include a hardware module) that reads and interprets the scoreboard's output data, and / or interface via another type of interface. Data from the scoreboard can then be analyzed (e.g., in real-time during data generation, with normal latency due to data transmission and processing, or near real-time). Based on the analysis of the data, automatic triggers can be generated when certain conditions are met (e.g., the clock stops, and the home team's score increases by 1 after x seconds, etc.).

[0073] In some examples, audio during an event can be monitored, and automatic triggers can be generated in response to certain patterns, tones, and / or other audio events. In some illustrative examples, triggers can be automatically generated in response to the detection of a golf club hitting a ball, the detection of a baseball bat hitting a ball, the detection of a loud cheer from a crowd, the detection of a referee's whistle, the detection of a keyword of a voice command, and / or in response to other audio events.

[0074] In some examples, automatic triggers can be generated based on the detection of an athlete's jersey number, face, or other identifying features involved in the event. For example, a user can enter the jersey number of a favorite athlete into an app or website. In another example, a user can enter a photo of a favorite athlete. Other identifying features of a favorite athlete can also be entered into the app or website. Video sequences including the athlete can be detected based on the detection of identifying features (e.g., using computer vision techniques such as object and / or face detection to detect jersey numbers or faces). In an illustrative example, during a baseball game, a camera can be positioned behind home plate, with the athlete hitting the ball at home plate in a position where the jersey number is clearly visible and stable. This setup allows a computer vision algorithm to analyze video frames of the batter, identify the jersey number, and generate an automatic trigger based on the identified jersey number, causing the camera to capture video of the batter during the batting incident and send it to media system 302.

[0075] The characteristics of a user of a device located on-site during an event or a user remotely observing the event may include biometric events associated with the user (e.g., changes in blood pressure, temperature, increased sweating, increased and / or rapid breathing, and / or other biometric events), the user's movements (e.g., clapping, jumping, high-fiving, and / or other movements), the user's audio levels (e.g., volume associated with shouting, etc.), any combination thereof, and / or other suitable user characteristics. Biometric sensors (e.g., wearable biometric sensing devices, smartwatches, fitness trackers, user device 306, media capture device 304, or other suitable devices' biometric sensors), motion sensors of the device (e.g., user device 306, media capture device 304, wearable device, or other suitable devices' motion sensors), microphones or other audio sensors of the device (e.g., user device 306, media capture device 304, wearable device, or other suitable devices' microphones), or any other means of measuring the user's characteristics may be used.

[0076] As described above, biometric-based triggers can be generated based on biometric events associated with users of devices present at the event or users remotely observing the event. For example, a user's heart rate, pulse, temperature, respiratory rate, and / or other indicators (e.g., indicators selected by the user) can be detected, and an automatic trigger can be generated to capture the moment that elicits a bodily response when a change corresponding to a predetermined pattern is detected. In some cases, a default pattern can be used initially, and these patterns can be trained for the user's body as more biometric data is obtained (e.g., using a wearable device). In some implementations, biometric-based triggers can be generated based on biometric indicators of other individuals associated with the event, such as athletes participating in the event (e.g., basketball players during a basketball game), coaches of teams participating in the event, or others present at the event or observing it remotely.

[0077] In some examples, an automatically generated trigger can be generated based on the user's clapping pattern. The clapping pattern can be detected by a wearable device (e.g., a smartwatch, fitness tracker, or other device) that can be paired with an application or website associated with the provider of media system 302.

[0078] Media system 302 (e.g., using trigger detection engine 312) can associate detected triggers with events, corresponding times in the events, and / or one or more media capture devices 304 or other media sources. Trigger data for detected triggers can be stored for later use (e.g., for selecting candidate media segments, for transmission to media capture devices to select candidate media segments, and / or for other purposes). For example, media system 302 can associate detected triggers with events and can also associate detected triggers with times in the events corresponding to the time when the trigger was generated. In an illustrative example, the event could include a football match, and a trigger could be generated in response to a user selecting a button on one or more user devices 306 (e.g., a virtual button in the graphical interface of an application associated with the provider of media system 302 installed on the user device) when a goal is scored. Media system 302 can detect the trigger and associate it with the time corresponding to the goal score.

[0079] Media system 302 is also configured to receive media content (e.g., video, audio, images, and / or other media content) from one or more media capture devices 304. In some cases, media system 302 receives media only from certified recording devices at the event location. For example, media system 302 may receive media captured for that event only from media capture devices that have been certified to media system 302 and / or have been certified for the event. In some examples, one or more media capture devices 304 may record and / or buffer the captured media. In such examples, media system 302 may upload all or a portion of the stored media at a later point in time. For example, a media capture device among one or more media capture devices 304 may send the entire media content to media system 302 (e.g., automatically or in response to media system 302 requesting the entire captured media content from a media capture device). In another example, media system 302 may request a specific clip of the captured media content from a media capture device or from multiple media capture devices among one or more media capture devices 304 (e.g., based on one or more detected triggers, as described in more detail below). In some cases, media capture devices may capture media content based solely on a trigger. For example, one or more media capture devices in media capture device 304 may capture video, audio, and / or other media of an event upon receiving a trigger. In an illustrative example, one or more media capture devices in media capture device 304 may capture the media at the time the trigger is generated (e.g., recording video, audio, and / or other media) as well as media for a certain duration (e.g., x seconds) before and / or after the generation of the trigger. The time window in which the media capture device captures the media content (based on the time of the generation of the trigger plus the duration before and / or after) may be referred to herein as the media window size. In some examples, media system 302 may receive media streams (e.g., video streams and audio-video streams, audio streams, etc.) from a media capture device or multiple media capture devices in one or more media capture devices 304, in which case the media capture device or multiple media capture devices may record and / or buffer the media content, or may not record and / or buffer the media content. While media content is being captured by one or more media capture devices 304, the media stream may be provided to media system 302.

[0080] In some examples, device data may also be transmitted by one or more media capture devices 304 to media system 302, either separately from or together with the media content. Device data associated with a media capture device may include device identifiers (e.g., device name, device serial number, or other numeric identifier, or other suitable identifiers), identifiers of the user of the media capture device (if the device has a unique user) (e.g., user's name, user's numeric identifier, or other suitable user identifier), sensor data that can be used to determine the stability of the media capture device (e.g., whether the media capture device 304 is mounted (stable) rather than handheld), and / or any other suitable data. Sensor data may come from any suitable type of sensor in the media capture device, such as an accelerometer, gyroscope, magnetometer, and / or other types of motion sensors. Media system 302 may store device data for later use. In some cases, media system 302 may associate device data with a corresponding time in an event.

[0081] In some implementations, media system 302 may receive additional data from one or more user devices 306, a scoreboard at the location hosting the event, and / or other devices at that location, and may associate the additional data with corresponding times in the event. Media system 302 may also store additional data for later use. Figure 4 This is a diagram illustrating an example of a scoreboard 400 at a sporting event. Scoreboard data from scoreboard 400 can be provided to media system 302. The scoreboard data can be used to detect significant events occurring during the event (e.g., score changes, athlete substitutions, significant time changes such as a quarter, half, or the end of the event itself, penalties, etc.), which can be used to automatically generate one or more triggers.

[0082] In some examples, a camera (e.g., a camera in one or more media capture devices 304 and / or a camera in one or more user devices 306) may be aligned with the scoreboard 400 to capture images or video of the scoreboard 400. In some cases, an application or website associated with the provider of the media system 302 (e.g., installed on and / or performed by the media capture device 304 and / or user device 306) may analyze changes in the scoreboard image to detect significant events occurring during an event. In some cases, the application or website may send images (and / or scoreboard data) of the scoreboard to the media system 302, and the media system 302 may analyze changes in the scoreboard data (or images) to detect significant events occurring during an event. For example, computer vision techniques may be used to analyze the images to determine significant events occurring during an event. In some examples, the scoreboard may be integrated into an application or website associated with the provider of media system 302 (e.g., installed on and / or performed by the media capture device and / or user device). In these examples, events may be synchronized with the scoreboard integrated into the application or website. For example, as described above, triggers may be based on data from a third-party system (e.g., the scoreboard) that may be indirectly paired with the event. In an illustrative example, a third-party system may be paired with a "location" object in the system, and a location may be paired with an event for a given date and time. When an event is active in a given "location," all devices and / or systems paired with that "location" may provide data and / or triggers used in the context of that event.

[0083] In some examples, the scoreboard interface device can connect to the scoreboard and interpret events occurring on the scoreboard, such as clock stop, clock start, clock value, home team score, away team score, etc. In some cases, the scoreboard interface can be a standalone hardware device. The scoreboard interface device can connect to media system 302, user devices in one or more user devices 306, and / or media capture devices in one or more media capture devices 304 via a network, and can relay scoreboard information to these devices. Media system 302, user devices, and / or media capture devices can generate triggers based on the scoreboard information.

[0084] Candidate media segments can be selected from the captured media based on triggers detected by trigger detection engine 312. A media segment can include a portion of an entire media content item captured by one or more media capture devices 304. For example, a media content item can include video captured by a media capture device, and a media segment can include a portion of that video (e.g., the first ten seconds of the entire video, or other portions of the video). The portion of a media content item corresponding to a media segment can be based on the time the trigger was generated. For example, a trigger can be generated at a point in time during an event (e.g., generated by a user device, a media capture device, a triggering device, or automatically). In response to this trigger, a specific portion of the media content item (corresponding to the time the trigger was generated) can be selected as a candidate media segment. For example, a portion of the media content item used as a candidate media segment can include media captured at the time the trigger was generated, plus media captured for a certain duration (e.g., x seconds) before and / or after the time the trigger was generated, resulting in a candidate media segment with a duration approximately equal to that duration. In an illustrative example, a candidate media segment could include media captured at the time the trigger was generated, plus media captured ten seconds before the time the trigger was generated, resulting in a 10-11 second media segment. In another illustrative example, a portion of the media content item used for a candidate media segment may include media captured at the time of generation triggering, plus media captured ten seconds before the time of generation triggering and media captured ten seconds after the time of generation triggering, resulting in a 20-21 second media segment.

[0085] Coordination between the trigger time, the time in the media item (e.g., in the video), and the time in the event can be based on an absolute time value (e.g., provided by the network) and a time reference of a local device (e.g., one or more media capture devices 304, one or more user devices 306, and / or one or more trigger devices 308). For example, in some cases, the trigger time is based on an absolute time value, which may be provided by a network server or other device (e.g., a cellular base station, an internet server, etc.). The absolute time value can be represented by a timestamp. Network-connected devices are designed to provide a reliable and accurate time reference via a network (e.g., a cellular network, a WiFi network, etc.), which is then available for use by the application. The media capture device 304 can receive the absolute time value of the trigger and the media window size (e.g., trigger time - 5s, trigger time + 20s, or other window sizes) from the server and can capture or provide the corresponding media. When the media capture device has an accurate time reference and is also running an application associated with the media system 302, the absolute timestamp provided by the network server can be directly mapped to the local video using the timestamp of the local media capture device.

[0086] In one illustrative example, one or more media capture devices 304 may be smartphones running an application associated with the provider of media system 302. In such an example, there is no lag or minimal lag between the video being captured (and then available to the application) and its absolute time reference provided by the network. In another illustrative example, where the camera is physically separated from the device or server running the application associated with the provider of media system 302, the video needs to be transmitted to the server running the application. In such an example, when the application receives a trigger with an absolute timestamp, the application can use the local server time reference to execute the trigger. However, the application may not automatically compensate for the delay caused by the transmission of video and / or trigger from the camera to the server. Without adjusting for such an offset, the result is that the moment captured in a video clip appears to be captured earlier in time than expected. Techniques for correcting this delay are provided. In one illustrative example, a fixed offset can be input to the server (e.g., manually, automatically, etc.) and applied to all video (and / or other media) from one or more media capture devices 304 for a given event, or to the video on all or more media capture devices 304. For example, when extracting a video clip based on a obtained trigger, the server can apply the offset to all videos. In another illustrative example, a recognition pattern (e.g., QR code, barcode, image, or other recognition pattern) showing an available instantaneous absolute time value can be captured by any of the one or more media capture devices 304, and the server engine of media system 302 can automatically detect the recognition pattern, interpret the corresponding time value of the recognition pattern, calculate the required offset, and define it as the offset of the camera feed source or the offset of all camera feed sources for the corresponding event. As a result, when a trigger is generated, the application will receive the absolute timestamp and the calculated offset value to crop the correct video clip, taking into account the latency specific to a particular media stream.

[0087] In another illustrative example, artificial intelligence applications or algorithms (e.g., machine learning using neural networks) can be used to determine the offset to be applied to one or more media content items (e.g., video) when extracting a media segment based on an acquired trigger. For example, using the techniques described below, the motion of one or more objects captured in a media segment can be analyzed relative to an acquired trigger to obtain the offset between the trigger and the motion that may have generated it. An estimated offset can be generated using the results of analysis of multiple captured media segments from a media stream. For example, if most of the motion captured in a media segment occurs at approximately the same time relative to the corresponding trigger in each captured media segment, that time can be defined as the offset for that media stream. The offset relative to the corresponding trigger can be further analyzed to further improve the accuracy of the offset or to adjust for drift in the offset for a specific media stream.

[0088] In some implementations, media system 302 may receive media captured by one or more media capture devices 304 (e.g., as a media stream, or after media has been stored by one or more media capture devices 304), and may extract candidate media segments from the received media using detected triggers. For example, media segment analysis engine 314 may extract media segments (e.g., video segments from captured video) from media received from one or more media capture devices 304, wherein the extracted video segments correspond to triggers from one or more trigger sources (e.g., triggers generated by one or more user devices 306, triggers generated by one or more media capture devices 304, triggers generated by one or more trigger devices 308, and / or automatically generated triggers).

[0089] As described above, media system 302 can associate a detected trigger with a time in an event corresponding to the time when the trigger was generated. Media segment analysis engine 314 can extract portions of the received media based on the time associated with the trigger. For example, the time associated with the trigger can be used by media segment analysis engine 314 to identify portions of the received media to be extracted as candidate media segments. Using the example above, a portion of the media content to be selected as a candidate media segment could include media captured at the time the trigger was generated, as well as media captured for a certain duration (e.g., x seconds) before and / or after the time the trigger was generated, resulting in a candidate media segment with a length approximately equal to that duration.

[0090] In some implementations, media system 302 may transmit detected triggers to one or more media capture devices 304 and may receive media segments corresponding to the triggers. In such implementations, one or more media capture devices 304 may use the triggers to select candidate media segments. For example, media system 302 may send one or more detected triggers to some or all of one or more media capture devices 304, and a media capture device receiving a trigger from media system 302 may use the trigger to select candidate media segments. In some cases, media system 302 may also send the time associated with the trigger (e.g., a first time associated with a first trigger, a second time associated with a second trigger, etc.). One or more media capture devices 304 may select candidate media segments in the same manner as described above for media system 302. For example, a media capture device in one or more media capture devices 304 may use the time associated with the trigger to identify portions of the received media to be extracted as candidate media segments. Once a media capture device has selected one or more candidate media segments (based on one or more received triggers), the media capture device may send the selected one or more media segments to media system 302 for analysis. If multiple media capture devices extract candidate media segments from different media content items (e.g., different videos) captured by different media capture devices, all media capture devices can send their respective candidate media segments to media system 302 for analysis.

[0091] Media system 302 can associate received media content and associated device data with corresponding events. For example, media system 302 can store video content received by the media capture device along with the device data of the media capture device. The video content and device data can be associated with the event of video capture. As described above, device data may include a device identifier, an identifier of the user of the device (if the device has a unique user), sensor data that can be used to determine the stability of the media capture device, any combination thereof, and / or any other suitable data.

[0092] Media system 102 can evaluate candidate media segments to determine which segments should be included in the media content aggregate. For example, quality metric engine 316 can analyze candidate video segments to determine their quality. Quality data can be associated with candidate video segments (e.g., quality data for a first video segment can be stored in association with the first video segment, quality data for a second video segment can be stored in association with the second video segment, and so on). The highest quality candidate media segments can be selected for inclusion in the media content aggregate.

[0093] Quality data for media clips can be based on one or more quality metrics determined for the media clip. These quality metrics can be based on factors indicating the degree of interest the media clip might have for a user. An example of a quality metric for a video clip (or other media clip) is the amount of motion of one or more objects captured in the media clip. For example, media capture device 304 can capture video of a sporting event, including athletes on the field. During a sporting event, athletes may move on the field. Candidate video clips may include a portion of the captured video (e.g., a ten-second clip of the captured video). Candidate video clips may include multiple video frames in which athletes move between frames. The amount of motion of the athletes in the video clip can be determined and used as a quality metric. For example, more motion (e.g., above a motion threshold) is associated with higher quality because more motion indicates that the video has captured one or more objects of interest (e.g., athletes). No motion or little motion (e.g., below a motion threshold) may be associated with lower quality because no motion or very little motion indicates that the video has not captured one or more objects of interest. For example, the absence of motion in a video clip can indicate the absence of an athlete in the video clip. In this case, moments of interest during the sporting event (such as those indicated by triggers that cause the video clip to be extracted as candidate video clips) were not captured by the video clip.

[0094] The amount of motion of one or more objects in a video clip can be determined by the quality metric engine 316 using any suitable technique. For example, optical flow between frames of a candidate video clip can be used to determine the movement of pixels between frames, which can indicate the amount of motion occurring in the video clip. In some cases, an optical flow map can be generated at each frame of the video clip (e.g., starting from the second frame of the video clip). The optical flow map can include a vector for each pixel in the frame, each vector indicating the movement of the pixel between frames. For example, optical flow can be calculated between adjacent frames to generate optical flow vectors, and these optical flow vectors can be included in the optical flow map. Each optical flow map can include a two-dimensional (2D) vector field, where each vector is a displacement vector indicating the movement of a point from the first frame to the second frame.

[0095] The quality metric engine 316 can use any suitable optical flow process to determine the motion or movement of one or more objects in a video clip. In an illustrative example, a pixel I(x,y,t) in frame M of a video clip may move a distance (Δx, Δy) in the next frame M+t after a time interval Δt. Assuming the pixels are identical and the intensity remains constant between frame M and the next frame M+t, the following assumptions can be made:

[0096] I(x, y, t)=I(x+Δx, y+Δy, t+Δt) Formula (1).

[0097] By approximating the right side of equation (1) using a Taylor series, then removing the common terms and dividing by Δt, the optical flow equation can be derived as follows:

[0098] f x u+f y v+f t =0, Equation (2),

[0099] in:

[0100]

[0101]

[0102]

[0103] as well as

[0104]

[0105] The image gradient f can be obtained using the optical flow equation (Equation (2)). x and f y and the gradient over time (denoted as f) t The terms u and v are the velocity or x and y components of the optical flow of I(x, y, t), and are unknown. The optical flow equation cannot be solved due to the presence of two unknown variables; in such cases, any suitable estimation technique can be used to estimate the optical flow. Examples of such estimation techniques include difference methods (e.g., Lucas-Kanade estimation, Horn-Schunck estimation, Buxton-Buxton estimation, or other suitable difference methods), phase correlation, block-based methods, or other suitable estimation techniques. For example, Lucas-Kanade assumes that the optical flow (the displacement of an image pixel) is small and approximately constant in a local neighborhood of pixel I, and uses the least squares method to solve the fundamental optical flow equation for all pixels in that neighborhood.

[0106] Other techniques can also be used to determine the amount of motion of one or more objects in a video clip. For example, color-based techniques can be used to negatively determine motion. In an illustrative example, if media capture device 304 is filming one side of an ice rink, and that side is empty because athletes are on the other side, the quality metric engine 316 can identify that white is dominant in these video frames. Based on the detected dominance of white, the quality metric engine 316 can determine that there is no motion (or very little motion) in the video being captured. The amount of white (e.g., based on pixel intensity values) can be compared to a threshold to determine whether white is dominant, thereby determining whether motion is present. In an illustrative example, the threshold can include a threshold number of pixels with white pixel values ​​(e.g., a value of 255 in the pixel value range of 0-255). For example, if the amount of white is greater than the threshold, it can be determined that there is no motion or very little motion in the scene. When the amount of white is less than the threshold, the quality metric engine 316 can determine that motion is present in the video being captured. Although white is used as an example of color, any suitable color (e.g., based on the type of event) can be used. For example, for other types of events (e.g., football, rugby, hockey, lacrosse, etc.), the color can be set to green.

[0107] In some examples, the motion of one or more objects captured in a media clip can be quantified as a motion score. The motion score can be a weighted motion score. For example, when extracting a media clip from a video that predates a certain duration (e.g., x seconds) before the trigger is generated, lower weights can be used for earlier portions of the clip, and higher weights for later portions, because most of the action corresponding to the trigger typically occurs at the end of the clip, which is closer to the trigger's generation time. In such an example, motion detected in earlier portions of the video clip would be given less weight than motion detected in later portions. In an illustrative example, a video clip could include 300 frames (denoted as frames 1-300, corresponding to a ten-second video clip at 30 frames per second). A weight of 0.15 can be assigned to each frame in frames 1-100, a weight of 0.35 can be assigned to each frame in frames 101-200, and a weight of 0.5 can be assigned to each frame in frames 201-300. A motion score can be calculated for this video clip by multiplying the detected motion between frames by the weights that change over time. In another example, when a media segment is captured a certain duration (e.g., x seconds) before the generation trigger time and a certain duration (e.g., x seconds) after the generation trigger time, lower weights can be used for the beginning and end of the media segment, while higher weights can be used for the middle of the media segment, which is closer to the generation trigger time.

[0108] Another example of a quality metric used for video clips (or other media clips) is the motion of one or more media capture devices 304 while capturing video clips. For example, the prioritization of media capture devices in a media content aggregation could be based on the stability of the media capture devices. As mentioned above, the device data stored for the media capture devices can include sensor data that can be used to determine the stability of the media capture devices. Sensor data can be used to determine whether the media capture device is stable (e.g., due to being mounted on a tripod or other support, as opposed to handheld). The stability of the media capture devices can be used in conjunction with a quality metric based on the motion of one or more objects in the video clip, because a handheld media capture device (e.g., a mobile phone with a camera) may shake while shooting video, which would also result in a high motion score. Stability-based prioritization can help offset such problems. For example, if it is determined that the media capture device is stable, and it is determined that a video clip extracted from the video captured by the media capture device has a high motion score, then the quality metric engine 316 can determine that the motion detected for that media clip is the actual motion of the objects, rather than the movement of the media capture device itself.

[0109] In some cases, a stability threshold can be used to determine whether a media capture device is stable or unstable. For example, sensor data can be obtained by the quality metric engine 316 from one or more sensors of the media capture device (e.g., accelerometers, gyroscopes, magnetometers, combinations thereof, and / or another type of sensor). In some cases, an application associated with a provider of media system 302 can obtain sensor data from the media capture device and can send the sensor data to media system 302. In an illustrative example, sensor data may include pitch, roll, and yaw, which correspond to rotations about the x, y, and z axes in three-dimensional space. In some implementations, sensor data corresponding to a time period of a media segment can be obtained (e.g., readings from one or more sensors within the time period of the media segment). The quality metric engine 316 can compare the sensor data to a stability threshold to determine whether the media capture device is stable or unstable. If the sensor data indicates movement below the stability threshold (indicating little to no movement), the media capture device can be determined to be stable. On the other hand, if the sensor data indicates movement above the stability threshold (indicating more movement than expected), the media capture device can be determined to be unstable. Compared to other media capture devices determined to be unstable (e.g., determined based on sensor data), stable media capture devices are associated with higher quality. Media capture devices determined to be stable (and therefore of higher quality) will be preferred over those considered unstable.

[0110] In some cases, device sensors can expose real-time data that is interpreted by the application in real time (e.g., as the data is generated and received) and converted into scores (e.g., motion scores). When a trigger is executed, the application can retrieve the latest score value (e.g., motion score value), verify whether the score is above or below a quality threshold, and upload the corresponding video clip with an associated quality or priority tag. In some cases, the tags may include the score. When the server application processes all received videos for sorting and subsequently processes media content aggregations (e.g., highlights), the server application can read these tags to prioritize each video.

[0111] Another example of a quality metric used for a media segment (e.g., a video segment or any other media segment) is the number of triggers associated with that segment. For example, a first video segment corresponding to a larger number of triggers compared to a second video segment indicates that the first video segment is more interesting than the second. In such an example, the first video segment could be considered to have higher trigger-based quality than the second video segment.

[0112] Another example of quality metric is based on the presence of a specific object in a media clip. Such an object may be referred to herein as an object of interest. For example, an athlete in a particular sporting event may be identified as an object of interest, and media clips with media associated with that athlete may be prioritized over media clips not associated with that athlete. In an illustrative example, video clips that include the athlete during the event may be prioritized over other video clips that do not include the athlete. In some cases, the object of interest may be determined based on user input. For example, a user may provide input (e.g., using a user device in one or more user devices 306, a media device in one or more media devices 304, or other devices) indicating that the user is interested in a specific object. In response to user input, the quality metric engine 316 may tag or otherwise identify media clips that include the object of interest. In some cases, the object of interest may be identified automatically. For example, an athlete in a team may be predefined as an object of interest, and media clips with that athlete may be assigned a higher priority than media clips that do not include that athlete.

[0113] Any suitable technique can be used to identify objects of interest in media clips, such as object recognition for video, audio recognition for audio, and / or any other suitable type of identification or labeling technique. In an illustrative example, an athlete can be designated as the object of interest. Facial recognition can be performed on a video frame to identify the presence of the athlete within the frame. In another example, object recognition can be performed to determine the jersey number of an athlete in a video frame, and the jersey number assigned to the athlete can be identified based on the result of the object recognition. Object recognition can include object character recognition algorithms, or any other type of object recognition that can identify numbers in a video frame.

[0114] The quality metric engine 316 can associate candidate media segments with data from all data sources, such as media segment quality data (e.g., motion, stability, number of triggers, and / or other quality data in the field of view), device data (e.g., user identifiers, motion sensor data, and / or other device data), time-based data (e.g., scoreboard data or other time-based data), and trigger-based data (e.g., user identifier for each trigger, total number of triggers, and / or other trigger-based data).

[0115] As mentioned above, the quality of the identified different media segments can be used to determine which media segments will be included in the media content compilation (e.g., an event highlight). For example, multiple media capture devices 304 may simultaneously capture video of the event. Figure 5 This is an example diagram showing location 500 where an event is taking place. As shown, the event is a football match, and four video capture devices 504a, 504b, 504c, and 504d are arranged at different locations around location 500. Video capture devices 504a, 504b, 504c, and 504d provide different camera angles and viewpoints of the event. [The last sentence appears to be incomplete and possibly refers to a different scenario.] Figure 5 The time point during the football match shown is used to generate a trigger. Based on the generation of this trigger, candidate video clips corresponding to this time point and a certain duration before and / or after the time of the trigger generation can be extracted from the videos captured by all video capture devices 504a, 504b, 504c, and 504d.

[0116] The quality metrics determined for the extracted media clips can be compared to determine which candidate media clips from different video capture devices 504a-504d will be included in the media content collection corresponding to the event (e.g., a highlight reel of the event, a set of video clips, and / or other suitable media collections). For example, as Figure 5As shown, the athlete is located in the lower half of the playing field, which falls within the field of view of two video capture devices 504b and 504d. There is no athlete in the upper half of the playing field, which falls within the field of view of two video capture devices 504a and 504c. Since the athlete is within the field of view of both video capture devices 504b and 504d, the amount of motion captured from the athlete in video clips from video capture devices 504b and 504d will be higher than the motion of any object in video clips from video capture devices 504a and 504c. Consequently, the video clips from video capture devices 504b and 504d will have a higher motion score than the video clips from video capture devices 504a and 504c. Video clips with higher motion scores can be prioritized for use in media content aggregation.

[0117] The stability of video capture devices 504a, 504b, 504c, and 504d can also be determined. For example, sensor data from video capture devices 504a, 504b, 504c, and 504d can be analyzed by a quality metric engine 316, which can determine whether the sensor data indicates movement above or below a stability threshold. In an illustrative example, video capture devices 504b and 504c can be considered stable (e.g., on a tripod or other stabilizing mechanism) based on sensor data indicating movement below a stability threshold, while video capture devices 504a and 504d can be considered unstable (e.g., because the device is being handheld) based on sensor data indicating movement above a stability threshold. Stable video capture devices 504b and 504c can be prioritized for media content aggregation compared to less stable video capture devices 504a and 504d.

[0118] The media aggregation generation engine 318 can then select video clips from video capture devices 504a, 504b, 504c, and 504d for inclusion in the media content aggregation based on motion scores and stability-based priority. For example, since video capture devices 504b and 504c are stable, they will be prioritized over video capture devices 504a and 504d. Since video clips from video capture device 504b have higher motion scores than those from video capture device 504d, the media aggregation generation engine 318 can then select video capture device 504b over video capture device 504c. This will result in high-quality video clips (e.g., with better image stability) associated with the generated triggers.

[0119] Media system 302 can generate a media content aggregate that includes media segments selected based on data associated with candidate media segments. For example, media aggregate generation engine 318 can automatically create aggregates using media segment quality data (based on determined quality metrics). In an illustrative example, media aggregate generation engine 318 can identify a subset of the most relevant video segments based on the total number of triggers, and can select the best video segment from that subset for each trigger based on media segment quality data.

[0120] A media content aggregation can include multiple shortened clips (corresponding to media segments) of captured media content from different media capture devices located around the event. The selected group of media segments can have different characteristics, which can be based on the characteristics or settings of the media capture devices when capturing the media segments. For example, video segments in a media content aggregation can come from different points in time (corresponding to triggers generated throughout the event), can come from different viewpoints on site (corresponding to different media capture devices used to capture the video segments), can have different scaling or other display characteristics, can have different audio characteristics (e.g., based on different audio captured at different points in time), can have any suitable combination thereof, and / or can have any other suitable variations between the media segments in the aggregation.

[0121] In some examples, media content aggregations may include time-based highlights. Time-based highlights can be user-generated and may include media clips corresponding to all triggers generated by that user's device during the event. For example, if a user generates ten triggers by selecting trigger icons ten times in an app installed on their device, a time-based highlight may include ten media clips generated based on those ten triggers. These ten media clips will correspond to moments in the events that occurred around the time the triggers were generated and can be displayed chronologically (e.g., the earliest media clip is displayed first, followed by the next media clip in time, and so on). For each moment in time, the highest quality media clip will be selected for the time-based highlight (determined using the quality metrics described above).

[0122] In some examples, media content aggregation can include event-based highlights. For example, media segments corresponding to the moments with the highest number of triggers during an event can be selected for inclusion in an event-based highlight. In another example, media segments can be selected for inclusion in an event-based highlight based on how many times a user has shared these media segments with other users. For example, a given user can share a media segment with other users by: sending an email to another user with the media segment, sending a text message to another user with the media segment (e.g., using SMS or other messaging platforms), sending a message to another user with the media segment, tagging another user for the media segment, and / or using a sharing mechanism. Media segments that are shared more frequently than other media segments can be selected for inclusion in an event-based highlight. In some cases, media segments can be selected for inclusion in an event-based highlight based on the presence of a specific object of interest within the media segment. For example, an athlete of interest can be identified (e.g., based on user input, the frequency of the object's appearance in the video segment, and / or based on other factors), and the media segment can be analyzed to determine the presence of an object of interest. Compared to other clips that do not include the object of interest, you can choose to include media clips that include the object of interest in the event-based highlight reel.

[0123] In some cases, metadata associated with media clips can also be considered when generating event-based highlights. For example, while a media clip corresponding to a specific moment during an event may have a small number of associated triggers (e.g., only a single trigger, two triggers, etc.), users can input metadata into their devices to indicate that the moment corresponds to a goal during a sporting event. In another example, scoreboard data can be used to determine when a media clip corresponds to a goal in a match. Based on the metadata indicating that the moment corresponds to a goal, the corresponding media clip can be automatically included in the event-based highlights for that match. Any combination of the above factors can be analyzed to determine which video clips should be included in the event-based highlights.

[0124] Once a media content aggregation is generated, media system 302 can then provide the media content aggregation (e.g., a highlight reel) to an authenticated user. For example, an authenticated user with an app associated with the provider of media system 302 installed on their device can access the media content aggregation through the app's interface. In another example, a user can access the media content aggregation through a website associated with the provider of media system 302. In some examples, a user can store and / or share the media content aggregation within the app or website, or share it outside the app or website (e.g., on a social networking site or app or other suitable platform) with other users associated with that user (e.g., the user's "friends").

[0125] Figure 6 This diagram illustrates an example use case of a media system 602 used to generate media content for an event. The event may include a hockey match. Triggers generated by authenticated users at the event location and / or by authenticated users remotely observing the event can be received by media system 602. For example, triggers can be generated by a first trigger source 632 and a second trigger source 634. In some cases, triggers may also be generated automatically during the event, as previously described. Media content can be captured by authenticated media sources at the event location (e.g., media capture device 304) and can be received by media system 602. For example, a first media source 636 and a second media source 638 can be used to capture media content (e.g., video, audio, and / or images) for the event.

[0126] In some cases, to create an event highlight 640, media system 602 can retrieve media clips associated with one or more moments in the event, users, athletes, teams, locations, and / or other characteristics of the event. Media system 602 can select candidate media clips based on triggers generated by a specific user's device (first trigger source 632 or second trigger source 634), triggers generated by devices of all users authenticated for the event (e.g., first trigger source 632 and second trigger source 634), a specific time in the event, objects in the field of view (e.g., specific athletes participating in the event), etc. Media system 602 can then select the best media capture device (e.g., first media source 636 or second media source 638) for each media clip in the highlight. For example, as described above, media clips can be prioritized or ordered based on motion in the field of view of the media capture device used to capture the media clip (e.g., time-weighted), the motion or stability of the media capture device, the presence of specific objects in the media clip, and / or other factors. Highlights 640 may include event-based highlights or time-based highlights. In some cases, event-based highlights may be generated for the event, and time-based highlights may be generated for each authenticated user who causes one or more triggers to occur during the event (e.g., by clicking a trigger button in an application or website associated with the provider of media system 602).

[0127] Media system 602 may also provide media clips 642 corresponding to some or all of the media segments captured during the event. For example, video clip 644 corresponds to a video segment that occurred at 7:31 a.m. during the event. Media clips 642 and highlights 640 may be made available to all users authenticated for the event.

[0128] Using the media generation techniques described above, high-quality media content compilations can be generated for events. For example, the most relevant and highest-quality media clips can be automatically selected for inclusion in event highlights.

[0129] Figure 7 This is a flowchart illustrating an example of a process 700 for generating media content using the techniques disclosed herein. At block 702, process 700 includes (e.g., via a server computer or other device or system) detecting a trigger from a device. The trigger is associated with an event at the scene. For example, in some cases, an event marker may be included in the trigger to indicate which event the trigger is associated with. In some cases, the server computer or other device or system may be part of a media system (e.g., media system 302). The device does not capture media segments for use with the server computer or other device or system. For example, the device may include a user device (e.g., user device 306) or a triggering device (e.g., triggering device 308) that does not provide media content to the server computer (e.g., media system 302). In some examples, the device is located at the scene. In some examples, the device is located at a location remote from the scene. In some examples, the trigger is generated in response to user input obtained by the device. For example, a user may press or otherwise select a virtual button on a graphical interface of an application or website associated with the server computer, and the device may generate a trigger in response to the selection of the virtual button. In other examples, users can select a physical button, perform one or more gesture inputs, speak one or more voice inputs, and / or use another selection mechanism to generate a trigger.

[0130] In some examples, triggers may be automatically generated by the device based on at least one or more of the following: detection at a specific moment during the event and characteristics of the user during the event. For example, moments or events during the event that might lead to the automatic generation of a trigger could include score changes during a sporting event, athlete substitutions during a sporting event, penalties during a sporting event, significant time changes during the event, or audio levels during the event. Characteristics of users of the device located on-site during the event or users remotely observing the event could include biometric events associated with the user, user movement, user audio levels, any combination thereof, and / or other suitable user characteristics.

[0131] At box 704, process 700 includes (e.g., via a server computer or other device or system) obtaining media segments captured by multiple media capture devices located in the field. In some examples, as described above, process 700 may include (e.g., via a server computer or other device or system) obtaining one or more media segments captured by a single media capture device located in the field. At least one of the obtained media segments corresponds to a detected trigger. For example, the at least one media segment may be captured, recorded, and / or transmitted to a server computer in response to a detected trigger. In some cases (e.g., via a server computer or other device or system), media segments that do not correspond to a trigger are not obtained. For example, in some cases (e.g., via a server computer or other device or system), only media segments for which triggers have been generated are obtained.

[0132] In some examples, media segments can be obtained by extracting media segments from media streams received from multiple media capture devices (or from a single media capture device) (e.g., via a server computer or other device or system). For example, obtaining a media segment may include: receiving a first media stream captured by a first media capture device among multiple media capture devices; receiving a second media stream captured by a second media capture device among multiple media capture devices; and extracting a first media segment from the first media stream and a second media segment from the second media stream. In another example, obtaining a media segment may include: receiving a first media stream captured by a first media capture device; receiving a second media stream captured by the first media capture device; and extracting a first media segment from the first media stream and a second media segment from the second media stream.

[0133] In some examples, a media segment can be obtained in response to a trigger sent to a media capture device (e.g., via a server computer or other device or system). For example, in response to receiving a trigger (e.g., from a media system, media capture device, user equipment, and / or triggering device), the media capture device can send a media segment to a server computer (or other device or system). In one illustrative example, the trigger can be a first trigger, in which case process 700 may also include detecting a second trigger associated with an event in the field. In such an example, a first media segment (e.g., obtained by a server computer or other device or system) can correspond to a first trigger, and a second media segment (e.g., obtained by a server computer or other device or system) can correspond to a second trigger. Continuing with this example, obtaining a media segment can include (e.g., via a server computer or other device or system) sending a first trigger to a first media capture device among a plurality of media capture devices. The first trigger causes a first media segment to be extracted from media captured by the first media capture device. The server computer (or other device or system) may also send a second trigger to at least one of a first media capture device and a second media capture device, wherein the second trigger causes a second media segment to be extracted from media captured by at least one of the first and second media capture devices. For example, the second trigger may be sent to the first media capture device, the second media capture device, or both the first and second media capture devices. The first media capture device, the second media capture device, or both the first and second media capture devices may extract the second media segment from the captured media. The server computer (or other device or system) may receive the extracted first media segment and the extracted second media segment from at least one of the first and second media capture devices.

[0134] In some examples, process 700 may (e.g., via a server computer or other device or system) associate the first and second triggers with times in the events on site. In these examples, the obtained media segments correspond to the times associated with the first and second triggers. The times associated with the triggers can be used to determine the specific media segments to be extracted based on those triggers.

[0135] In box 706, process 700 includes (e.g., via a server computer or other device or system) determining one or more quality metrics for each media segment in the acquired media segments. The quality metric for the media segment may be determined based on a first motion of an object captured in the media segment, a second motion of the media capture device used to capture the media segment, or both the first and second motions. In some cases, in addition to or as an alternative to the first and / or second motion, the quality metric for the media segment may be based on the presence of an object in the media segment (e.g., the presence of a person in the media segment, an athlete's jersey number, or the presence of other objects in the media segment).

[0136] In box 708, process 700 includes selecting a subset of media segments from the acquired media segments. The subset of media segments can be selected based on one or more quality metrics determined for each media segment in the acquired media segments. For example, the quality metrics of the different acquired media segments can be analyzed to determine which media segments have the highest quality. In some cases, the subset of media segments is also selected based on the number of triggers associated with each media segment in the subset of media segments.

[0137] In some examples, process 700 can select a subset of media segments from the obtained media segments by obtaining a first media segment of an event captured by a first media capture device and a second media segment of the event captured by a second media capture device. In some examples, the first and second media segments can be captured by the first media capture device. The first and second media segments can be captured simultaneously from different perspectives in the event (e.g., by first and second media capture devices located at different locations). Process 700 can determine that the motion of one or more objects captured in the first media segment is greater than the motion of one or more objects captured in the second media segment. Process 700 can select the first media segment based on the fact that the motion of one or more objects captured in the first media segment is greater than the motion of one or more objects captured in the second media segment. In some cases, as described above, motion scores can be used to select a subset of media segments. For example, media segments with higher motion scores can be preferentially used in media content aggregation.

[0138] In box 710, process 700 includes generating a media segment collection that includes a subset of media segments. In some examples, process 700 may include providing the media segment collection to one or more devices.

[0139] In some examples, the process of generating media content using the techniques disclosed herein can be performed more efficiently than the methods described above. Figure 7The described operations are fewer or more operations. For example, in some examples, the process may include (e.g., via a server computer or other device or system) acquiring one or more media segments captured by a media capture device located on-site or captured by multiple media capture devices. One or more triggers may be acquired (e.g., via a server computer or other device or system), or they may not be acquired. When a trigger is detected, at least one media segment among the acquired media segments may correspond to the detected trigger. Furthermore, when a trigger is acquired, media segments that do not correspond to the trigger may not be acquired. The process may include determining one or more quality metrics for each media segment among the acquired media segments.

[0140] Any quality metric disclosed herein can be determined (e.g., via a server computer or other device or system). In one illustrative example, at least one quality metric of a media segment can be determined based on at least one of a first motion of an object captured in the media segment and a second motion of a media capture device used to capture the media segment. In some cases, the process may include selecting a subset of media segments from the obtained media segments based on one or more quality metrics determined for each media segment of the obtained media segments. For example, as described above, a subset of media segments with the highest quality can be selected from the media segments. The process may include (e.g., via a server computer or other device or system) generating a media segment collection that includes this subset of media segments.

[0141] The above (for example, for) Figure 7 The examples described above (and the other examples above) can be implemented individually or in any combination.

[0142] In some examples, process 700 and / or other processes described herein may be executed by a computing device or apparatus. For example, process 700 may be executed by a server computer or other system disclosed as capable of performing operations for generating media content. In an illustrative example, process 700 and / or other processes described herein may be executed by... Figure 3The media system 302 shown (e.g., one or more server computers or other devices or systems included in media system 302) performs the execution. The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a wearable device, a server (e.g., in a Software as a Service (SaaS) system or other server-based system), and / or any other computing device with the resource capability to perform process 700 and / or other processes described herein. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more output devices, and / or other components of the device configured to perform the steps of process 700 and / or other processes described herein. The computing device may include memory configured to store data (e.g., media data, trigger data, authentication data, and / or any other suitable data) and one or more processors configured to process that data. The computing device may also include one or more network interfaces configured to transmit data. The network interface may be configured to transmit network-based data (e.g., Internet Protocol (IP) based data or other suitable network data). In some embodiments, the computing device may also include a display.

[0143] Components of a computing device can be implemented in circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. A computing device may also include a display (as an example of or supplement to an output device), a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP)-based data or other types of data.

[0144] Process 700 is illustrated as a flowchart or logic flow diagram, whose operations represent a series of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media, which perform the operations when executed by one or more processors. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the operations can be combined in any order and / or in parallel to implement the processing.

[0145] Furthermore, process 700 can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors, hardware, or a combination thereof. As described above, the code can be stored, for example, in the form of a computer program comprising multiple instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.

[0146] Figure 8An architecture of a computing system 800 is illustrated, wherein components of system 800 communicate electrically with each other using connections such as a bus 805. The exemplary system 800 includes a processing unit (CPU or processor) 810 and a system connection 805 that couples various system components, including system memories 815 such as read-only memory (ROM) 820 and random access memory (RAM) 825, to the processor 810. System 800 may include a cache of high-speed memory that is directly connected to, closely adjacent to, or integrated into the processor 810. System 800 can copy data from memory 815 and / or storage device 830 to cache 812 via processor 810 for fast access. In this way, the cache can provide performance improvements, thereby avoiding latency for processor 810 while waiting for data. These and other modules can control or be configured to control processor 810 to perform various actions. Other system memories 815 may also be available. Memory 815 may include various different types of memory with different performance characteristics. Processor 810 may include any general-purpose processor and hardware or software services, such as services 1 832, 2 834, and 3 836 stored in storage device 830, which are configured to control processor 810 and dedicated processors, wherein software instructions are incorporated into the actual processor design. Processor 810 may be a completely independent computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0147] To enable users to interact with computing device 800, input device 845 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 835 can also be one or more of many output mechanisms known to those skilled in the art. In some instances, a multimodal system allows users to provide multiple types of input to communicate with computing device 800. Communication interface 840 typically governs and manages user input and system output. There are no limitations on operation on any particular hardware arrangement, therefore, the basic features here can be readily replaced with improved hardware or firmware arrangements during their development.

[0148] Storage device 830 is a non-volatile memory and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as magnetic tape, flash memory card, solid-state storage device, digital universal disk, cassette tape, random access memory (RAM) 825, read-only memory (ROM) 820 and combinations thereof.

[0149] Storage device 830 may include services 832, 834, and 836 for controlling processor 810. Other hardware or software modules are conceivable. Storage device 830 may be connected to system connection 805. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium in conjunction with necessary hardware components (e.g., processor 810, connection 805, output device 835, etc.) to perform that function.

[0150] For clarity of explanation, in some cases, this technology may be presented as including various functional blocks, including functional blocks that are methods embodied in software or a combination of hardware and software, including devices, device components, steps or routines.

[0151] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0152] The methods according to the examples above can be implemented using computer-executable instructions stored in or otherwise accessible from a computer-readable medium. These instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a particular function or group of functions. Some computer resources may be accessible via a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that can be used to store instructions, used information, and / or information created during the methods according to the examples include disks or optical discs, flash memory, USB devices with non-volatile memory, network storage devices, etc.

[0153] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have generally been described above according to their function. Whether this function is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art can implement the described functions in different ways for each specific application; however, such implementation decisions should not be construed as causing a departure from the scope of this application.

[0154] Devices implementing the methods or processes disclosed herein may include hardware, firmware, and / or software, and may take any of a variety of form factors. Typical examples of such form factors include laptop computers, smartphones, small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. As a further example, such functionality may also be implemented on circuit boards in different chips or different processes executed in a single chip.

[0155] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, handheld wireless communication devices, or integrated circuit devices with multiple uses, including applications in handheld wireless communication devices and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) like synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively, the technology may be implemented at least in part through a computer-readable communication medium that carries or transmits program code in the form of instructions or data structures (e.g., propagated signals or waves) that can be accessed, read, and / or executed by a computer.

[0156] The program code can be executed by a processor, which may include one or more processors (e.g., one or more digital signal processors (DSPs)), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternative embodiments, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor), multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0157] Instructions, the medium for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are the means for providing the functionality described in these disclosures.

[0158] Although various examples and other information are used to interpret aspects of the scope of the appended claims, the specific features or arrangements in such examples should not imply any limitation on the claims, as those skilled in the art will be able to derive a wide variety of implementations from these examples. Furthermore, although some subjects may have been described in language specific to structural features and / or method steps, it should be understood that the subjects defined in the appended claims are not necessarily limited to these described features or actions. For example, such functionality may be distributed differently in or performed in components other than those identified herein. Specifically, the described features and steps are disclosed as examples of components of systems and methods within the scope of the appended claims.

[0159] The claim language, or other language, that states "at least one" and / or "one or more" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language stating "at least one of A and B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of a set" and / or "one or more of a set" does not limit the set to items listed in the set. For example, the claim language stating "at least one of A and B" can mean A, B, or A and B, and can also include items not listed in the set of A and B.

[0160] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (such as microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0161] Those skilled in the art will recognize that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this description.

Claims

1. A method for generating media content, the method comprising: The computing device receives media content and data associated with multiple triggers from multiple user devices configured to capture media data, the multiple triggers being associated with events at locations where the multiple user devices are located; The computing device determines at least one candidate media segment based on the media content received from the plurality of user devices from the plurality of triggers; Based on the delay in receiving the media content from at least one of the plurality of user devices, the offset time of at least one of the plurality of triggers is determined; The computing device determines the media segment duration of the at least one candidate media segment based on the offset time determined for the at least one trigger and a time window relative to the offset time determined for the at least one trigger, the time window including at least one of a first time before the offset time and a second time after the offset time; as well as The computing device generates a media segment having the duration of the media segment based on the offset time determined for the at least one trigger and the time window relative to the offset time.

2. The method of claim 1, wherein the trigger of the plurality of triggers is generated by a user device of the plurality of user devices in response to user input.

3. The method of claim 1, wherein the trigger of the plurality of triggers is automatically generated by a user device of the plurality of user devices based on at least one of the detection of a time during the event and a characteristic of the user during the event.

4. The method of claim 3, wherein the user's characteristics include at least one of biometric events associated with the user, the user's movement, and the user's audio level.

5. The method of claim 3, wherein the moment is detected based on the detection of objects in the media segment.

6. The method of claim 1, wherein the time window is determined based on user input.

7. The method according to claim 1, wherein the time window is automatically determined.

8. The method of claim 7, wherein the time window is automatically determined based on motion associated with the media segment during the time window.

9. The method of claim 1, wherein generating the media segment comprises capturing the media segment by the computing device based on the at least one trigger and the duration of the media segment.

10. The method according to claim 1, further comprising: The media segment is generated at least in part by extracting the at least one candidate media segment based on the at least one trigger and the duration of the media segment.

11. The method of claim 1, wherein the offset time is a fixed offset of media content received from at least one media capture device.

12. The method of claim 1, wherein a machine learning system is used to determine the offset time.

13. The method according to claim 1, wherein, The multiple triggers are associated with multiple media content items, and the method further includes: Based on the multiple triggers, determine one or more media segments from each of the multiple media content items; and Generate a media segment set that includes the one or more media segments.

14. A system for generating media content, comprising: The memory is configured to store media data; as well as One or more processors are coupled to the memory and configured to: Media content and data associated with multiple triggers are received from multiple user devices configured to capture media data, the multiple triggers being associated with events at locations where the multiple user devices are located; Based on the multiple triggers, at least one candidate media segment is determined from the media content received from the multiple user devices; Based on the delay in receiving the media content from at least one of the plurality of user devices, the offset time of at least one of the plurality of triggers is determined; Based on the offset time determined for the at least one trigger and a time window relative to the offset time determined for the at least one trigger, the media segment duration of the at least one candidate media segment is determined, the time window including at least one of a first time before the offset time and a second time after the offset time; as well as A media segment having the duration of the media segment is generated based on the offset time determined for the at least one trigger and the time window relative to the offset time.

15. The system of claim 14, wherein the trigger of the plurality of triggers is generated by a user device of the plurality of user devices in response to user input.

16. The system of claim 14, wherein the trigger of the plurality of triggers is automatically generated by a user device of the plurality of user devices based on at least one of the detection of a time during the event and a characteristic of a user during the event.

17. The system of claim 14, wherein the time window is determined based on user input.

18. The system of claim 14, wherein the time window is automatically determined.

19. The system according to claim 14, wherein, The plurality of triggers are associated with a plurality of media content items, and wherein one or more processors are configured to: Based on the multiple triggers, determine one or more media segments from each of the multiple media content items; and Generate a media segment set that includes the one or more media segments.

20. A non-transitory computer-readable medium for a server computer, the non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: Media content and data associated with multiple triggers are received from multiple user devices configured to capture media data, the multiple triggers being associated with events at locations where the multiple user devices are located; Based on the multiple triggers, at least one candidate media segment is determined from the media content received from the multiple user devices; Based on the delay in receiving the media content from at least one of the plurality of user devices, the offset time of at least one of the plurality of triggers is determined; Based on the offset time determined for the at least one trigger and a time window relative to the offset time determined for the at least one trigger, the media segment duration of the at least one candidate media segment is determined, the time window including at least one of a first time before the offset time and a second time after the offset time; as well as A media segment having the duration of the media segment is generated based on the offset time determined for the at least one trigger and the time window relative to the offset time.