Audio processing for detecting crowd noise in sports event television programs
The system detects crowd noise events through spectrogram analysis to generate metadata for enhanced interactive television applications, addressing the lack of real-time crowd noise detection in existing systems and improving highlight generation in sporting events.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-10
AI Technical Summary
Existing television systems lack the ability to automatically detect and utilize crowd noise events in real-time to enhance interactive television applications and generate metadata for highlights in sporting events.
A system and method for constructing a spectrogram of audio data, identifying spectral magnitude peaks, and generating vectors to detect significant crowd noise events, which are then used to create metadata for enhancing highlight generation in sporting events.
Enables real-time detection and tracking of crowd noise events, allowing for automated metadata generation and enhanced interactive television applications, including synchronized highlight creation.
Smart Images

Figure 2026063043000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority from U.S. Provisional Application No. 62 / 680,955, filed on June 5, 2018, titled "Audio Processing for Detecting Occurrences of Crowd Noise in Sporting Event Television Programming" (Attorney Docket No. THU007 - PROV), the entire content of which is incorporated herein by reference.
[0002] This application claims priority from U.S. Provisional Application No. 62 / 712,041, filed on July 30, 2018, titled "Audio Processing for Extraction of Variable Length Disjoint Segments from Television Signal" (Attorney Docket No. THU006 - PROV), the entire content of which is incorporated herein by reference.
[0003] This application claims priority from U.S. Provisional Application No. 62 / 746,454, filed on October 16, 2018, titled "Audio Processing" for Detecting Occurrences of Loud Sound Characterized by Short - Time Energy Bursts, the entire content of which is incorporated herein by reference.
[0004] This application claims priority from U.S. Utility Application No. 16 / 421,391, filed on May 23, 2019, titled "Audio Processing for Detecting Occurrences of Crowd Noise" in Sporting Event Television Programming, the entire content of which is incorporated herein by reference.
[0005] This application relates to U.S. Utility Model Application No. 13 / 601,915, “Generating Excitement Levels for Live Performances,” filed on 31 August 2012 and issued on 16 June 2015 as U.S. No. 9,060,210, which is incorporated herein by reference in its entirety.
[0006] This application relates to U.S. Utility Model Application No. 13 / 601,927, “Generating Alerts for Live Performances,” filed on 31 August 2012 and issued on 23 September 2014 as U.S. No. 8,842,007, which is incorporated herein by reference in its entirety.
[0007] This application relates to U.S. Utility Model Application No. 13 / 601,933, “Generating Teasers for Live Performances,” filed on 31 August 2012 and issued on 26 November 2013 as U.S. No. 8,595,763, which is incorporated herein by reference in its entirety.
[0008] This application relates to U.S. Utility Model Application No. 14 / 510,481 (THU001), filed on 9 October 2014, for "Generating a Customized Highlight Sequence Depicting an Event," which is incorporated herein by reference in its entirety.
[0009] This application relates to U.S. Utility Model Application No. 14 / 710,438 (THU002), filed on 12 May 2015, for "Generating a Customized Highlight Sequence Depicting Multiple Events," which is incorporated herein by reference in its entirety.
[0010] This application relates to U.S. Utility Model Application No. 14 / 877,691 (THU004), filed on 7 October 2015, for "Customized Generation of Highlight Show with Narrative Component," which is incorporated herein by reference in its entirety.
[0011] This application relates to U.S. Utility Model Application No. 15 / 264,928 (THU005), filed on 14 September 2016, for "User Interface for Interaction with Customized Highlight Shows," which is incorporated herein by reference in its entirety.
[0012] This application relates to U.S. Utility Model Application No. 16 / 411,704 (THU009), filed on 14 May 2019, for "Video Processing for Enabling Sports Highlights Generation," which is incorporated herein by reference in its entirety.
[0013] This application relates to U.S. Utility Model Application No. 16 / 411,710 (THU010), filed on 14 May 2019, for "Machine Learning for Recognizing and Interpreting Embedded Information Card Content," which is incorporated herein by reference in its entirety.
[0014] This application relates to U.S. Utility Model Application No. 16 / 411,713 (THU012), filed on 14 May 2019, for "Video Processing for Embedded Information Card Localization and Content Extraction," which is incorporated herein by reference in its entirety.
[0015] This paper relates to a technology for identifying multimedia content, as well as related information on television devices and video servers that distribute multimedia content, enabling embedded software applications to utilize the multimedia content and provide content and services synchronized with that multimedia content. Various embodiments relate to a method and system for providing automated audio analysis to identify and extract information from television program content depicting a sporting event and create metadata related to video highlights for in-game and post-game viewing. [Background technology]
[0016] Enhanced television applications such as interactive advertising, as well as enhanced program guides with pre-game, in-game, and post-game interactive applications, have long been envisioned. Existing cable systems, originally designed for broadcast television, are now required to support new applications and services, including interactive television services and enhanced (interactive) program guides.
[0017] Several frameworks have been standardized to enable augmented television applications. Examples include the OpenCable® augmented TV application messaging specification and the Tru2way specification, which refer to interactive digital cable services delivered over cable video networks and include features such as interactive program guides, interactive advertising, and games. Furthermore, cable operators' "OCAP" programs offer interactive services such as e-commerce shopping, online banking, electronic program guides, and digital video recording. These efforts have enabled the first generation of video synchronization applications, which are synchronized with video content delivered by program producers / broadcasters, providing additional data and interactivity to television programs.
[0018] Recent advancements in video / audio content analysis technologies and corresponding mobile devices have opened up a range of new possibilities in developing sophisticated applications that operate in sync with live TV program events. These new technologies, along with advancements in audio signal processing and computer vision, and the improved computing power of modern processors, enable the real-time generation of sophisticated program content highlights with metadata currently lacking in television and other media environments. [Overview of the project]
[0019] A system and method are presented for detecting, selecting, and tracking significant crowd noise (e.g., audience cheers) by enabling the automated, real-time processing of audio data, such as audio streams extracted from television program content of sporting events.
[0020] In at least one embodiment, a spectrogram of audio data is constructed, and a significant collection of spectral magnitude peaks is identified at each position in a sliding two-dimensional time-frequency area window. A spectral indicator is generated for each position in the analysis window, forming a vector of spectral indicators with the associated time positions. In a subsequent processing step, runs of selected indicator-position pairs with narrow time intervals are identified as potential events of interest. For each run, the internal indicator values are sorted to obtain the maximum magnitude indicator value with the associated time position. Furthermore, the time position (start / center) and duration (count of indicator-position pairs) are extracted for each run. A preliminary event vector is formed containing a triplet of parameters (M, P, D) representing the maximum indicator value, start / center time position, and run duration for each event. This preliminary event vector is then processed to generate a final crowd noise event vector corresponding to the desired event interval, event loudness, and event duration.
[0021] In at least one embodiment, when cluster noise event information is extracted, it is automatically added to sports event metadata associated with highlights of a sports event video and can then be used in connection with the automatic generation of highlights.
[0022] In at least one embodiment, a method for extracting metadata from an audiovisual stream of an event can include storing audio data extracted from the audiovisual stream in a data store, using a processor to automatically identify one or more portions of the audio data indicative of crowd excitement at the event, and storing in the data store metadata that includes at least a time index indicating the time at which each of the portions occurs within the audiovisual stream. Alternatively, the audio data can be extracted from an audio stream or from previously stored audiovisual content or audio content.
[0023] The audiovisual stream may be a broadcast of an event. The event may be a sports event or another type of event. The metadata may be related to highlights that are considered of particular interest to one or more users.
[0024] The method may further include using an output device to present the metadata during viewing of a highlight by one of one or more users to indicate the level of crowd excitement associated with the highlight.
[0025] The method may further include using the time index to identify the start and / or end of a highlight. As described below, the start and / or end of a highlight can be adjusted based on an offset.
[0026] The method may further include using an output device to present a highlight to one of one or more users during automatic identification of one or more portions.
[0027] This method may further include preprocessing the audio data by resampling the audio data to a desired sampling rate before automatic identification of one or more portions.
[0028] This method may further include preprocessing the audio data by filtering the audio data to reduce or remove noise before automatic identification of one or more portions.
[0029] This method may further include preprocessing the audio data to generate a spectrogram (two-dimensional time-frequency representation) for at least a portion of the audio data before automatic identification of one or more portions.
[0030] Automatically identifying one or more portions may include identifying peaks in the magnitude of the spectrum at each position of a sliding two-dimensional time-frequency analysis window of the spectrogram.
[0031] Automatically identifying one or more portions may further include generating a spectral indicator for each position of the analysis window and using the spectral indicators to form a vector of spectral indicators having associated time portions.
[0032] This method may further include identifying a run of selected pairs of spectral indicators and analysis window positions, capturing the identified run with a set of R vectors, and obtaining one or more maximum magnitude indicators using the set of R vectors.
[0033] This method may further include extracting a time index from each of the R vectors.
[0034] This method may further include generating a preliminary event vector by replacing each R vector with a parameter triplet representing the maximum magnitude indicator, the time index, and the run length of one of the runs.
[0035] This method may further involve processing preliminary event vectors to generate crowd noise event information, including time indices.
[0036] Further details and variations are described herein. [Brief explanation of the drawing]
[0037] The attached drawings, along with their descriptions, illustrate several embodiments. Those skilled in the art will recognize that the specific embodiments shown in the drawings are merely illustrative and not intended to limit the scope.
[0038] [Figure 1A] This block diagram shows a hardware architecture in a client / server configuration where event content is delivered via a network-connected content provider. [Figure 1B] This block diagram shows a hardware architecture in another client / server configuration where event content is stored on a client-based storage device. [Figure 1C] This block diagram shows the hardware architecture in a standalone embodiment. [Figure 1D] This block diagram shows an overview of a system architecture according to one embodiment. [Figure 2] This schematic block diagram shows an example of a data structure that can be incorporated into the audio data, user data, and highlight data in Figures 1A, B, and 1C, according to one embodiment. [Figure 3A]An example of an audio waveform graph showing the occurrence of crowd noise events (e.g., crowd cheers) in an audio stream extracted from television program content of a sports event in the time domain, according to one embodiment, is shown. [Figure 3B] An example of a spectrogram corresponding to the audio waveform graph in Figure 3A in the time-frequency domain, according to one embodiment, is shown. [Figure 4] This flowchart shows a method for performing on-the-fly processing of audio data to extract metadata, according to one embodiment. [Figure 5] This flowchart shows a method, according to one embodiment, for analyzing audio data in the time-frequency domain to detect clustering of spectral magnitude peaks associated with long-duration crowd cheers. [Figure 6] This is a flowchart showing a method for generating crowd noise event vectors according to one embodiment. [Figure 7] This is a flowchart showing a method for internal processing of each R vector according to one embodiment. [Figure 8] This flowchart shows a method for further selecting a desired crowd noise event according to one embodiment. [Figure 9] This flowchart shows a method for further selecting a desired crowd noise event according to one embodiment. [Figure 10] This flowchart shows a method for further selecting a desired crowd noise event according to one embodiment. [Modes for carrying out the invention]
[0039] definition The following definitions are provided for illustrative purposes only and are not intended to limit the scope. ●Events: In this description, the term “event” refers to a game, session, match, series, performance, program, concert, etc., or a part thereof (act, period, quarter, half, inning, scene, chapter, etc.). An event may be a sporting event, an entertainment event, or a specific performance by a single individual or a subset of individuals within a larger group of participants in an event. Examples of non-sporting events include television shows, news flashes, socio-political events, natural disasters, films, plays, radio shows, podcasts, audiobooks, online content, and musical performances. An event can be of any length. For illustrative purposes, this technology is often described herein in terms of sporting events. However, those skilled in the art will recognize that this technology can also be used in other contexts, including highlight shows for audiovisual, audio, visual, graphics-based, interactive, non-interactive, or text-based content. Accordingly, the use of the term “sporting event” and other sport-specific terms in this description is intended to illustrate one possible embodiment, but not to limit the scope of the technology described to that one embodiment. Rather, such terminology should be considered to extend to appropriate contexts beyond sports, as is appropriate for this technology. For the sake of clarity, the term “event” is also used to refer to accounts or representations of an event, such as audiovisual recordings of the event, or other content items that include accounting, explanation, or depiction of the event. ● Highlights: An excerpt or portion of an event, or an excerpt or portion of content related to an event that is considered to be of particular interest to one or more users. Highlights can be of any length. In general, the technologies described herein provide a mechanism for identifying and presenting a customized set of highlights (which may be selected based on specific characteristics and / or user preferences) for any suitable event. "Highlights" may also be used to refer to an account or representation of a highlight, such as an audiovisual recording of the highlight, or other content items that include an accounting, explanation, or depiction of a highlight. Highlights do not have to be limited to a depiction of the event itself, but may include other content related to the event. For example, in the case of a sporting event, highlights may include audio / video during the game, as well as other content such as pre-game, in-game, and post-game interviews, analysis, and commentary. Such content may be recorded from linear television (for example, as part of an audiovisual stream depicting the event itself) or obtained from any number of other sources. Various types of highlights can be provided, including, for example, occurrences (plays), strings, possessions, and sequences, all of which are defined below. Highlights do not need to have a fixed duration, but they can incorporate start and / or end offsets, as described below. ● Clip: A portion of an event's audio, visual, or audiovisual representation. A clip may correspond to or represent a highlight. In many contexts herein, the term "segment" is used interchangeably with "clip." A clip may be a portion of an audio stream, video stream, or audiovisual stream, or a portion of stored audio, video, or audiovisual content. ● Content delineator: One or more video frames that indicate the start or end of a highlight. ● Occurrence: Something that happens during an event. Examples include goals, plays, downs, hits, saves, shots on goal, baskets, steals, snaps or attempted snaps, near misses, fights, the start or end of a game, quarters, halves, periods or innings, pitches, penalties, injuries, dramatic incidents in entertainment events, songs, solos, etc. Occurrences can also be unusual, such as power outages or incidents involving unruly fans. The detection of such occurrences can be used as a basis for determining whether or not to designate a particular portion of an audiovisual stream as a highlight. Occurrences are also referred to as “plays” herein for ease of naming, but such usage should not be interpreted as limiting the scope. Occurrences may be of any length, and expressions of occurrences can vary in length. For example, as above, an extended expression of an occurrence may include scenes describing the time period immediately before and after the occurrence, while a simple expression may include only the occurrence itself. Any intermediate expressions can also be provided. In at least one embodiment, the choice of duration for representing an occurrence may depend on user preference, available time, determined level of excitement for the occurrence, importance of the occurrence, and / or any other arbitrary factors. ●Offset: An amount that adjusts the length of the highlight. In at least one embodiment, a start offset and / or end offset can be provided to adjust the start time and / or end time of the highlight, respectively. For example, if the highlight depicts a goal, the highlight may be extended by a few seconds (via the end offset) to include the cheers and / or fan reaction that follow the goal. The offset can be configured to change automatically or manually based, for example, the amount of time available for the highlight, the importance and / or excitement level of the highlight, and / or other preferred factors. ● String: A series of occurrences that are linked or related to each other in some way. Occurrences may occur within a possession (as defined below) or span multiple possessions. Occurrences may occur within a sequence (as defined below) or span multiple sequences. Occurrences may be linked or related because they have some thematic or narrative connection to each other, or because one leads to another, or for other reasons. An example of a string is a set of passes that lead to a goal or basket. This should not be confused with "text string," which has the meaning usually assigned to it in computer programming techniques. ●Possession: A time-bound portion of an event. The start / end time boundaries of possession may depend on the type of event. In certain sporting events where one team may be offensive and the other defensive (e.g., basketball or soccer), possession can be defined as the period of time when one team possesses the ball. In sports where possession of the puck or ball is more fluid, such as hockey or soccer, possession can be considered to extend to the period of time when one team can substantially control the puck or ball, ignoring momentary contact by the other team (e.g., blocked shots or saves). In baseball, possession is defined as a half-inning. In soccer, possession can include several sequences in which the same team possesses the ball. In other types of sporting events and non-sporting events, the term "possession" may be somewhat misnomer, but it is still used here for illustrative purposes. Examples in non-sporting contexts include chapters, scenes, and acts. For example, in the context of a music concert, possession may correspond to the performance of a single song. Possession can include any number of occurrences. ● Sequence: A time-bound portion of an event that includes a single continuous period of action. For example, in a sporting event, a sequence may begin when an action (such as a face-off or tip-off) starts and end when a whistle is blown to indicate an interruption of the action. In sports such as baseball or soccer, a sequence may be equivalent to a play, which is a form of occurrence. A sequence can contain or be a portion of any number of possessions. ● Highlight Show: A set of highlights arranged for presentation to the user. A highlight show may be presented linearly (e.g., in an audiovisual stream) or it may be presented in a way that allows the user to select which highlights to view and in what order (e.g., by clicking on links or thumbnails). The presentation of a highlight show may be non-interactive or interactive, allowing the user to pause, rewind, skip, fast forward, and communicate their like or dislike preferences. A highlight show may be, for example, a condensed game. A highlight show may contain any number of consecutive or non-consecutive highlights from a single event or multiple events, and may even contain highlights from various types of events (e.g., various sports, as well as / or combinations of highlights from sporting and non-sporting events). ●User / Viewer: The terms “User” or “Viewer” are interchangeable to refer to any individual, group, or other entity viewing, listening to, or otherwise experiencing an event, one or more highlights of an event, or a highlight show. The terms “User” or “Viewer” may also refer to any individual, group, or other entity that, at some future time, may view, listen to, or otherwise experience any of the event, one or more highlights of an event, or a highlight show. The term “Viewer” may be used for descriptive purposes, but since an event does not need to have a visual component, a “Viewer” may instead be a listener or other consumer of the content. ● Excitement Level: A measure of how exciting or interesting an event or highlight is expected to be for a particular user or for users in general. The excitement level can also be determined in relation to a particular occurrence or player. Various methods for measuring or evaluating the excitement level are described in the relevant applications described above. As described, the excitement level may depend on occurrences within the event, as well as other factors such as the overall context or importance of the event (e.g., playoff games, pennant implications, rivalries). In at least one embodiment, the excitement level can be associated with each occurrence, string, possession, or sequence within the event. For example, the excitement level of a possession can be determined based on the occurrences that occur within that possession. The excitement level may be measured differently for different users (e.g., a fan of one team vs. a neutral fan) and may differ depending on each user's personal characteristics. ● Metadata: Data that is related to and stored in association with other data. Primary data may be media such as sports programs or highlights. ● Video data. The length of a video, which may be in digital or analog format. Video data can be stored on a local storage device or received in real time from a source such as a TV broadcast antenna, cable network, or computer server, in which case it may be called a “video stream.” Video data may or may not include audio components, and if it does, it may be called “audiovisual data” or an “audiovisual stream.” ● Audio data. The length of audio, which may be in digital or analog format. Audio data can be the audio component of audiovisual data or an audiovisual stream, and can be separated by extracting audio data from audiovisual data. Audio data can be stored in local storage or received in real time from a source such as a TV broadcast antenna, cable network, or computer server, in which case it may be called an "audio stream". ● Stream. Audio stream, video stream, or audiovisual stream. ● Time Index: An indicator of time within audio, video, or audiovisual data related to an event occurring or otherwise associated with a specified segment such as a highlight. ● Spectrogram. A visual representation of the frequency spectrum of a signal, such as an audio stream, which changes over time. ● Analysis window. A specified subset of video data, audio data, audiovisual data, spectrograms, streams, or other processed versions of streams or data, in which one step of analysis is focused. Audio data, video data, audiovisual data, or spectrograms can be analyzed within segments using, for example, moving analysis windows and / or a series of analysis windows that cover various segments of the data or spectrogram.
[0040] overview According to various embodiments, methods and systems are provided for automatically creating time-based metadata related to highlights of television programs such as sporting events, wherein such video highlights and associated metadata are generated in sync with the television broadcast of the sporting event, or while the video content of the sporting event is being streamed from a storage device through a video server after the television broadcast of the sporting event.
[0041] In at least one embodiment, an automated video highlighting and associated metadata generation application can receive a live broadcast audiovisual stream or a digital audiovisual stream received via a computer server. The application can then process audio data, such as an audio stream extracted from the audiovisual stream, using, for example, digital signal processing techniques, to detect crowd noise, such as crowd cheers.
[0042] In alternative embodiments, the techniques described herein may be applied to other types of source content. For example, audio data does not need to be extracted from an audiovisual stream, but rather may be a radio broadcast or other audio description of a sporting event or other event. Alternatively, the techniques described herein may be applied to stored audio data describing an event, which may or may not be extracted from stored audiovisual data.
[0043] An interactive television application enables a user watching a television program on either a primary television display or a secondary display such as a tablet, laptop, or smartphone to be presented with highlighted television program content in a timely and appropriate manner. In at least one embodiment, a set of clips representing highlights of television broadcast content is generated and / or stored in real time, along with a database containing time-based metadata that describes in more detail the events presented by the highlight clips. As will be described in more detail herein, the start and / or end times of such clips can be determined, at least in part, based on an analysis of extracted audio data.
[0044] In various embodiments, metadata associated with a clip may be any information, such as text information, images, and / or any type of audiovisual data. One type of metadata relevant to both in-game and post-game video content highlights presents events detected by real-time processing of audio data extracted from a television broadcast of a sporting event. In various embodiments, the systems and methods described herein enable automated metadata generation and video highlight processing, and the start and / or end times of highlights can be detected and determined by analyzing digital audio data, such as an audio stream. For example, event information can be extracted by analyzing such audio data to detect cheering crowd noise following a particular exciting event, audio announcement, music, etc., and such information can be used to determine the start and / or end times of highlights.
[0045] In at least one embodiment, real-time processing is performed on audio data, such as an audio stream extracted from television program content of a sporting event, to detect, select, and track significant crowd noise (such as crowd cheers).
[0046] In at least one embodiment, the system and method receive compressed audio data, read the compressed audio data, decode it, and resample it to a desired sampling rate. Pre-filtering can be performed for noise reduction, click removal, and selection of the target frequency band. Any of several interchangeable digital filtering stages can be used.
[0047] Spectrograms can be constructed for audio data. A significant collection of peaks in spectral magnitude can be identified at each position in a sliding two-dimensional time-frequency area window.
[0048] A spectral indicator may be generated for each position in the analysis window, and a vector of spectral indicators with the associated time positions may be formed.
[0049] Runs of selected indicator and position pairs with narrow time intervals are identified, and a set of vectors is generated.
number
[0050] The time position (start / center) and length (duration) of a run (count of indicator-position pairs) can be extracted from each R vector.
[0051] A preliminary event vector can be formed by replacing each R vector with a parameter triplet (M, P, D) representing the maximum indicator value, start / mid-time position, and run-length (duration), respectively.
[0052] By processing preliminary event vectors, a final crowd noise event vector can be generated according to the desired event interval, event loudness, and event duration.
[0053] The extracted crowd noise event information can be automatically added to the sports event metadata associated with the highlights of the sports event video.
[0054] In another embodiment, the system and method perform real-time processing of an audio stream extracted from a television broadcast of a sporting event to detect, select, and track significant crowd noise. This system and method may include capturing television broadcast content, extracting and processing digital audio data such as a digital audio stream to detect significant crowd noise events, generating a time-frequency audio spectrogram, performing a combined time-frequency analysis of the audio data to detect areas of high spectral activity, generating spectral indicators for overlapping spectrogram areas, forming vectors of selected indicator-location pairs, identifying runs of selected indicator-location pairs with narrow time intervals, forming a set of vectors with identified runs, forming at least one preliminary event vector with parameter triplets (M, P, D) derived from each run of selected indicator-location pairs, and modifying at least one preliminary event vector to generate at least one final crowd noise event vector with a desired event interval, event loudness, and event duration.
[0055] Initial preprocessing of the decoded audio data may be performed for at least one of the following: noise reduction, removal of clicks and other spurious sounds, and selection of frequency bands of interest by selecting interchangeable digital filtering stages.
[0056] A spectrogram can be constructed to analyze audio data in the spectral region. In at least one embodiment, the size of the analysis window is selected along with the size of the overlapping region of the analysis window. In at least one embodiment, the analysis window is slid along the spectrogram, and at each analysis window position, the normalized average size of the analysis window is calculated. In at least one embodiment, the average size is determined as the spectral indicator at each analysis window position. In at least one embodiment, the initial event vector is occupied by calculated pairs of analysis window indicators and their associated positions. In at least one embodiment, the initial event vector indicator is subject to thresholding so as to retain only indicator-position pairs that have indicators exceeding a threshold.
[0057] Each run may include a variable count of indicators of unequal sizes. In at least one embodiment, for each run, the indicators may be internally sorted by indicator value to obtain the largest size indicator.
[0058] For each run, the start / mid-time position and run duration can be extracted.
[0059] The preliminary event vector can be formed using a parameter triplet (M, P, D). In at least one embodiment, the triplet (M, P, D) represents the maximum indicator value, the start / mid-time position, and the run duration, respectively.
[0060] The preliminary event vector can be modified to generate the final crowd noise event vector according to the desired event interval, event loudness, and event duration. In various embodiments, the preliminary event vector is modified by selecting an acceptable event distance, an acceptable event duration, and / or an acceptable event loudness.
[0061] Crowd noise event information can be further processed and automatically added to metadata related to highlights of sports events on television broadcasts.
[0062] System Architecture According to various embodiments, this system can be implemented on any electronic device or set of electronic devices equipped to receive, store, and present information. Such electronic devices may include, for example, desktop computers, laptop computers, televisions, smartphones, tablets, music players, audio devices, kiosks, set-top boxes (STBs), game systems, wearable devices, and home electronic devices.
[0063] While this system is described herein in relation to its implementation in a particular type of computing device, those skilled in the art will recognize that the technology described herein can be implemented in other contexts and in any suitable device capable of actually receiving and / or processing user input and presenting output to the user. Therefore, the following description is intended to illustrate various embodiments as examples, rather than to limit its scope.
[0064] Referring here to Figure 1A, a block diagram is shown illustrating the hardware architecture of system 100 for automatically extracting metadata based on event audio data in a client / server embodiment. Event content, such as audiovisual streams containing audio content, may be provided via a network-connected content provider 124. An example of such a client / server embodiment is a web-based implementation, where each of one or more client devices 106 runs a browser or app that provides a user interface for interacting with content from various servers 102, 114, 116, including a data provider server 122 and / or a content provider server 124, via a communication network 104. The transmission of content and / or data in response to requests from client devices 106 is handled using hypertext markup languages (HTML), Java, Objective-Language, etc. This can be done using any known protocol and language, such as C, Python, or JavaScript.
[0065] The client device 106 may be any electronic device, such as a desktop computer, laptop computer, television, smartphone, tablet, music player, audio device, kiosk, set-top box, game system, wearable device, or home electronic device. In at least one embodiment, the client device 106 has several hardware components well known to those skilled in the art. The input device 151 may be any component that receives input from the user 150, including, for example, a handheld remote control, keyboard, mouse, stylus, touch-sensitive screen (touchscreen), touchpad, gesture receptor, trackball, accelerometer, 5-way switch, microphone, etc. The input may be provided via any preferred mode, including, for example, one or more of pointing, tapping, typing, dragging, gesture, tilting, shaking, and / or speech. The display screen 152 may be any component that graphically displays information, video, content, etc., including depiction and highlighting of events. Such outputs may also include, for example, audiovisual content, data visualizations, navigation elements, graphical elements, queries requesting information and / or parameters for content selection. In at least one embodiment, if only a portion of the desired output is presented at a time, dynamic controls such as a scrolling mechanism may be available via the input device 151 to select the currently displayed information and / or change how the information is displayed.
[0066] The processor 157 may be a conventional microprocessor for performing operations on data under the direction of software, in accordance with well-known art. The memory 156 may be random-access memory having structures and architectures known in the art, used by the processor 157 in the process of running software for performing the operations described herein. The client device 106 may also include local storage (not shown), which may be a hard drive, flash drive, optical or magnetic storage device, or web-based (cloud-based) storage.
[0067] Any preferred type of communication network 104, such as the Internet, television network, cable network, or cellular network, can be used as a mechanism for transmitting data between a client device 106 and various servers 102, 114, 116 and / or content provider 124 and / or data provider 122, according to any preferred protocol and technology. In addition to the Internet, other examples include cellular telephone networks, EDGE, 3G, 4G, Long-Term Evolution (LTE), Session Initiation Protocol (SIP), Short Message Peer-to-Peer Protocol (SMPP), SS7, Wi-Fi, Bluetooth, ZigBee, Hypertext Transfer Protocol (HTTP), Secure Hypertext Transfer Protocol (SHTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), and / or any combination thereof. In at least one embodiment, the client device 106 sends a request for data and / or content over the communication network 104 and receives a response from servers 102, 114, 116 containing the requested data and / or content.
[0068] In at least one embodiment, the system of Figure 1A operates in connection with a sporting event. However, it should be understood that the teachings herein also apply to non-sporting events, and the techniques described herein are not limited to applications to sporting events. For example, the techniques described herein can be used to operate in connection with television shows, movies, news events, game shows, political activities, business shows, dramas, and / or other episodic content, or for two or more such events.
[0069] In at least one embodiment, system 100 identifies highlights of a broadcast event by analyzing audio content representing the event. This analysis can be performed in real time. In at least one embodiment, system 100 includes one or more web servers 102 coupled to one or more client devices 106 via a communication network 104. The communication network 104 may be a public network, a private network, or a combination of public and private networks such as the Internet. The communication network 104 may be a LAN, WAN, wired, wireless, and / or a combination of the above. In at least one embodiment, the client devices 106 may connect to the communication network 104 via either a wired or wireless connection. In at least one embodiment, the client devices may also include a recording device capable of receiving and recording events, such as a DVR, PVR, or other media recording device. Such a recording device may be part of or external to the client device 106. In other embodiments, such a recording device may be omitted. Figure 1A shows one client device 106, but system 100 can be implemented using a single type or any number of client devices 106 of multiple types.
[0070] The web server 102 may include one or more physical computing devices and / or software capable of receiving requests from client devices 106, responding to those requests with data, and sending unilateral alerts and other messages. The web server 102 may employ various strategies for fault tolerance and scalability, such as load balancing, caching, and clustering. In at least one embodiment, the web server 102 may include caching techniques, as known in the art, for storing client requests and information related to events.
[0071] The web server 102 may maintain or designate one or more application servers 114 to respond to requests received from client devices 106. In at least one embodiment, the application server 114 provides access to business logic for use by client application programs within client devices 106. The application server 114 may be located in the same location as, shared with, or co-managed by, the web server 102. The application server 114 may also be located separately from the web server 102. In at least one embodiment, the application server 114 interacts with one or more analytics servers 116 and one or more data servers 118 to perform one or more operations of the disclosed technology.
[0072] One or more storage devices 153 can act as “data stores” by storing data related to the operation of system 100. This data may include, for example, audio data 154 representing one or more audio signals. The audio data 154 may be extracted, for example, from audiovisual streams or stored audiovisual content representing sporting events and / or other events.
[0073] Audio data 154 may include an audio stream accompanying a video image, a processed version of the audiovisual stream, and any information related to the audio embedded in the audiovisual stream, such as metrics and / or vectors related to the audio data 154, such as the time index, duration, magnitude, and / or other parameters of the event. User data 155 may include any information describing one or more users 150, such as demographics, purchasing behavior, audiovisual stream viewing behavior, interests, and preferences. Highlight data 164 may include highlights, highlight identifiers, time indicators, categories, excitement levels, and other data related to the highlights. Audio data 154, user data 155, and highlight data 164 will be described in more detail later.
[0074] In particular, many components of system 100 may be or include computing devices. Each such computing device may have an architecture similar to that of client device 106, as shown and described above. Thus, any of the communication network 104, web server 102, application server 114, analysis server 116, data provider 122, content provider 124, data server 118, and storage device 153 may include one or more computing devices, each of which may optionally have an input device 151, display screen 152, memory 156, and / or processor 157, as described above in relation to client device 106.
[0075] In an exemplary operation of system 100, one or more users 150 of client device 106 view content from content provider 124 in the form of an audiovisual stream. The audiovisual stream may represent an event, such as a sporting event. The audiovisual stream may be a digital audiovisual stream that can be easily processed using known computer vision techniques.
[0076] When an audiovisual stream is displayed, one or more components of System 100, such as the client device 106, the web server 102, the application server 114, and / or the analysis server 116, analyze the audiovisual stream, identify highlights within the audiovisual stream, and / or extract metadata from the audiovisual stream, for example, from the audio component of the stream. This analysis may be performed in response to the receipt of a request to identify highlights and / or metadata of the audiovisual stream. Alternatively, in another embodiment, the highlights and / or metadata may be identified without any specific request made by the user 150. In yet another embodiment, the analysis of the audiovisual stream may be performed without the audiovisual stream being displayed.
[0077] In at least one embodiment, user 150 can specify specific parameters for the analysis of audio data 154 via the input device 151 of client device 106 (e.g., which events / games / teams to include, how much time user 150 has available to view the highlights, which metadata is desired, and / or other parameters). To customize the analysis of audio data 154 without necessarily requiring user 150 to specify preferences, user preferences can also be extracted from storage, such as from user data 155 stored in one or more storage devices 153. In at least one embodiment, user preferences can be determined based on observed behavior and actions of user 150, such as by observing website visit patterns, television viewing patterns, music listening patterns, online purchases, previous highlight identification parameters, highlights and / or metadata actually viewed by user 150, etc.
[0078] Additionally or alternatively, user preferences can be retrieved from previously remembered preferences explicitly provided by user 150. Such user preferences may indicate which teams, sports, players, and / or event types are of interest to user 150, and / or they may indicate what types of metadata or other information related to highlights would be of interest to user 150. Thus, such settings can be used to guide the analysis of audiovisual streams to identify highlights and / or extract highlight metadata.
[0079] The analysis server 116, which may include one or more computing devices as described above, can analyze live and / or recorded feeds of commentary statistics related to one or more events from the data provider 122. Examples of data providers 122 include, but are not limited to, real-time sports information providers such as STATS®, Perform (available from Opta Sports in London, UK), and SportRadar in St. Gallen, Switzerland. In at least one embodiment, the analysis server 116 generates a set of different excitement levels for an event. Such excitement levels can then be stored together with highlights identified or received by the system 100 according to the techniques described herein.
[0080] The application server 114 can analyze the audiovisual stream to identify highlights and / or extract metadata. Additionally or alternatively, such analysis may be performed by the client device 106. The identified highlights and / or extracted metadata may be specific to user 150. In such cases, it may be advantageous to identify the highlights within the client device 106 that are associated with a particular user 150. The client device 106 may receive, retain, and / or retrieve applicable user preferences for highlight identification and / or metadata extraction, as described above. Additionally or alternatively, highlight generation and / or metadata extraction may be performed globally (i.e., using objective criteria applicable to the entire user population, regardless of the preferences of a particular user 150). In such cases, it may be advantageous to identify the highlights and / or extract the metadata within the application server 114.
[0081] Content that facilitates highlight identification, audio analysis, and / or metadata extraction may come from any preferred source, including content providers 124, which may include websites such as YouTube and MLB.com, sports data providers, television stations, client-based or server-based DVRs, etc. Alternatively, the content may come from a local source, such as a DVR or other recording device associated with (or embedded in) the client device 106. In at least one embodiment, the application server 114 generates a customized highlight show with highlights and metadata available to the user 150 as downloadable or streaming content, or on-demand content, or in any other way.
[0082] As described above, it may be advantageous for user-specific highlight identification, audio analysis, and / or metadata extraction to be performed on a specific client device 106 associated with a particular user 150. Such an embodiment can avoid the need to unnecessarily transmit video content or other high-bandwidth content over the communication network 104, particularly when the content is already available on the client device 106.
[0083] For example, referring here to Figure 1B, an example of System 160 is shown in which at least a portion of the audio data 154 and highlight data 164 are stored in a client-based storage device 158, which can be any form of local storage device available to the client device 106. One example is a DVR, which can record events such as video content of a complete sporting event. Alternatively, the client-based storage device 158 can be any magnetic, optical, or electronic storage device for data in digital format. Examples include flash memory, magnetic hard drives, CD-ROMs, DVD-ROMs, or other devices integrated with or communicatively coupled to the client device 106. Based on information provided by the application server 114, the client device 106 can extract metadata from the audio data 154 stored in the client-based storage device 158 and store the metadata as highlight data 164 without needing to search for other content from the content provider 124 or other remote sources. Such a configuration can save bandwidth and make effective use of existing hardware that may already be available to the client device 106.
[0084] Returning to Figure 1A, in at least one embodiment, the application server 114 can identify different highlights and / or extract different metadata for different users 150, depending on the individual user's preferences and / or other parameters. The identified highlights and / or extracted metadata can be presented to the user 150 via any suitable output device, such as the display screen 152 of the client device 106. If desired, multiple highlights can be identified and compiled into a highlight show along with their associated metadata. Such a highlight show can be assembled into a “highlight reel” or set of highlights that is accessed via a menu and / or played for the user 150 according to a predetermined sequence. In at least one embodiment, the user 150 can control the highlight playback and / or delivery of the associated metadata via the input device 151, for example, to: ● Select specific highlights and / or metadata to display. ● Pause, rewind, fast forward, ●Skip to the next highlight ● Return to the beginning of the previous highlight within the highlight show, as well as / or ● Perform other actions.
[0085] Further details regarding these features are provided in the relevant U.S. patent applications cited above.
[0086] In at least one embodiment, one or more data servers 118 are provided. The data server 118 can respond to requests for data from any of the servers 102, 114, 116 to retrieve or provide, for example, audio data 154, user data 155, and / or highlight data 164. In at least one embodiment, such information can be stored in any preferred storage device 153 accessible by the data server 118 and can come from any preferred source, such as the client device 106 itself, a content provider 124, a data provider 122, etc.
[0087] Referring here to Figure 1C, System 180 is shown in an alternative embodiment in which System 180 is implemented in a standalone environment. Similar to the embodiment shown in Figure 1B, at least some of the audio data 154, user data 155, and highlight data 164 may be stored in a client-based storage device 158 such as a DVR. Alternatively, the client-based storage device 158 may be flash memory or a hard drive, or it may be another device integrated with or communicatively coupled to the client device 106.
[0088] User data 155 may include the preferences and interests of user 150. Based on such user data 155, system 180 may extract metadata from audio data 154 and present it to user 150 in the manner described herein. Additionally or alternatively, metadata may be extracted based on objective criteria that are not based on information specific to user 150.
[0089] Referring here to Figure 1D, an overview of System 190 having an architecture according to an alternative embodiment is shown. In Figure 1D, System 190 includes a broadcast service such as a content provider 124, a content receiver in the form of a client device 106 such as a television set with an STB, a video server such as an analysis server 116 that can capture and stream television program content, and / or other client devices 106 such as mobile devices and laptops, all connected via a network such as a communication network 104, which can receive and process television program content. A client-based storage device 158 such as a DVR can be connected to any of the client devices 106 and / or other components and can store audiovisual streams, highlights, highlight identifiers, and / or metadata to facilitate the identification and presentation of highlights and / or extracted metadata via any of the client devices 106.
[0090] The specific hardware architectures shown in Figures 1A, 1B, 1C, and 1D are for illustrative purposes only. Those skilled in the art will recognize that the teachings described herein can be implemented using other architectures. Many of the components shown therein are optional and can be omitted, integrated with other components, and / or replaced by other components.
[0091] In at least one embodiment, the system can be implemented as software written in any suitable computer programming language, whether standalone or in a client / server architecture. Alternatively, it can be implemented and / or embedded in hardware.
[0092] data structure Figure 2 is a schematic block diagram showing an example of a data structure that can be incorporated into audio data 154, user data 155, and highlight data 164 according to one embodiment.
[0093] As shown, audio data 154 may include recordings for each of the multiple audio streams 200. Although audio streams 200 are shown for illustrative purposes, the techniques described herein can be applied to any type of audio data 154 or content, whether streamed or stored. In addition to the audio streams 200, the recordings of audio data 154 may include other data generated in accordance with or useful for the analysis of the audio streams 200. For example, for each audio stream 200, audio data 154 may include a spectrogram 202, one or more analysis windows 204, a vector 206, and a time index 208.
[0094] Each audio stream 200 may exist in the time domain. Each spectrogram 202 can be calculated for the corresponding audio stream 200 in the time-frequency domain. By analyzing the spectrogram 202, audio events at desired frequencies, such as crowd noise, can be more easily identified.
[0095] The analysis window 204 may be a designation of a predetermined time and / or frequency interval for the spectrogram 202. Computationally, the spectrogram 202 can be analyzed using a single moving (i.e., "sliding") analysis window 204, or a series of displaced (optionally overlapping) analysis windows 204 can be used.
[0096] Vector 206 may be a dataset containing intermediate and / or final results from the analysis of audio stream 200 and / or the corresponding spectrogram 202.
[0097] Time index 208 can indicate the time when a significant event occurs within audio stream 200 (and / or the audiovisual stream from which audio stream 200 is extracted). For example, time index 208 could be the time in a broadcast when crowd noise increases or decreases. Thus, in the context of a sporting event, time index 208 could indicate the beginning or end of a particularly interesting part of the audiovisual stream, such as an important or impressive play.
[0098] As further shown, user data 155 may include records related to user 150, each of which may include demographic data 212, preferences 214, browsing history 216, and purchase history 218 of a particular user 150.
[0099] Demographic data 212 may include any type of demographic data, including but not limited to age, sex, place, nationality, religious affiliation, and education level.
[0100] Preferences 214 may include choices made by User 150 regarding their preferences. Preferences 214 may be directly related to the collection and / or viewing of highlights and metadata, or they may be more general in nature. In either case, preferences 214 can be used to facilitate the identification and / or presentation of highlights and metadata to User 150.
[0101] Browsing history 216 may list television programs, audiovisual streams, highlights, web pages, search queries, sporting events, and / or other content searched and / or viewed by user 150.
[0102] Purchase history 218 can list products or services purchased or requested by user 150.
[0103] As further shown, the highlight data 164 may include records for j highlights 220, each record may include an audiovisual stream 222 and / or metadata 224 for a particular highlight 220.
[0104] The audiovisual stream 222 may contain a video depicting the highlight 220, which may be obtained from one or more audiovisual streams of one or more events (for example, by cropping the audiovisual stream to contain only the audiovisual stream 222 related to the highlight 220). Within the metadata 224, the identifier 223 may contain a time index (such as the time index 208 of the audio data 154) and / or other markers indicating where the highlight 220 resides within the audiovisual stream of the event from which the highlight 220 is obtained.
[0105] In some embodiments, each recording of highlight 220 may include only one of the audiovisual stream 222 and identifier 223. Highlight playback may be performed by playing the audiovisual stream 222 for the user 150, or by using identifier 223 to play only the highlighted portion of the audiovisual stream from which the highlight 220 is derived. Storage of identifier 223 is optional. In some embodiments, identifier 223 may be used only to extract the audiovisual stream 222 for highlight 220, and then this highlight 220 may be stored instead of identifier 223. In either case, the time index 208 of highlight 220 may be extracted from audio data 154 and added to highlight 220, or at least temporarily stored as metadata 224 added to the audiovisual stream from which the highlight 220 is derived.
[0106] In addition to, or as an alternative to, identifier 223, metadata 224 may include information about the highlight 220, such as the date of the event, the season, and the groups or individuals involved in the event or audiovisual stream from which the highlight 220 is derived, such as teams, players, coaches, anchors, broadcasters, and fans. Among the information, metadata 224 for each highlight 220 may include phase 226, clock 227, score 228, frame number 229, excitement level 230, and / or crowd excitement level 232.
[0107] Phase 226 can be a phase of an event related to Highlight 220. More specifically, Phase 226 can be a stage of a sporting event in which the start, middle, and / or end of Highlight 220 are located. For example, Phase 226 could be "third quarter," "second inning," "bottom of the inning," etc.
[0108] Clock 227 may be the game clock associated with Highlight 220. More specifically, Clock 227 may be the state of the game clock at the start, middle, and / or end of Highlight 220. For example, Clock 227 may be "15:47" relative to Highlight 220, which marks the start, end, or span of a period of a sporting event where 15 minutes and 47 seconds are displayed on the game clock.
[0109] A score of 228 could be the game score related to Highlight 220. More specifically, a score of 228 could be the score at the beginning, end, and / or middle of Highlight 220. For example, a score of 228 could be "45-38", "7-0", "30-love", etc.
[0110] Frame number 229 may be the number of a video frame in the audiovisual stream from which highlight 220 is obtained, or in the audiovisual stream 222 associated with highlight 220, relating to the start, middle, and / or end of highlight 220.
[0111] The excitement level 230 may be a measure of how exciting or interesting an event or highlight is expected to be for a particular user 150 or for users in general. In at least one embodiment, the excitement level 230 may be calculated as shown in the relevant application above. Additionally or alternatively, the excitement level 230 may be determined by an analysis of audio data 154, which may be components extracted from the audiovisual stream 222 and / or audio stream 200, at least in part. For example, audio data 154 containing higher levels of crowd noise, announcements, and / or uptempo music may indicate a higher excitement level 230 for the relevant highlight 220. The excitement level 230 does not have to be static with respect to the highlight 220, but instead may change during the course of the highlight 220. Thus, the system 100 may further refine the highlight 220 to show the user only the portion that exceeds a threshold excitement level 230.
[0112] The crowd excitement level 232 may be a measure of how excited the crowd participating in the event appears to be. In at least one embodiment, the crowd excitement level 232 may be determined based on an analysis of audio data 154. In other embodiments, visual analysis may be used to measure crowd excitement or to supplement the results of the audio data analysis.
[0113] For example, if analysis of the audio stream 200 of highlight 220 detects intense crowd noise, the crowd excitement level 232 for highlight 220 may be considered relatively high. Similar to the excitement level 230, the crowd excitement level 232 can change throughout the highlight 220. Therefore, the crowd excitement level 232 may include multiple indicators corresponding to specific times within highlight 220.
[0114] The data structure shown in Figure 2 is for illustrative purposes only. Those skilled in the art will recognize that some of the data in Figure 2 may be omitted or replaced with other data in the performance of highlight identification and / or metadata extraction. Additionally or alternatively, data not specifically shown in Figure 2 or described in this application may be used in the performance of highlight identification and / or metadata extraction.
[0115] Audio data 154 In at least one embodiment, the system performs several stages of analysis of audio data 154, such as an audio stream, in the time-frequency domain to detect crowd noise, such as crowd cheers, chants, and fan support, during a depiction of a sporting event or another event. The depiction may be a television broadcast, an audiovisual stream, an audio stream, a stored file, etc.
[0116] First, the compressed audio data 154 is read, decoded, and resampled to the desired sampling rate. Next, the resulting PCM stream is pre-filtered for noise reduction, click removal, and / or selection of the desired frequency band using one of several interchangeable digital filtering steps. Subsequently, a spectrogram is constructed for the audio data 154. A significant collection of peaks in spectral magnitude is identified at each position in a sliding two-dimensional time-frequency area window. A spectral indicator is generated for each position in the analysis window, forming a vector of spectral indicators with the associated time positions.
[0117] Next, runs of selected indicator-and-position pairs with narrow time intervals are identified, and a set of vectors is generated.
number
[0118] Figure 3A shows an example of an audio waveform graph 300 in an audio stream 310 extracted from sports event television program content in the time domain, according to one embodiment. The highlighted area 320 shows an exemplary noise event, such as crowd cheering. The amplitude of the captured audio is relatively high in the highlighted area 320, which may represent a relatively large portion of the audio stream 310.
[0119] Figure 3B shows an example of a spectrogram 350 corresponding to the audio waveform graph 300 in Figure 3A in the time-frequency domain, according to one embodiment. In at least one embodiment, the detection and marking of events of interest are performed in the time-frequency domain, and the timing boundaries of the events are presented in real time to the video highlighting and metadata generation application. This may enable the generation of corresponding metadata 224, identifiers 223 that identify the start and / or end of the highlight 220, the crowd excitement level occurring during the highlight 220, and so on.
[0120] Audio Data Analysis and Metadata Extraction Figure 4 is a flowchart of Method 400 performed by an application (running on one of the client device 106 and / or analysis server 116) that receives an audiovisual stream 222 and performs on-the-fly processing of audio data 154 to extract metadata 224 corresponding to, for example, highlights 220. According to Method 400, audio data 154 such as audio stream 310 can be processed to detect crowd noise audio events, music events, announcement events, and / or other audible events related to the generation of television program content highlights.
[0121] In at least one embodiment, method 400 (and / or other methods described herein) is performed on audio data 154 extracted from an audiovisual stream or other audiovisual content. Alternatively, the techniques described herein can be applied to other types of source content. For example, the audio data 154 does not need to be extracted from an audiovisual stream, but rather may be a radio broadcast or other audio description of a sporting event or other event.
[0122] In at least one embodiment, Method 400 (and / or other methods described herein) may be performed by a system such as System 100 in Figure 1A. However, alternative systems, including but not limited to System 160 in Figure 1B, System 180 in Figure 1C, and System 190 in Figure 1D, may be used instead of System 100 in Figure 1A. Furthermore, the following description assumes that crowd noise events are identified. However, it will be understood that different types of audible events may be identified and used to extract metadata according to methods similar to those described herein.
[0123] Method 400 in Figure 4 can begin with step 410, in which audio data 154, such as an audio stream 200, is read. If the audio data 154 is in a compressed format, it can optionally be decoded. In step 420, the audio data 154 can be resampled to a desired sampling rate. In step 430, the audio data 154 can be filtered using one of several interchangeable digital filtering steps. Then, in step 440, a spectrogram 202 can optionally be generated for the filtered audio data 154 by, for example, calculating a Short-Time Fourier Transform (STFT) for a 1-second chunk of the filtered audio data 154. The time-frequency coefficients of the spectrogram 202 can be stored in a two-dimensional array for further processing.
[0124] In particular, in some embodiments, step 440 may be omitted. Instead of performing an analysis on the spectrogram 202, further analysis can be performed directly on the audio data 154. Figures 5 to 10 below assume that step 440 has been performed and the remaining analysis steps are performed on the spectrogram 202 corresponding to the audio data 154 (for example, after decoding, resampling, and / or filtering the audio data 154 as described above).
[0125] Figure 5 is a flowchart illustrating a method 500 for analyzing audio data 154, such as an audio stream 200, in the time-frequency domain by, for example, analyzing a spectrogram 202 to detect clustering of spectral magnitude peaks associated with long-duration crowd cheers (crowd noise), according to one embodiment. First, in step 510, a two-dimensional rectangular time-frequency analysis window 204 of size (F × T) is selected, where T is a value of several seconds (typically about 6 seconds) and F is the frequency range to be considered (typically 500 Hz to 3 kHz). Next, in step 520, a window overlap region N between adjacent analysis windows 204 is selected and a window sliding step S = (TN) is calculated (typically about 1 second). The method proceeds to step 530, where the analysis window 204 slides along the spectral time axis. In step 540, the normalized magnitude is calculated at each position of the analysis window 204, followed by the calculation of the average peak magnitude of the analysis window 204. The calculated average spectral peak magnitudes represent the event indicators associated with each position in the analysis window 204. In step 550, thresholds are applied to each indicator value to generate the initial event vector 206, which includes indicator-position pairs as its elements.
[0126] As established above, the initial event vector may include a set of indicator-position pairs selected by thresholding in step 550. This vector can then be analyzed to identify densely packed groups of indicators with narrow spacing between adjacent elements. This process is shown in Figure 6.
[0127] Figure 6 is a flowchart of a method 600 for generating crowd noise event vectors according to one embodiment. In step 610, the initial vector of selected events may be read using a set of indicator / position pairs. In step 620, all selected indicator position runs with a position interval of S seconds between adjacent vector elements are set to the vector set
number
[0128] Figure 7 is a flowchart of a method 700 for internal processing of each R vector according to one embodiment. In step 710, the elements of R may be sorted in descending order by indicator value. The maximum indicator value can be extracted as M parameters for the event. In step 720, the start / center time may be recorded as parameter P for each of each vector R. In step 730, for each vector R, the number of elements may be counted and recorded as the duration parameter D for each vector R. For each event, a triplet (M, P, D) may be formed that describes the event's intensity (loudness), start / center position, and / or duration. These triplets may replace the R vectors as new derived elements that fully convey the information desired about the crowd noise event. As shown in the flowchart of Figure 7, subsequent processing may include, in step 740, combining the M, P, and D parameters of each R to form a new vector with the (M, P, D) triplet as its element. The event vector is passed through the process of selecting the event interval, event duration, and event loudness (magnitude indicator) to form the final timeline of the detected crowd noise events.
[0129] Figure 8 is a flowchart of a method 800 for further selecting desired crowd noise events according to one embodiment. According to one embodiment, method 800 can remove event vector elements that are spaced below a minimum time distance between adjacent events. Method 800 may begin in step 810, in which system 100 steps through event vector elements one at a time. In query 820, the time distance to the previous event position can be tested. According to query 820, if this time distance is below a threshold, the position may be skipped in step 830. If the time distance is not below the threshold, the position may be accepted in step 840. In either case, method 800 may proceed to query 850. According to query 850, if the end of the event vector is reached, a modified event vector may be generated, and vector elements that are considered too close to each other are removed. If the end of the event vector is not reached, step 810 can be continued, and additional vector elements may be removed as needed.
[0130] Figure 9 is a flowchart of a method 900 for further selecting desired crowd noise events according to one embodiment. Method 900 can remove event vector elements whose crowd noise duration falls below a desired level. Method 900 may begin in step 910, in which system 100 steps through the duration component of the event vector. In query 920, the duration component of an event vector element can be tested. According to query 920, if this duration falls below a threshold, the event vector element may be skipped in step 940. If the duration does not fall below the threshold, the event vector element may be accepted in step 930. In either case, method 900 may proceed to query 950. According to query 950, if the end of the event vector is reached, a modified event vector may be generated, and the vector element is considered to represent insufficient duration crowd noise and is removed. If the end of the event vector is not reached, step 910 can be continued, and additional vector elements may be removed as needed.
[0131] Figure 10 is a flowchart of a method 1000 for further selecting desired crowd noise events according to one embodiment. Method 1000 can remove event vector elements where the crowd size indicator falls below a desired level. Method 1000 may begin from step 1010, and system 100 steps through the event vectors and subsequent selections. In query 1020, the size of a crowd noise event can be tested. According to query 1020, if this size falls below a threshold, the event vector element may be skipped in step 1040. If the size does not fall below the threshold, its position may be accepted in step 1030. In either case, method 1000 may proceed to query 1050. According to query 1050, if the end of the event vector is reached, a modified event vector may be generated, and the vector element is removed as it is considered to have insufficient crowd noise size. If the end of the event vector is not reached, step 1010 is continued, and additional vector elements may be removed as needed.
[0132] The event vector post-processing steps described in Figures 8, 9, and 10 can be performed in any desired order. The steps shown can be performed in any combination with each other, and some steps can be omitted. At the end of event vector processing, a new final event vector can be generated, containing the desired event timeline of the sporting event.
[0133] In at least one embodiment, an automated video highlight and associated metadata generation application receives a live broadcast audiovisual stream containing audio and video components, or a digital audiovisual stream received via a computer server, and processes audio data 154 extracted from the audiovisual stream using digital signal processing techniques to detect distinct crowd noise (e.g., audience cheers), as described above. These events can be sorted and selected using the techniques described herein. The extracted information can then be added to sports event metadata 224 associated with a sports event television program video and / or video highlight 220. Such metadata 224 can be used, for example, to determine the start / end times of segments used in highlight generation. As described herein and in the related applications above, the start and / or end times of the highlight can be adjusted based on an offset, which can be based on the amount of time available for the highlight, the importance and / or excitement level of the highlight, and / or any other preferred factors. Additionally or alternatively, the metadata 224 can be used to provide information to a user 150 while viewing the audiovisual stream or the highlight 220, such as a corresponding excitement level 230 or crowd excitement level 232.
[0134] This system and method have been described in particular detail with respect to possible embodiments. Those skilled in the art will understand that this system and method may be implemented in other embodiments. First, the specific names, capitalization of terms, attributes, data structures, or other programming or structural aspects of components are neither essential nor important, and mechanisms and / or features may differ in name, format, or protocol. Furthermore, this system may be implemented through a combination of hardware and software, or entirely with hardware elements, or entirely with software elements. Also, the specific division of functions among the various system components described herein are illustrative and not essential. Functions performed by a single system component may instead be performed by multiple components, and functions performed by multiple components may instead be performed by a single component.
[0135] Any reference in this specification to “one embodiment” or “embodiment” means that a particular feature, structure, or characteristic described in relation to the embodiment is included in at least one embodiment. The phrases “in one embodiment” or “in at least one embodiment” appearing in various places in this specification do not necessarily all refer to the same embodiment.
[0136] Various embodiments may include any number of systems and / or methods for performing the above techniques, either individually or in any combination. Another embodiment includes a computer program product comprising a non-temporary computer-readable storage medium and computer program code encoded on the medium for causing a processor in a computing device or other electronic device to perform the above techniques.
[0137] Some of the above sections are presented in terms of algorithms and symbolic representations of operations on data bits in the memory of computing devices. Descriptions and representations of these algorithms are means used by those skilled in the art to most effectively communicate the substance of their work to others skilled in the art. An algorithm is generally considered here as a self-consistent set of steps (instructions) leading to a desired result. A procedure is one that requires the physical manipulation of physical quantities. These quantities, though not always, take the form of electrical, magnetic, or optical signals that can be stored, transferred, combined, compared, and otherwise manipulated. For reasons of common usage, it is sometimes convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. Furthermore, without loss of generality, it is sometimes convenient to refer to a particular arrangement of steps requiring the physical manipulation of physical quantities as a module or code device.
[0138] However, it should be noted that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. As will be evident from the following explanation, unless otherwise specifically stated, explanations throughout the explanation using terms such as “processing,” “calculating,” “calculating,” “displaying,” or “determining” should be understood to refer to the actions and processes of a computer system or similar electronic computing module and / or device that manipulates and transforms data represented as physical (electronic) quantities within the memory or registers or other such information storage, transmission, or display devices of a computer system.
[0139] Certain embodiments include process steps and instructions described herein in the form of algorithms. Process steps and instructions can be embodied in software, firmware, and / or hardware. If embodied in software, it should be noted that they can be downloaded, reside on different platforms used by various operating systems, and operate from them.
[0140] This document also relates to apparatus for performing the operations described herein. Such apparatus may comprise a general-purpose computing device that can be specifically constructed for a required purpose or selectively activated or reconfigured by a computer program stored in the computing device. Such computer programs may be stored in computer-readable storage media, including but not limited to floppy disks, optical disks, CD-ROMs, DVD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROMs, EEPROMs, flash memory, solid-state drives, magnetic or optical cards, application-specific integrated circuits (ASICs), or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus. The program and its associated data may also be hosted and run remotely, for example, on a server. Furthermore, the computing devices referred to herein may include a single processor or may be architectures employing multiple processor designs to enhance computing power.
[0141] The algorithms and representations presented herein are not inherently related to any particular computing device, virtualization system, or other apparatus. Various general-purpose systems may also be used with the programs taught herein, or it may be more convenient to construct dedicated apparatuses to perform the necessary method steps. The structures required for these various systems will be evident from the descriptions provided herein. Furthermore, this system and method are not described with reference to any particular programming language. Various programming languages may be used to implement the teachings described herein, and it will be understood that the above references to specific languages are provided for the purpose of enabling and disclosing best modes of use.
[0142] Accordingly, various embodiments include computer systems, computing devices, or other electronic devices, or software, hardware, and / or other elements for controlling any combination or combination thereof. Such electronic devices may include, for example, processors, input devices (such as keyboards, mice, touchpads, trackpads, joysticks, trackballs, microphones, and / or any combination thereof) according to the art, output devices (such as screens, speakers), memory, long-term storage (such as magnetic storage and optical storage), and / or network connectivity. Such electronic devices may be portable or non-portable. Examples of electronic devices that can be used to implement the described systems and methods include desktop computers, laptop computers, televisions, smartphones, tablets, music players, audio devices, kiosks, set-top boxes, game systems, wearable devices, home electronic devices, and server computers. Electronic devices may use any operating system, not limited to, for example, Linux, Microsoft Windows (available from Microsoft Corporation in Redmond, Washington), Mac OS X (available from Apple Inc. in Cupertino, California), iOS (available from Apple Inc. in Cupertino, California), Android (available from Google, Inc. in Mountain View, California), and / or other operating systems suitable for use on the device.
[0143] While a limited number of embodiments have been described herein, those skilled in the art who appreciate the merits of the above description will understand that other embodiments may be conceived. Furthermore, it should be noted that the language used herein has been chosen primarily for readability and teaching purposes and may not have been chosen to describe or limit the subject matter. Accordingly, this disclosure is intended to illustrate the scope but not to limit it.
Claims
1. A method for extracting metadata from event descriptions, The processor receives audio data for one or more events, The processor identifies one or more portions of the audio data that include crowd excitement data, Identifying the peaks in spectral magnitude at each position within the time-frequency analysis window of the spectrogram of the aforementioned audio data, To generate spectral indicators at each position in the aforementioned time-frequency analysis window, The process includes using the spectral indicator to form a vector of spectral indicators having a relevant time portion, wherein the crowd excitement data is identified to correspond to audio data at a threshold excitement level. The processor extracts the crowd excitement data from the audio data, The processor determines the start and end times of one or more segments of the audio data based on the crowd excitement data, The processor generates one or more highlights based on the start time and end time of one or more segments, A method comprising the processor storing one or more highlights in a data store.
2. The method according to claim 1, further comprising the processor adding the crowd excitement data to the event metadata associated with the audio data.
3. The method according to claim 1, further comprising the processor causing the one or more highlights to be displayed on an output device.
4. The method according to claim 1, further comprising the processor preprocessing the audio data by resampling the audio data to a desired sampling rate.
5. The method according to claim 1, further comprising the processor preprocessing the audio data by filtering the audio data to reduce or remove noise.
6. The method according to claim 1, wherein the processor identifies one or more portions of the audio data that include the crowd excitement data, and includes analyzing visual data corresponding to the audio data.
7. A non-temporary computer-readable medium containing one or more sequences of instructions, wherein the instructions are executed by a processor. The aforementioned processor receives audio data for one or more events, The processor identifies one or more portions of the audio data that include crowd excitement data, Identifying the peaks in spectral magnitude at each position within the time-frequency analysis window of the spectrogram of the aforementioned audio data, To generate spectral indicators at each position in the aforementioned time-frequency analysis window, Identifying, including forming a vector of spectral indicators having a relevant time portion using the spectral indicators, The processor extracts the crowd excitement data from the audio data, wherein the crowd excitement data corresponds to the audio data corresponding to the threshold excitement level. The processor determines the start and end times of one or more segments of the audio data based on the crowd excitement data, The processor generates one or more highlights based on the start time and end time of one or more segments, A non-temporary computer-readable medium that causes a computing system to perform an operation including storing one or more highlights in a data store.
8. The non-temporary computer-readable medium according to claim 7, further comprising the processor adding the crowd excitement data to event metadata associated with the audio data.
9. The non-temporary computer-readable medium according to claim 7, further comprising the processor causing the one or more highlights to be displayed on an output device.
10. The non-temporary computer-readable medium according to claim 7, further comprising the processor preprocessing the audio data by resampling the audio data to a desired sampling rate.
11. The non-temporary computer-readable medium according to claim 7, further comprising the processor preprocessing the audio data by filtering the audio data to reduce or remove noise.
12. The non-temporary computer-readable medium according to claim 7, wherein the processor identifies one or more portions of the audio data that include the crowd excitement data, and includes analyzing visual data corresponding to the audio data.
13. A computer system, Processor and The system comprises memory in which program instructions are stored, and during execution by the processor, Receiving audio data from one or more events, Identifying one or more portions of the aforementioned audio data that include crowd excitement data, Identifying the peaks in spectral magnitude at each position within the time-frequency analysis window of the spectrogram of the aforementioned audio data, To generate spectral indicators at each position in the aforementioned time-frequency analysis window, Identifying, including forming a vector of spectral indicators having a relevant time portion using the spectral indicators, Extracting the crowd excitement data from the aforementioned audio data, Based on the crowd excitement data, the start and end times of one or more segments of the audio data are determined, wherein the crowd excitement data corresponds to audio data corresponding to a threshold excitement level. Based on the start time and end time of one or more segments, one or more highlights are generated. A computer system that causes a computing system to perform an operation including storing one or more of the aforementioned highlights in a data store.
14. The computer system according to claim 13, further comprising adding the crowd excitement data to the event metadata associated with the audio data.
15. The computer system according to claim 13, further comprising the operation of displaying one or more highlights on an output device.
16. The computer system according to claim 13, further comprising preprocessing the audio data by resampling the audio data to a desired sampling rate.
17. The computer system according to claim 13, further comprising preprocessing the audio data by filtering the audio data to reduce or remove noise.
18. The method according to claim 1, further comprising the processor adjusting the start time or end time of one or more segments based on one or more offsets, wherein the one or more offsets are based on the amount of time available for one or more highlights.
19. The operation further comprises the processor adjusting the start time or end time of one or more segments based on one or more offsets, wherein the one or more offsets are based on the amount of time available for one or more highlights, according to claim 7 of the non-temporary computer-readable medium.
20. The computer system according to claim 13, wherein the operation further includes adjusting the start time or end time of one or more segments based on one or more offsets, the one or more offsets being based on the amount of time available for the one or more highlights.