Video processing for embedded information card location and content extraction

The method and system automatically detect and interpret information cards in video frames to generate synchronized metadata for sports highlights, improving interactive television applications with timely user information.

JP2026001126APending Publication Date: 2026-01-06STATS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025163102
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-05-14
Filing Date
2025-09-30
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing television systems lack the ability to automatically detect and interpret embedded information cards within video frames of sports broadcasts, limiting the generation of synchronized metadata and interactive applications.

Method used

A method and system for automatically locating and interpreting information cards within video frames using computer vision techniques, extracting metadata from detected text strings, and generating synchronized metadata for highlights.

Benefits of technology

Enables real-time generation of metadata for highlights, enhancing interactive television applications by providing timely and relevant information to users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026001126000001_ABST
    Figure 2026001126000001_ABST
Patent Text Reader

Abstract

Metadata of one or more highlights of a video stream may be extracted from one or more card images embedded in the video stream.SOLUTION: A highlight may be a segment of a video stream having a particular interest, such as a broadcast of a sporting event. According to one method, video frames of a video stream are stored. The one or more information cards embedded in the decoded video frame may be detected by analyzing one or more predetermined video frame regions. Image segmentation, edge detection, and / or identification of closed contours may then be performed on the identified video frame regions. Further processing may include obtaining the smallest rectangular surrounding area that encloses all remaining segments, which may then be further processed to determine the precise boundaries of the information card. The card image may be analyzed to obtain metadata, which may be stored in association with at least one of the video frames.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 673,412 (Attorney Docket No. THU010-PROV), filed May 18, 2018, for "Machine Learning for Recognizing and Interpreting Embedded Information Card Content," which is hereby incorporated by reference in its entirety.

[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 673,411 (Attorney Docket No. THU009-PROV), filed May 18, 2018, for "Video Processing for Enabling Sports Highlights Generation," which is incorporated herein by reference in its entirety.

[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 673,413 (Attorney Docket No. THU012-PROV), filed May 18, 2018, for "Video Processing for Embedded Information Card Localization and Content Extraction," which is hereby incorporated by reference in its entirety.

[0004] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 680,955 (Attorney Docket No. THU007-PROV), filed June 5, 2018, for "Audio Processing for Detecting Occurrences of Crowd Noise in Sporting Event Television Programming," which is incorporated herein by reference in its entirety.

[0005] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 712,041 (Attorney Docket No. THU006-PROV), filed July 30, 2018, for "Audio Processing for Extraction of Variable Length Disjoint Segments from Television Signal," which is hereby incorporated by reference in its entirety.

[0006] This application is filed on October 16, 2018, for Detecting Occurrences of Loud Sound This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 746,454 (Attorney Docket No. THU016-PROV) for "Suitable for Applications Characterized by Short-Time Energy Bursts," which is incorporated herein by reference in its entirety.

[0007] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,710 (Attorney Docket No. THU010), filed May 14, 2019, for "Machine Learning for Recognizing and Interpreting Embedded Information Card Content," which is hereby incorporated by reference in its entirety.

[0008] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,704 (Attorney Docket No. THU009), filed May 14, 2019, for "Video Processing for Enabling Sports Highlights Generation," which is incorporated herein by reference in its entirety.

[0009] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,713 (Attorney Docket No. THU012), filed May 14, 2019, for "Video Processing for Embedded Information Card Localization and Content Extraction," which is hereby incorporated by reference in its entirety.

[0010] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,915, filed August 31, 2012, and issued June 16, 2015, as U.S. Patent No. 9,060,210, for "Generating Excitement Levels for Live Performances," which is hereby incorporated by reference in its entirety.

[0011] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,927, filed August 31, 2012, and issued September 23, 2014, as U.S. Patent No. 8,842,007, for "Generating Alerts for Live Performances," which is hereby incorporated by reference in its entirety.

[0012] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,933, filed August 31, 2012, and issued November 26, 2013, as U.S. Patent No. 8,595,763, for "Generating Teasers for Live Performances," which is hereby incorporated by reference in its entirety.

[0013] This application is related to U.S. Utility Patent Application Serial No. 14 / 510,481 (Attorney Docket No. THU001), filed October 9, 2014, for "Generating a Customized Highlight Sequence Depicting an Event," which is hereby incorporated by reference in its entirety.

[0014] This application is related to U.S. Utility Patent Application Serial No. 14 / 710,438 (Attorney Docket No. THU002), filed May 12, 2015, for "Generating a Customized Highlight Sequence Depicting Multiple Events," which is hereby incorporated by reference in its entirety.

[0015] This application is related to U.S. Utility Patent Application Serial No. 14 / 877,691 (Attorney Docket No. THU004), filed October 7, 2015, for "Customized Generation of Highlight Show with Narrative Component," which is hereby incorporated by reference in its entirety.

[0016] This application is related to U.S. Utility Patent Application Serial No. 15 / 264,928 (Attorney Docket No. THU005), filed September 14, 2016, for "User Interface for Interaction with Customized Highlight Shows," which is hereby incorporated by reference in its entirety.

[0017] This document relates to techniques for identifying multimedia content and associated information on television devices or video servers that deliver the multimedia content, and enabling embedded software applications to utilize the multimedia content to provide content and services in synchronization with the multimedia content. Various embodiments relate to methods and systems for providing automated video and audio analysis used to identify and extract information within sports television video content and create metadata associated with video highlights for in-game and post-game review of the sports television video content. [Background technology]

[0018] Enhanced television applications such as interactive advertising and enhanced program guides with pre-game, in-game, and post-game interactive applications have long been envisioned. Existing cable systems, originally designed for broadcast television, are being called upon to support a host of new applications and services, including interactive television services and enhanced (interactive) programming guides.

[0019] Several frameworks have been standardized to enable enhanced television applications, for example OpenCable (商標) These include the Enhanced TV Application Messaging specification and the Tru2way specification, which refer to interactive digital cable services delivered over cable video networks and include features such as interactive program guides, interactive advertising, and games. Additionally, cable operators' "OCAP" programs offer interactive services such as e-commerce shopping, online banking, electronic program guides, and digital video recording. These efforts enable the first generation of video-synchronized applications that synchronize with video content delivered by program producers / broadcasters, providing additional data and interactivity for television programming.

[0020] Recent developments in video and audio content analysis technologies and compatible mobile devices have opened up a range of new possibilities for developing advanced applications that operate in sync with live TV programming events. These new technologies, along with advances in computer vision and video processing, and the improved computing power of modern processors, make it possible to generate high-quality program content highlights accompanied by metadata in real time. Summary of the Invention

[0021] A method and system are presented for automatically locating information cards ("card images"), such as information scoreboards, within a video frame or multiple video frames in a sports television broadcast program. Also described are methods and systems for identifying text strings within various fields of the located card images and for reading and interpreting text information from various fields of the located card images.

[0022] In at least one embodiment, the detection, location, and reading of the card images is performed synchronously with respect to the presentation of the sports television program content. In at least one embodiment, an automated process is provided for receiving a digital video stream, analyzing one or more frames of the digital video stream, and automatically detecting and locating card image quadrilaterals. In another embodiment, an automated process is provided for analyzing one or more located card images, recognizing and extracting text strings (e.g., within text boxes), and reading information from the extracted text boxes.

[0023] In yet another embodiment, detected text strings associated with specific fields within the card image are interpreted, thus providing instant in-game information related to the content of a televised sporting event. The extracted in-frame information may be used to generate metadata associated with automatically created custom video content, such as a set of highlights of the broadcast television program content associated with the audiovisual and textual data.

[0024] In at least one embodiment, a method for extracting metadata from a video stream may include storing at least one portion of the video stream in a data store. In the processor, one or more card images embedded in at least one of the video frames may be automatically identified and extracted by at least one of identifying predetermined locations within the video frames that define video frame regions that include the card images and sequentially processing multiple regions of the video frames to identify video frame regions that include the card images. In the processor, the card images may be analyzed to obtain metadata, which may be stored in the data store in association with at least one of the video frames.

[0025] In at least one embodiment, the video stream may be a broadcast of a sporting event, the video frames may constitute highlights deemed to be of particular interest to one or more users, and the metadata may describe the status of the sporting event during the highlights.

[0026] In at least one embodiment, the method may further include presenting, at an output device, the metadata during display of the highlight. Automatically identifying and extracting card images and analyzing the card images to obtain the metadata may be performed on the highlight during display of the highlight.

[0027] In at least one embodiment, the method may further include locating and extracting the card image from the video frame region. Locating and extracting the card image from the video frame region may include cropping the video frame to isolate the video frame region. Alternatively or additionally, locating and extracting the card image from the video frame region may include segmenting the video frame region or a processed version of the video frame region to generate a segmented image and modifying pixel values ​​of the segment adjacent to a boundary of the segmented image. Alternatively or additionally, locating and extracting the card image from the video frame region may include removing a background from the video frame region or the processed version of the video frame region. Alternatively or additionally, locating and extracting the card image from the video frame region may include generating an edge image based on the video frame region, finding a contour in the edge image, approximating the contour as a polygon, and extracting an area enclosed by a smallest rectangular perimeter that encompasses all of the contour to generate a bounding rectangle image.

[0028] In at least one embodiment, the method may further include iteratively counting color-corrected pixels for each edge of the surrounding rectangular image and moving any boundary edges inward with a number of color-corrected pixels exceeding a threshold.

[0029] In at least one embodiment, the method may further include verifying the detected quadrangle within the region by counting a first number of pixels within the video frame region, a second number of pixels within the surrounding rectangle image, and a third number of pixels within the adjusted surrounding rectangle image. The first, second, and third numbers may be compared to determine whether the proposed quadrangle within the region is feasible.

[0030] In at least one embodiment, locating and extracting the card image from the video frame area may include adjusting a left border (or another border) of the card image.

[0031] Further details and variations are described herein. [Brief explanation of the drawings]

[0032] The accompanying drawings, together with the description, illustrate several embodiments. Those skilled in the art will recognize that the specific embodiments shown in the drawings are merely exemplary and are not intended to be limiting in scope. [Figure 1A] FIG. 1 is a block diagram depicting a hardware architecture according to a client / server embodiment, where event content is provided via networked content providers. [Figure 1B] 1 is a block diagram depicting a hardware architecture according to another client / server embodiment, where event content is stored on a client-based storage device. [Figure 1C] FIG. 2 is a block diagram illustrating a hardware architecture according to a stand-alone embodiment. [Figure 1D] FIG. 1 is a block diagram illustrating an overview of a system architecture, according to one embodiment. [Figure 2] FIG. 1 is a schematic block diagram illustrating an example of a data structure that may be incorporated into card images, user data, and highlight data, according to one embodiment. [Figure 3A] 1 is a screenshot view of an example video frame from a video stream showing an information card image ("card image") embedded within the frame, such as might be found in television program content of a sporting event. [Figure 3B] 10A-10C are a series of screenshot diagrams depicting additional examples of video frames with embedded card images. [Figure 4]4 is a flowchart depicting a method performed by an application to receive a video stream and perform on-the-fly processing of video frames to locate and extract card images and associated metadata, such as the card images of FIG. 3 and their associated match status information, according to one embodiment. [Figure 5] 5 is a flowchart illustrating in more detail the steps for processing a predetermined region of a video frame to detect a viable card image from FIG. 4; [Figure 6] 1 is a flowchart depicting a method for top-level processing for valid card image quadrangle detection in a specified area of ​​a decoded video frame, according to one embodiment. [Figure 7] 1 is a flowchart illustrating a method for more accurate card image quadrilateral determination, according to one embodiment. [Figure 8] 10 is a flowchart depicting a method for adjusting the quadrilateral boundary of an enclosure that encompasses all detected contours, according to one embodiment. [Figure 9] 1 is a flowchart depicting a method for card image quadrilateral verification, according to one embodiment. [Figure 10] 10 is a flowchart depicting a method for optional stabilization of the left border of a very elongated card image shape, according to one embodiment. [Figure 11] 1 is a flowchart illustrating a method for performing text extraction from a card image 207, according to one embodiment. [Figure 12] 1 is a flowchart depicting a method for performing processing and interpretation of a text string, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0033] definition The following definitions are provided for illustrative purposes only and are not intended to limit the scope. Event: For purposes of this description, the term "event" refers to a game, session, matchup, series, performance, program, and / or concert, etc., or portions thereof (such as an act, period, quarter, half, inning, scene, or chapter). An event may be a sporting event, an entertainment event, or a specific performance of a single individual or a subset of individuals within a larger group of event participants. Examples of non-sporting events include television shows, breaking news, sociopolitical events, natural disasters, movies, plays, radio programs, podcasts, audiobooks, online content, and / or musical performances. An event can be of any length. For illustrative purposes, the technology is often described herein in terms of sporting events, but those skilled in the art will recognize that the technology can be used in other contexts, including highlight shows of any audiovisual, audio, qualification, graphics-based, interactive, non-interactive, or text-based content. Thus, the use of the term “sporting event” and any other sports-specific terminology in this description is intended to illustrate one contemplated embodiment, but is not intended to limit the scope of the described technology to that one embodiment. Rather, such terms should be considered to extend to any suitable non-sporting context appropriate to this technology. For ease of explanation, the term “event” is also used to refer to a report or representation of an event, such as an audiovisual recording of the event, or any other content item that includes a report, description, or depiction of the event. Highlights: An excerpt or portion of an event, or content associated with an event, that is deemed to be of particular interest to one or more users. Highlights can be of any length. Generally, the technology described herein provides a mechanism for identifying and presenting a customized set of highlights (which may be selected based on specific characteristics and / or user preferences) for any suitable event.The term "highlight" is also used to refer to a report or representation of a highlight, such as an audiovisual recording of a highlight, or any other content item that includes a report, description, or depiction of a highlight. Highlights need not be limited to depictions of the event itself, but can include other content associated with the event. For example, in the case of a sporting event, highlights can include audio / video from the game as well as other content, including pre-game, in-game, and post-game interviews, analysis, and / or commentary. Such content can be recorded from linear television (e.g., as part of a video stream depicting the event itself) or can be derived from any number of other sources. Various types of highlights can be provided, including, for example, occurrences (plays), strings, possessions, and sequences, all of which are defined below. Highlights need not be of a fixed duration, but can incorporate a start offset and / or end offset, as described below. Content delineator: One or more video frames that indicate the beginning or end of a highlight. Occurrence: Something that occurs during an event. Examples include a goal, a play, a down, a hit, a save, a shot on goal, a basket, a steal, a snap or snap attempt, a near miss, a fight, the start or end of a game, a quarter, half, period, or inning, a pitch, a penalty, an injury, a dramatic event at an entertainment event, a song, and / or a solo. An occurrence can also be an unusual incident, such as a power outage and / or an incident with an unruly fan. Detection of such an occurrence can be used as the basis for determining whether to designate a particular portion of a video stream as a highlight. An occurrence is also referred to herein as a "play" for ease of naming, although such usage should not be construed as limiting in scope. An occurrence may have any length, and representations of occurrences may have varying lengths.For example, as described above, an extended representation of an occurrence may include footage depicting the time period immediately before and after the occurrence, while a simple representation may include only the occurrence itself. Optional intermediate representations may also be provided. In at least one embodiment, the selection of a duration for representing an occurrence may vary depending on user preferences, available time, a determined excitement level for the occurrence, the importance of the occurrence, and / or any other factors. Offset: An amount by which the length of a highlight is adjusted. In at least one embodiment, a start offset and / or end offset may be provided to adjust the start time and / or end time of the highlight, respectively. For example, if a highlight depicts a goal, the highlight may be extended by several seconds (via an end offset) to include the celebration and / or fan reaction following the goal. The offset may be configured to vary automatically or manually based, for example, on the time available for the highlight, the importance and / or excitement level of the highlight, and / or any other suitable factors. String: A series of occurrences that are linked or related to each other in some way. An occurrence may occur within a possession (defined below) or across multiple possessions. An occurrence may occur within a sequence (defined below) or across multiple sequences. Occurrences may be linked or related because they have some thematic or narrative connection to one another, or because one leads to another, or for any other reason. An example of a string is a set of passes leading to a goal or basket. This should not be confused with "string of text," which has the meaning typically assigned to it in computer programming. Possession: Any time-bound portion of an event. The distinction between the start and end times of a possession may vary depending on the type of event. For certain sporting events (e.g., basketball or soccer) where one team can be offensive while the other team is defensive, a possession can be defined as the period of time one team has the ball.In sports where possession of the puck or ball is more fluid, such as hockey or soccer, possession is considered to extend to the period of time when one of the teams has substantial control of the puck or ball, ignoring momentary contact by the other team (such as a blocked shot or save). In baseball, a possession is defined as a half-inning. In soccer, a possession can include several sequences in which the same team has the ball. For other types of sporting and non-sporting events, the term "possession" may be somewhat misleading, but is still used herein for illustrative purposes. Examples in non-sporting contexts include a chapter, scene, act, or television segment. For example, in the context of a music concert, a possession might correspond to the performance of a single song. A possession can include any number of occurrences. Sequence: A time-delimited portion of an event that encompasses the time period of one continuous action. For example, in a sporting event, a sequence may begin at the start of an action (such as a face-off or tip-off) and end when the whistle is blown to indicate a stoppage of action. In sports such as baseball or soccer, a sequence may equate to a play, which is a form of occurrence. A sequence can include any number of possessions or may be a portion of a possession. Highlight show: A set of highlights arranged for presentation to a user. A highlight show may be presented linearly (e.g., a video stream) or in a way that allows the user to choose which highlights to watch and in what order (e.g., by clicking links or thumbnails). The presentation of a highlight show may be non-interactive or interactive, allowing, for example, the user to pause, rewind, skip, fast-forward, and / or communicate preferences. A highlight show may be, for example, a condensed game.A highlight show can include any number of consecutive or non-consecutive highlights from a single event or from multiple events, and can even include highlights from different types of events (e.g., a combination of highlights from different sports and / or sporting and non-sporting events). User / Viewer: The terms "user" or "viewer" refer interchangeably to an individual, group, or other entity that watches, listens to, or otherwise experiences an event, one or more highlights from an event, or a highlight show. User or viewer can also refer to an individual, group, or other entity that watches, listens to, or otherwise experiences an event, one or more highlights from an event, or a highlight show at some future point in time. While the term "viewer" is sometimes used for descriptive purposes, an event need not include a visual component, and therefore a "viewer" may instead be a listener or any other consumer of content. Narrative: A coherent story that links a set of highlight segments in a particular order. Excitement Level: A measure of how exciting or interesting an event or highlight will be to a particular user or users in general. Excitement level can also be determined with respect to a particular occurrence or player. Various techniques for measuring or assessing excitement levels are described in the related applications referenced above. As described, excitement levels may vary depending on the occurrence within an event and other factors, such as the overall context or importance of the event (e.g., playoff matches, pennant influence, and / or rivalries). In at least one embodiment, an excitement level may be associated with each occurrence, string, possession, or sequence within an event. For example, the excitement level of a possession may be determined based on the occurrences occurring within that possession. Excitement levels may be measured differently by different users (e.g., fans of a team and neutral fans) and may vary depending on each user's personal characteristics.Metadata: Data that relates to and is stored in association with other data. Primary data may be media such as sports programming or highlights. Card Image: An image within a video frame that provides data about something depicted in the video, such as an event, a depiction of an event, or a portion thereof. Exemplary card images include match scores, match clocks, and / or other statistics from a sporting event. Card images may appear momentarily or for the entire duration of the video stream, and those that appear momentarily may be specifically related to the portion of the video stream in which they appear. A "card image" is a modified or processed version of the actual card image that appears within a video frame. · Character image: A portion of an image believed to be associated with a single character. A character image may include an area surrounding a character. For example, a character image may include a roughly rectangular bounding box surrounding a character. · Character: A word, a number, or a symbol that can be part of the representation of a word or number. A character can include letters, numbers, and special characters and may be in any language. · String: A set of characters grouped in a way that indicates they relate to a single piece of information, such as the names of teams playing in a sporting event. Often, English strings of characters are arranged horizontally and read from left to right. However, strings may be arranged differently in English than in other languages. · Video frame region: A portion of a video frame believed to contain a card image based either on knowledge of predetermined locations where card images are expected to appear within the video frame or on sequential analysis of multiple regions of the video frame to identify which regions are likely to contain card images.

[0034] overview According to various embodiments, methods and systems are provided for automatically creating time-based metadata associated with highlights of a television program of a sporting event. The highlights and associated intra-frame time-based information may be extracted synchronously with respect to the television broadcast of the sporting event, or may be extracted while the video content of the sporting event is being streamed from a backup device via a video server after the television broadcast of the sporting event.

[0035] In at least one embodiment, a software application operates synchronously with the playback and / or reception of television program content to provide informational metadata associated with highlights of the content. Such software may run, for example, on the television device itself, on an associated STB, on a video server capable of receiving and subsequently streaming program content, or on a mobile device capable of receiving a video feed containing live programming.

[0036] In video management and processing systems, and in the context of interactive (enhanced) program guides, a set of video clips representing television broadcast content highlights can be automatically generated and / or stored in real time, along with a database containing time-based metadata that more fully describes the events presented in the highlights. The metadata accompanying the video clips can include any information, such as, for example, text information, images, and / or any type of audiovisual data. In this manner, interactive television applications can provide timely and relevant content to users viewing program content on either a primary television display or a secondary display, such as a tablet, laptop, or smartphone.

[0037] One type of metadata associated with highlights of in-game and post-game video content conveys real-time information about sports game parameters extracted directly from the live program content by reading information cards (“card images”) embedded in one or more of the video frames of the program content. In various embodiments, the systems and methods described herein enable this type of automatic metadata generation.

[0038] In at least one embodiment, a system and method automatically detects and locates card images embedded in one or more decoded video frames of a television broadcast of a sporting event program or in a sporting event video streamed from a playback device. A number of predetermined regions of interest within the decoded video frames are analyzed, and card image quadrangles are located and processed in real time using computer vision techniques to convert information from the identified card images into a set of metadata describing the status of the sporting event.

[0039] In another embodiment, an automated process is described in which a digital video stream is received and one or more video frames of the digital video stream are analyzed for the presence of card image rectangles. A text box is then located within the identified card image and the text present within the text box is interpreted to create a metadata file that associates the card image content with video highlights in the analyzed digital video stream.

[0040] In yet another embodiment, multiple text strings (text boxes) are identified, and the position and size of the image of each character within the string of characters associated with the text box are detected. The multiple text strings from various fields of the card image are then processed and interpreted to form corresponding metadata that provides multiple pieces of information related to the portion of the sporting event associated with the processed card image and the analyzed video frame.

[0041] The automated metadata generation video system presented herein can operate in connection with live broadcast video streams or digital video streamed via a computer server. In at least one embodiment, the video stream can be processed in real time using computer vision techniques to extract metadata from embedded card images.

[0042] System Architecture According to various embodiments, the system can be implemented in any electronic device or set of electronic devices equipped to receive, store, and present information, such as, for example, a desktop computer, a laptop computer, a television, a smartphone, a tablet, a music player, an audio device, a kiosk, a set-top box (STB), a gaming system, a wearable device, and / or a consumer electronic device.

[0043] Although the system is described herein with reference to implementation on a particular type of computing device, those skilled in the art will recognize that the techniques described herein may be implemented in other contexts, and indeed on any suitable device capable of receiving and / or processing user input and presenting output to a user. Accordingly, the following description is intended to illustrate various embodiments by way of example, rather than to limit the scope.

[0044] 1A, a block diagram is shown depicting the hardware architecture of a system 100 for automatically extracting metadata from card images embedded in a video stream of an event, according to a client / server embodiment. Event content, such as a video stream, may be provided via a network-connected content provider 124. An example of such a client / server embodiment is a web-based implementation, in which one or more client devices 106 each run a browser or app that provides a user interface for interacting with content from various servers 102, 114, 116, including data provider(s) server(s) 122 and / or content provider(s) server(s) 124, via a communications network 104. Transmission of content and / or data in response to requests from the client devices 106 may be performed using any known protocols and languages, such as Hypertext Markup Language (HTML), Java, Objective C, Python, and / or JavaScript.

[0045] The client device 106 may be a desktop computer, a laptop computer, a television, a smartphone, a tablet, a music player, an audio device, a kiosk, a set-top box, a gaming system, a wearable device, a consumer electronic device, and / or any electronic device, etc. In at least one embodiment, the client device 106 has several hardware components known to those skilled in the art. The input device(s) 151 may be any component(s) that receive input from the user 150, including, for example, a handheld remote control, a keyboard, a mouse, a stylus, a touch-sensitive screen (touch screen), a touchpad, a gesture receptor, a trackball, an accelerometer, a five-way switch, or a microphone. The input may be provided via any suitable mode, including, for example, one or more of pointing, tapping, typing, dragging, gesturing, tilting, shaking, and / or speech. The display screen 152 may be any component that graphically displays information, video, and / or content, including depictions of events and / or highlights, etc. Such output may also include, for example, audiovisual content, data visualization, navigational elements, graphical elements, or queries requesting information and / or parameters for content selection, etc. In at least one embodiment in which only some of the desired outputs are presented at a time, dynamic control, such as a scrolling mechanism, may be available via input device(s) 151 to select which information is currently displayed and / or to change how the information is displayed.

[0046] Processor 157 may be a conventional microprocessor for performing operations on data under the direction of software in accordance with well-known techniques. Memory 156 may be random access memory having a structure and architecture known in the art for use by processor 157 in the course of executing software to perform the operations described herein. Client device 106 may also include local storage (not shown), which may be a hard drive, flash drive, optical or magnetic storage device, and / or web-based (cloud-based) storage, etc.

[0047] Any suitable type of communications network 104, such as the Internet, a television network, a cable network, and / or a cellular network, may be used as a mechanism for transmitting data between the client device 106 and the various server(s) 102, 114, 116 and / or content provider(s) 124 and / or data provider(s) 122 according to any suitable protocols and techniques. In addition to the Internet, other examples include cellular networks, EDGE, 3G, 4G, Long Term Evolution (LTE), Session Initiation Protocol (SIP), Short Message Peer-to-Peer Protocol (SMPP), SS7, Wi-Fi, Bluetooth, ZigBee, Hypertext Transfer Protocol (HTTP), Secure Hypertext Transfer Protocol (SHTTP), and / or Transmission Control Protocol / Internet Protocol (TCP / IP), etc., and / or any combination thereof. In at least one embodiment, the client device 106 sends requests for data and / or content over the communications network 104 and receives responses from the servers 102, 114, 116 that include the requested data and / or content.

[0048] 1A operates in connection with a sporting event. However, it should be understood that the teachings herein apply to events other than sporting events, and the techniques described herein are not limited to application to sporting events. For example, the techniques described herein can be utilized to operate in connection with television shows, movies, news events, game shows, political campaigns, business shows, dramas, and / or other episodic content, or for two or more such events.

[0049] In at least one embodiment, system 100 identifies highlights of a broadcast event by analyzing a video stream of the event. This analysis can be performed in real time. In at least one embodiment, system 100 includes one or more web server(s) 102 coupled to one or more client devices 106 via a communications network 104. Communications network 104 may be a public network, a private network, or a combination of public and private networks, such as the Internet. Communications network 104 may be a LAN, a WAN, wired, wireless, and / or a combination of the above. Client device 106, in at least one embodiment, can connect to communications network 104 via either a wired or wireless connection. In at least one embodiment, client device 106 may also include a recording device, such as a DVR, PVR, or other media recording device, capable of receiving and recording the event. Such a recording device may be part of client device 106 or may be external. In other embodiments, such a recording device may be omitted. Although FIG. 1A shows one client device 106, the system 100 may be implemented with any number of client device(s) 106 of a single type or multiple types.

[0050] Web server(s) 102 may include one or more physical computing devices and / or software capable of receiving requests from client device(s) 106, responding to those requests with data, as well as sending unsolicited alerts and other messages. Web server(s) 102 may employ various strategies for fault tolerance and scalability, such as load balancing, caching, and clustering. In at least one embodiment, web server(s) 102 may include caching techniques, as known in the art, for storing information related to client requests and events.

[0051] The web server(s) 102 may maintain or otherwise designate one or more application server(s) 114 to respond to requests received from the client device(s) 106. In at least one embodiment, the application server(s) 114 provide access to business logic for use by client application programs in the client device(s) 106. The application server(s) 114 may be co-located, shared, or co-managed with the web server(s) 102. The application server(s) 114 may also be remote from the web server(s) 102. In at least one embodiment, the application server(s) 114 interact with one or more analytics server(s) 116 and one or more data server(s) 118 to perform one or more operations of the disclosed techniques.

[0052] One or more storage devices 153 may function as a "data store" by storing data related to the operation of system 100. This data may include, for example, but is not limited to, card data 154 related to card images embedded in a video stream presenting an event, such as a sporting event, user data 155 related to one or more users 150, and / or highlight data 164 related to one or more highlights of an event.

[0053] Card data 154 may include any information related to card images embedded in the video stream, such as the card images themselves, subsets thereof, such as character images, text extracted from the card images, such as characters and strings of characters, and any of the aforementioned attributes useful for extracting text and / or meaning. User data 155 may include any information describing one or more users 150, including, for example, demographics, purchasing behavior, video stream viewing behavior, interests, and / or preferences. Highlight data 164 may include highlights, highlight identifiers, time indexes, categories, excitement levels, and other data related to the highlights. Card data 154, user data 155, and highlight data 164 are described in more detail below.

[0054] Notably, many components of system 100 may be or include computing devices. Each of such computing devices may have an architecture similar to that of client device 106, as shown and described above. Accordingly, any of communication network 104, web server 102, application server 114, analytics server 116, data provider 122, content provider 124, data server 118, and storage device 153 may include one or more computing devices, which may optionally have input device 151, display screen 152, memory 156, and / or processor 157, as described above in connection with client device 106.

[0055] In an exemplary operation of the system 100, one or more users 150 of the client devices 106 view content from the content providers 124 in the form of a video stream. The video stream may depict an event, such as a sporting event. The video stream may be a digital video stream that can be easily processed with known computer vision techniques.

[0056] Once the video stream is displayed, one or more components of the system 100, such as the client device 106, the web server 102, the application server 114, and / or the analytics server 116, may analyze the video stream to identify highlights within the video stream and / or extract metadata from the video stream, for example, from embedded card images and / or other aspects of the video stream. This analysis may be performed in response to receiving a request to identify highlights and / or metadata from the video stream. Alternatively, in another embodiment, highlights may be identified without a specific request by the user 150. In yet another embodiment, analysis of the video stream may occur without the video stream being displayed.

[0057] In at least one embodiment, user 150 can specify certain parameters for the analysis of the video stream (e.g., which events / matches / teams to include, how much time user 150 has available to watch highlights, what metadata is desired, and / or any other parameters, etc.) via input device(s) 151 of client device 106. User preferences can also be retrieved from storage, such as from user data 155 stored on one or more storage devices 153, to customize the analysis of the video stream without necessarily requiring user 150 to specify preferences. In at least one embodiment, user preferences can be determined based on observed behavior and actions of user 150, for example, by observing website visiting patterns, television viewing patterns, music listening patterns, online purchases, prior highlight identification parameters, and / or highlights and / or metadata actually viewed by user 150, etc.

[0058] Additionally or alternatively, user preferences may be retrieved from pre-stored preferences explicitly provided by user 150. Such user preferences may indicate which teams, sports, players, and / or types of events are of interest to user 150, and / or they may indicate what type of metadata or other information associated with highlights will be of interest to user 150. Such preferences may therefore be used to guide analysis of the video stream to identify highlights and / or extract metadata for highlights.

[0059] Analysis server(s) 116, which may include one or more computing devices described above, can analyze live and / or recorded feeds of play-by-play statistics related to one or more events from data provider(s) 122. Examples of data provider(s) 122 include, but are not limited to, providers of real-time sports information such as STATS™, Perform (available from Opta Sports, London, UK), and SportRadar, St. Gallen, Switzerland. In at least one embodiment, analysis server(s) 116 generate a set of different excitement levels for the event. Such excitement levels can then be stored in association with highlights identified by system 100 according to the techniques described herein.

[0060] The application server(s) 114 may analyze the video stream to identify highlights and / or extract metadata. Additionally or alternatively, such analysis may be performed by the client device(s) 106. The identified highlights and / or extracted metadata may be specific to a user 150; in such cases, it may be advantageous to identify highlights within the client device 106 that are associated with the particular user 150. The client device 106 may receive, maintain, and / or acquire applicable user preferences for highlight identification and / or metadata extraction, as described above. Additionally or alternatively, highlight generation and / or metadata extraction may be performed globally (i.e., using objective criteria generally applicable to a user population, regardless of the preferences of a particular user 150). In such cases, it may be advantageous to identify highlights and / or extract metadata within the application server(s) 114.

[0061] The content facilitating highlight identification and / or metadata extraction may come from any suitable source, including content provider(s) 124, including websites such as YouTube® and MLB.com, sports data providers, television stations, and / or client- or server-based DVRs. Alternatively, the content may come from a local source, such as a DVR or other recording device associated with (or embedded in) the client device 106. In at least one embodiment, the application server(s) 114 generate a customized highlight show with highlights and metadata available to the user 150, either as download, or streaming content, or on-demand content, or in some other manner.

[0062] As noted above, it may be advantageous for user-specific highlight identification and / or metadata extraction to be performed at a particular client device 106 associated with a particular user 150. Such an embodiment may avoid the need for video content or other high-bandwidth content to be unnecessarily transmitted over the communications network 104, especially if such content is already available at the client device 106.

[0063] For example, referring now to FIG. 1B , an example of a system 160 according to one embodiment is shown in which card data 154 and at least some of highlight data 164 are stored on a client-based storage device 158, which may be any type of local storage device available to the client device 106. Examples include a DVR capable of recording events, such as video content of a complete sporting event. Alternatively, the client-based storage device 158 may be any magnetic, optical, or electronic storage device for data in digital form. Examples include flash memory, a magnetic hard drive, a CD-ROM, a DVD-ROM, or other devices integrated with or communicatively coupled to the client device 106. Based on information provided by the application server(s) 114, the client device 106 may extract metadata from the card data 154 stored on the client-based storage device 158 and store the metadata as highlight data 164, without having to retrieve other content from the content provider 124 or other remote sources. Such a configuration can conserve bandwidth and can make effective use of existing hardware that may already be available for the client device 106 .

[0064] Returning to FIG. 1A , in at least one embodiment, application server(s) 114 can identify different highlights and / or extract different metadata for different users 150 depending on individual user preferences and / or other parameters. The identified highlights and / or extracted metadata may be presented to the user 150 via any suitable output device, such as a display screen 152 of the client device 106. If desired, multiple highlights can be identified and organized into a highlight show along with associated metadata. Such highlight shows may be assembled into a “highlight reel” or set of highlights that are accessed via a menu and / or played for the user 150 according to a predetermined sequence. The user 150, in at least one embodiment, can control the highlight playback and / or delivery of associated metadata via input device(s) 151, for example, to: select particular highlights and / or metadata for display; pause, rewind, and fast-forward; skip to the next highlight; return to the beginning of the previous highlight within the highlight show; and / or perform other actions.

[0065] Additional details regarding such functionality are provided in the related US patent applications cited above.

[0066] In at least one embodiment, another data server(s) 118 is provided. The data server(s) 118 may respond to requests for data from any of the server(s) 102, 114, 116, for example, to obtain or provide card data 154, user data 155, and / or highlight data 164. In at least one embodiment, such information may be stored in any suitable storage device 153 accessible by the data server 118 and may come from any suitable source, such as the client device 106 itself, the content provider(s) 124, and / or the data provider(s) 122.

[0067] 1C, an alternative embodiment of system 180 is shown in which system 180 is implemented in a stand-alone environment. Similar to the embodiment shown in FIG. 1B, at least some of card data 154, user data 155, and highlight data 164 may be stored on a client-based storage device 158, such as a DVR. Alternatively, client-based storage device 158 may be a flash memory or hard drive, or other device integrated with or communicatively coupled to client device 106.

[0068] User data 155 may include preferences and interests of user 150. Based on such user data 155, system 180 can extract metadata in card data 154 and present it to user 150 in the manner described herein. Additionally or alternatively, metadata can be extracted based on objective criteria that are not based on information specific to user 150.

[0069] 1D , an overview of a system 190 having an architecture according to an alternative embodiment is shown. In FIG. 1D , the system 190 includes broadcast services, such as content provider(s) 124, content receivers in the form of client devices 106, such as television sets with STBs, video servers, such as analysis server(s) 116, that can ingest and stream television program content, and / or other client devices 106, such as mobile devices and laptops, that can receive and process television program content, all connected via a network, such as the communications network 104. A client-based storage device 158, such as a DVR, can be connected to any of the client devices 106 and / or other components and can store video streams, highlights, highlight identifiers, and / or metadata to facilitate identification and presentation of highlights and / or extracted metadata via any of the client devices 106.

[0070] The particular hardware architectures depicted in Figures 1A, 1B, 1C, and 1D are merely exemplary. Those skilled in the art will recognize that the techniques described herein can be implemented using other architectures. Many of the components depicted herein are optional and may be omitted, combined with, and / or replaced by other components.

[0071] In at least one embodiment, the system can be implemented as software written in any suitable computer programming language, whether in a stand-alone or client / server architecture, or it may be implemented and / or embedded in hardware.

[0072] Data Structure FIG. 2 is a schematic block diagram illustrating an example of data structures that may be incorporated into card data 154, user data 155, and highlight data 164, according to one embodiment.

[0073] As shown, card data 154 may include a record for each of multiple broadcast networks 202. For example, for each of distribution networks 202, card data 154 may include a predetermined card location 203 where the distribution network typically displays a card image within a video frame. The predetermined card location may be expressed as coordinates (e.g., Cartesian coordinates) that, for example, identify opposite corners of the location, identify a center, height, and width, and / or identify the location and / or size of the card image.

[0074] Additionally, card data 154 may include one or more video frame regions 204 that have been or are to be analyzed for card image extraction and interpretation. Each video frame region 204 may be extracted from a video frame of the video stream.

[0075] For each video frame region 204, card data 154 may also include one or more processed video frame regions 206, which may be generated by modifying the video frame region 204 in a manner that facilitates identification and / or extraction of the card image 207. For example, the processed video frame regions 206 may include one or more cropped, re-colored, segmented, expanded, or otherwise modified versions of each video frame region 204.

[0076] Each video frame region 204 may also have a card image 207 identified within and / or extracted from the video frame region 204. Each card image 207 may include text that may be interpreted to provide metadata associated with a particular time within the video stream.

[0077] The card data 154 may also include one or more interpretations 208 for each video frame region 204. Each interpretation 208 may be specific text that is believed to be represented in the associated card image 207 after some analysis has been performed to recognize and interpret characters that appear in the card image 207. The interpretations 208 may be used to obtain metadata from the card image 207.

[0078] As further shown, user data 155 may include records associated with users 150, each of which may include demographic data 212, preferences 214, viewing history 216, and purchasing history 218 for a particular user 150.

[0079] Demographic data 212 may include any type of demographic data, including, but not limited to, age, gender, location, nationality, religious affiliation, and / or education level.

[0080] Preferences 214 may include selections made by user 150 regarding their preferences. Preferences 214 may relate directly to the collection and / or display of highlights and metadata, or may be more general in nature. In either case, preferences 214 may be used to facilitate the identification and / or presentation of highlights and metadata to user 150.

[0081] The viewing history 216 may list television programs, video streams, highlights, web pages, search queries, sporting events, and / or other content retrieved and / or viewed by the user 150 .

[0082] The purchase history 218 may list products or services purchased or requested by the user 150 .

[0083] As further shown, the highlight data 164 may include records of highlights 220 , each of which may include a video stream 222 , an identifier, and / or metadata 224 for a particular highlight 220 .

[0084] Video stream 222 may include video depicting highlight 220, which may be obtained from one or more video streams of one or more events (e.g., by trimming the video streams to include only the video stream 222 associated with highlight 220). Identifier 223 may include a time code and / or other indicator that indicates where highlight 220 resides within the video stream of the event from which it was obtained.

[0085] In some embodiments, each recording of highlight 220 may include only one of video stream 222 and identifier 223. Highlight playback may be performed by playing user 150's video stream 222 or by using identifier 223 to play only the highlighted portion of the video stream of the event from which highlight 220 was captured.

[0086] The metadata 224 may include information about the highlight 220, such as the date of the event, the season, and the groups or individuals involved in the event or video stream from which the highlight 220 was captured, such as the team, players, coaches, anchors, broadcasters, and / or fans. Among other information, the metadata 224 for each highlight 220 may include the time 225, the phase 226, the clock 227, the score 228, and / or the frame number 229.

[0087] Time 225 may be the time within video stream 222 at which highlight 220 is captured or the time within video stream 222 associated with highlight 220 for which metadata is available. In some examples, time 225 may be the play time within video stream 222 associated with highlight 220 at which card image 207 including metadata 224 is displayed.

[0088] Phase 226 may be a phase of an event associated with highlight 220. More specifically, phase 226 may be a stage of a sporting event during which card image 207 containing metadata 224 is displayed. For example, phase 226 may be "third quarter," "second inning," "bottom half," or the like.

[0089] Clock 227 may be the game clock associated with highlight 220. More specifically, clock 227 may be the state of the game clock when time card image 207 is displayed that includes metadata 224. For example, clock 227 may read "15:47" for card image 207 that is displayed with 15 minutes and 47 seconds displayed on the game clock.

[0090] The score 228 may be the match score associated with the highlight 220. More specifically, the score 228 may be the score at the time the card image 207 containing the metadata 224 is displayed. For example, the score 228 may be "45-38," "7-0," or "30-love," etc.

[0091] The frame number 229 may be the number of the video frame in the video stream from which the highlight 220 is captured, or the number of the video frame in the video stream 222 associated with the highlight 220 that is most directly associated with the highlight 220. More specifically, the frame number 229 may be the number of such a video frame in which the card image 207 containing the metadata 224 is displayed.

[0092] The data structures depicted in Figure 2 are merely exemplary. Those skilled in the art will recognize that in performing highlight identification and / or metadata extraction, some of the data in Figure 2 can be omitted or replaced with other data. Additionally or alternatively, data not shown in Figure 2 can be used in performing highlight identification and / or metadata extraction.

[0093] Card Image 3A, there is shown a screenshot illustration of an example video frame 300 from a video stream with embedded information in the form of a card image 207, as often occurs in television programs of sporting events. FIG. 3A depicts a card image 207 at the bottom right of the video frame 300 and a second card image 320 extending along the bottom of the video frame 300. The card images 207, 320 may include embedded information such as the game phase, the current clock, and the current score.

[0094] In at least one embodiment, information within the card image 207, 320 is located and processed for automatic recognition and interpretation of embedded text within the card image 207, 320. The interpreted text may then be assembled into text metadata that describes the status of a sports match at a particular point in time within the sporting event timeline.

[0095] In particular, card image 207 may relate to the currently shown sporting event, while second card image 320 may include information about a different sporting event. In some embodiments, only card images containing information deemed relevant to the currently playing sporting event are processed for metadata generation. Accordingly, without limiting scope, the following exemplary description assumes that only card image 207 is processed. However, in alternative embodiments, it may be desirable to process multiple card images in a given video frame 300, even including card images related to other sporting events.

[0096] 3A , the card image 207 may provide several different types of metadata 224, including team name 330, score 340, prior team performance 350, current match stage 360, match clock 370, playing status 380, and / or other information 390. Each of these may be extracted from within the card image 207 and interpreted to provide the metadata 224 corresponding to the highlights 220, including video frames 300, and more specifically, the video frames 300 in which the card image 207 is displayed.

[0097] 3B is a series of screenshot diagrams depicting additional examples of video frames 392, 394, 396, and 398, respectively, with embedded card images 393, 395, 397, and 399, to illustrate additional examples of embedded card image locations in sports television programming. Different television networks may have different types, shapes, and frame locations of such card images embedded in video frames of television program content of sporting events.

[0098] Card image location and extraction 4 is a flowchart depicting a method 400 performed by an application (e.g., running on one of the client devices 106 and / or the analytic server 116) that receives the video stream 222 and performs on-the-fly processing of the video frames 300 to locate and extract the card images 207 and associated metadata, such as the card images 207 of FIG. 3 and their associated match status information, according to one embodiment. The system 100 of FIG. 1A is referred to as the system that performs the method 400 and subsequent methods. However, alternative systems, including but not limited to the system 160 of FIG. 1B, the system 180 of FIG. 1C, and / or the system 190 of FIG. 1D, may be used in place of the system 100 of FIG. 1A.

[0099] 4 may include receiving a video stream 222. In step 410, one or more video frames 300 of the video stream 222 may be read and decoded, for example, by resizing the video frames 300 to a standard size. In query 420, step 430, step 440, and / or query 450, the video frames 300 may be processed to locate card images within the frames. In step 460, the detected card images 207 may be processed to extract information by reading and interpreting the card images 207. Metadata 224 may be generated based on the information extracted from the card images 207.

[0100] In at least one embodiment, detection of one or more card images 207 present in decoded video frame 300 is performed by analyzing a single predetermined frame area. Alternatively, such detection can be performed by analyzing multiple predetermined frame areas if the approximate locations of card images 207 within decoded video frame 300 are not known in advance. Thus, query 420 can determine whether the locations of card images 207 within video frame 300 are known. For example, some broadcast networks may always show card images 207 in the same location within video frame 300. If the broadcast network is known, the locations of card images 207 may also be known. Alternatively, the locations of card images 207 within video frame 300 may not be known and may need to be confirmed by system 100.

[0101] If the location of card image 207 within video frame 300 is known according to query 420, method 400 may proceed to step 430, where the known portion or region of the video frame may be processed to isolate a quadrilateral shape typically associated with card image 207. If the location of card image 207 within video frame 300 is not known, method 400 proceeds to step 440, where video frame 300 is divided into multiple regions, which may be predetermined regions of video frame 300. The regions of video frame 300 are sequentially analyzed to determine which regions contain card images similar to card images 207, 395, 397, and / or 399.

[0102] For example, the particular region(s) of the video frame 300 that contains the card image 207 may be known for each of the various broadcast networks. If the broadcast network is not known, the system 100 may sequentially step through each region of the video frame 300 that is known to be used by the broadcast network for displaying the card image 207 until one of the regions is found.

[0103] If a card image 207 is found according to query 450, method 400 may proceed to step 460, where the card image 207 is processed and information is extracted from the card image 207 to provide metadata 224. If a card image 207 is not found according to query 450, method 400 may return to step 410, where a new video frame may be loaded, decoded, and then analyzed for the presence of a card image 207.

[0104] As mentioned above, method 400 may, in some embodiments, be performed in real time while user 150 is watching a program (e.g., while video stream 222 corresponding to highlight 220 is being presented). Thus, method 400 may be performed in the background for each video frame 300 as it is being decoded for playback for user 150. There may be some delay as system 100 locates, extracts, and interprets card images 207. Thus, in this application, the presentation of metadata extracted from card images 207 is considered to be “real time” even if the presentation of metadata 224 lags behind the playback of the video frames 300 from which the metadata was obtained (e.g., by several video frames 300, a delay that is not perceptible or distracting to user 150).

[0105] Figure 5 is a flowchart illustrating in more detail step 440 for processing predetermined regions of video frame 300 to detect actionable card images 207 from Figure 4. Each predetermined region may represent an approximate location where a card image 207 may be present.

[0106] In at least one embodiment, the predetermined region within the decoded video frame 300 is generated based on knowledge of the approximate location of the card image 207 used by various television networks engaged in the broadcast of sporting event television programs, as previously described. Such television networks are known to use one or more regions of the video frame 300 to deliver visual and textual data within the frame via the card image 207.

[0107] At step 510, sequential processing of the regions may begin. At step 520, one of the regions may be processed to determine whether a valid card image 207 is present in that region. A query 530 may determine whether a card image 207 is found in that region. If found, the region may be further processed to extract the card image 207. The located card images 207 may be further processed for automatic recognition and interpretation of embedded text. Such interpreted text may then be further assembled into text metadata describing the status of a sporting event (e.g., a match) at a particular point in time on the sporting event timeline. In at least one embodiment, the options available for text rendering are based on the type of card image 207 detected within the video frame 300, which may be determined by the system 100 during the location and / or extraction of the card image 207. Additionally or alternatively, the options available for text rendering may be based on pre-assigned meanings of selected fields present within the particular type of detected card image 207.

[0108] If no card image 207 is found in the region, query 550 may check whether the region is the last region in video frame 300. If it is not the last region, system 100 may proceed to the next region in step 560 and then repeat the process for the next region according to step 520. If the region is the last region in video frame 300, video frame 300 may not contain a valid card image 207, and system 100 may proceed to the next video frame 300.

[0109] Automatic detection and location of card image quadrilaterals 6 is a flowchart illustrating a method 600 for top-level processing for valid card image quadrangle detection in a specified area of ​​a decoded video frame 300, according to one embodiment. Method 600 may be performed on a video frame region at a predetermined location, according to step 430 of FIG. 4, or on a video frame region identified via sequential processing of multiple regions of video frame 300, according to step 440 of FIGS. 4 and 5.

[0110] First, in step 610, the decoded video frame 300 may be cropped to a smaller area that includes a specified video frame region, providing a cropped image. In step 620, the cropped image may be segmented using any suitable segmentation algorithm, such as graph-based segmentation (e.g., “Efficient Graph-Based Image Segmentation,” P. Felzenszwalb, D. Huttenlocher, Int. Journal of Computer Vision, 2004, Vol. 59), and all generated segments may be color-coded and enumerated to provide a segmented image. Further processing of the segmented image may include removing background material surrounding the assumed quadrilateral defining the card image 207. In at least one embodiment, method 600 may proceed to step 630, where all pixels of segments adjacent to the boundary of the segmented cropped image are set to a black level. In step 640, all pixels of remaining inner segments of the segmented cropped image are set to a white level. In step 650, the two-color cropped image with the partially removed background may be passed on for further processing for accurate card image quadrilateral drawing.

[0111] FIG. 7 is a flowchart illustrating a method 700 for more accurate card image rectangle determination, according to one embodiment. First, in step 710, a cropped image with a partially removed background (e.g., generated in step 640 of FIG. 6 ) is converted to a gray image. The gray image may then be blurred, and in step 720, an edge detection process may be performed to generate an edge image with detected edges. Next, in step 730, the edge image may be processed for contour detection, and the resulting contour image may be further processed to approximate contours with closed polygons. Thereafter, in step 740, the contour / polygon image may be processed to determine a minimum rectangular perimeter that encloses all existing contours. The above steps may generate a rectangular enclosure that potentially contains the card image 207. However, this enclosure may be larger than the card image rectangle due to artifacts generated during the process of segmenting the cropped image. Therefore, in at least one embodiment, further adjustments are performed to compress this intermediate rectangular shape into the minimum rectangular area that contains the card image 207.

[0112] 8 is a flowchart depicting an example method 800 for adjusting the quadrilateral boundary of an enclosure that encompasses all detected contours (e.g., generated by step 740 of FIG. 7), according to one embodiment. The enclosure may enclose an interior area such that the majority of the surrounding pixels of the new, indented enclosure are the same pixel intensity (e.g., white in this particular example). Method 800 may remove undesirable interior region artifacts that extend outward, thus providing a new, more robust enclosure that may contain a valid card image 207.

[0113] The method of FIG. 8 may begin at step 810, where a contour perimeter is received. The contour perimeter may be a rectangular image. In step 820, system 100 may "walk" around the boundary of a rectangular enclosure image encompassing all detected contours (or quadrilateral images), and in step 830, count the black-level pixel values ​​for each boundary edge, i.e., top, bottom, left, and right. Next, in step 840, if any boundary edge contains more black-level pixels than a predetermined count, that edge is moved inward by one pixel, providing an adjusted quadrilateral area 850. This process continues until query 860 determines that the black-level pixel counts for all edges of the squeezed quadrilateral are below a predetermined threshold. The resulting squeezed rectangle represents the enclosure of a potential card image and may be verified in the processing steps described in connection with the flowchart of FIG. 9.

[0114] 9 is a flowchart depicting an exemplary method 900 for quadrilateral verification of card images, according to one embodiment. Method 900 may include analysis of three different image areas (pixel counts): a cropped image area (e.g., generated in step 610 of FIG. 6 ), a rectangular enclosure image area encompassing all detected contours (e.g., generated in step 740 of FIG. 7 ), and an area of ​​the image with an adjusted (pressed) quadrilateral boundary (e.g., generated in one or more iterations of step 840 of FIG. 8 ). Three parameters (A, B, C) may be generated in steps 910, 920, and 930, respectively, as follows: A = total pixel count of the cropped image area; B = total pixel count of the binary image around the contour; and C = black value pixel count of the binary image around the adjusted contour.

[0115] Next, in step 940, a weighted comparison of these three parameters can be performed so that if a valid card image rectangle is to be detected, the non-black pixel area of ​​the pressed rectangle must be present in a certain percentage with respect to the other two parameters. In step 950, if a valid card image 207 is detected based on the above weighted comparison, a flag is set to true. If, according to query 960, the flag is set to true, the system 100 can proceed to step 970, where the card image 207 (and / or a processed version of the card image 207) is passed to card image internal content processing. If, in step 950, a valid card image 207 is not detected, the flag is set to false, and according to query 960, the system 100 can proceed to the next specified frame region or to the next video frame 300 in step 980 to search for a valid card image 207 therein.

[0116] FIG. 10 is a flowchart depicting an exemplary method 1000 for optional stabilization of the left (or any other) boundary of a very elongated card image shape, according to one embodiment. The process may begin in step 1010, where system 100 extends the horizontal card image 207 to the cropped frame edge. In step 1020, system 100 may detect straight vertical lines within this extended image. The process may further include selecting detected vertical lines of a predetermined length and calculating selected sparse vertical line markers in step 1030. Finally, in step 1040, a marker (if any, found in the immediate vicinity) to the left of the original position of the card image 207 is selected, and in step 1050, the left edge of the rectangle is moved to the position of the marker further to the left of the original edge position of the detected card image rectangle. In step 1060, the card image rectangle is adjusted accordingly, and an updated region of interest (ROI) for the card is returned.

[0117] Internal processing of card images for information extraction In at least one embodiment, an automated process is implemented that includes receiving a digital video stream (which may include one or more highlights of a broadcast sporting event), analyzing one or more video frames of the digital video stream for the presence of a card image 207, extracting the card image 207, locating a text box within the card image 207, and interpreting the text present in the text box to create metadata 224 that associates content from the card image 207 with video highlights from the analyzed digital video stream.

[0118] FIG. 11 is a flowchart illustrating a method 1100 for performing text extraction from a card image 207, according to one embodiment. In step 1110, the extracted card image 207 may be resized to a standard size. Next, in step 1120, the resized card image 207 may be preprocessed using a series of filters, including, for example, contrast enhancement, bilateral and median filtering for noise reduction, and gamma correction followed by illumination compensation. In at least one embodiment, in step 1130, an "extreme region filter" with a two-stage classifier is created (e.g., L. Neumann, J. Matas, "Real-Time Scene Text Localization and Recognition," 5th IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, June 2012), and in step 1140, a cascade classifier is applied to each image channel of the card image 207. Next, in step 1150, character groups are detected and word box groups are extracted.

[0119] In at least one embodiment, multiple text strings (text boxes) are identified within the card image 207, and the position and size of each character within the string of characters associated with the text box is detected. The text strings from various fields of the card image 207 are then processed and interpreted to generate corresponding metadata 224, thus providing real-time information related to the current sporting event television program and the current timeline associated with the processed embedded card image 207.

[0120] 12 is a flowchart illustrating a method 1200 for processing and interpreting text strings according to one embodiment. In step 1210, the detected and extracted card image 207 may be processed, and interpreted text may be selected from a group of character bounding boxes within the card image 207. Next, in step 1220, text may be extracted, and the extracted text may be read and interpreted, for example, via optical character recognition (e.g., "An Overview of the Tesseract OCR Engine," R. Smith, Proceedings ICDAR '07, Vol. 02, September 2007). In step 1230, metadata 224 may be generated and structured. Next, intraframe information from the card image 207 is combined with the video highlight text and visual metadata.

[0121] The present system and method have been described in particular detail with respect to the envisioned embodiment. Those skilled in the art will appreciate that the system and method may be implemented in other embodiments. First, the specific naming of components, terminology capitalization, attributes, data structures, or any other programming or structural aspect is not required or important, and mechanisms and / or functions may differ in name, format, or protocol. Furthermore, the system may be implemented via a combination of hardware and software, entirely within hardware elements, or entirely within software elements. Also, the specific division of functionality among various system components described herein is merely exemplary and not required. Functionality performed by a single system component may instead be performed by multiple components, and functionality performed by multiple components may instead be performed by a single component.

[0122] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. The appearances of the phrases "in one embodiment" or "in at least one embodiment" in various places in the specification do not necessarily all refer to the same embodiment.

[0123] Various embodiments may include any number of systems and / or methods for implementing the above-described techniques, either alone or in any combination. Another embodiment includes a computer program product including a non-transitory computer-readable storage medium and computer program code encoded on the medium for causing a processor within a computing device or other electronic device to implement the above-described techniques.

[0124] Some portions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computing device's memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is herein generally conceived to be a self-consistent sequence of steps (instructions) leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It is sometimes convenient, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. Further, without loss of generality, it is also convenient to refer to specific arrangements of steps requiring physical manipulations of physical quantities as modules or code devices.

[0125] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, and as will be apparent from the description that follows, throughout this specification, descriptions utilizing terms such as "processing" or "computing" or "calculating" or "displaying" or "determining" will be understood to refer to operations and processes of a computer system or similar electronic computing module and / or device, and to mean manipulating and transforming data that are represented as physical (electronic) quantities in the computer system's memory or registers or other such storage, transmission, or display devices.

[0126] Certain aspects include process steps and instructions described herein in the form of an algorithm. It should be noted that the process steps and instructions may be embodied in software, firmware, and / or hardware, and that if embodied in software, may be downloaded to reside on and operate from a variety of platforms for use by a variety of operating systems.

[0127] This document also relates to apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, or may include a general-purpose computing device selectively activated or reconfigured by a computer program stored on the computing device. Such a computer program may be stored on a computer-readable storage medium, such as a floppy disk, optical disk, CD-ROM, DVD-ROM, magneto-optical disk, read-only memory (ROM), random-access memory (RAM), EPROM, EEPROM, flash memory, solid-state drive, magnetic or optical card, application-specific integrated circuit (ASIC), or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus. The program and its associated data may also be hosted and executed remotely, such as on a server. Furthermore, the computing devices referred to herein may include a single processor or may be architectures employing multiple processor designs to increase computing power.

[0128] The algorithms and displays presented herein are not inherently related to any particular computing device, virtualization system, or other apparatus. Various general-purpose systems may also be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the description provided herein. Moreover, the systems and methods are not described with reference to any particular programming language. It will be understood that a variety of programming languages ​​can be used to implement the teachings described herein, and any references above to specific languages ​​are provided for the purpose of enabling and best mode disclosure.

[0129] Accordingly, various embodiments include software, hardware, and / or other elements, or any combination or plurality of elements, for controlling a computer system, computing device, or other electronic device. Such an electronic device may include, for example, a processor, input devices such as a keyboard, mouse, touchpad, trackpad, joystick, trackball, microphone, and / or any combination thereof, output devices such as a screen and / or speaker, long-term storage such as memory, magnetic storage, and / or optical storage, and / or network connectivity, according to techniques known in the art. Such electronic devices may be portable or non-portable. Examples of electronic devices that can be used to implement the described systems and methods include desktop computers, laptop computers, televisions, smartphones, tablets, music players, audio devices, kiosks, set-top boxes, gaming systems, wearable devices, home electronic devices, and / or server computers, etc. The electronic device may use any operating system such as, but not limited to, Linux, Microsoft Windows available from Microsoft Corporation, Redmond, Washington, Mac OS X available from Apple Inc., Cupertino, California, iOS available from Apple Inc., Cupertino, California, Android available from Google Inc., Mountain View, California, and / or any other operating system adapted for use on the device.

[0130] While a limited number of embodiments have been described herein, those skilled in the art, having the benefit of the above description, will appreciate that other embodiments may be devised. Furthermore, it should be noted that the language used herein has been chosen primarily for ease of reading and educational purposes, and may not have been chosen to delineate or limit the subject matter. Accordingly, the present disclosure is intended to be illustrative, but not limiting, in scope.

Claims

1. 1. A method for extracting metadata from a video stream, said method comprising: storing video frames of the video stream in a data store; and automatically identifying and extracting, in a processor, a card image embedded in at least one of the video frames, wherein the identifying and extracting includes: identifying a predetermined location within the video frame that defines a video frame region that includes the card image; and identifying and extracting by performing at least one of sequentially processing a plurality of regions of the video frame to identify the region of the video frame that includes the card image; analyzing the card image in the processor to obtain metadata; storing the metadata in association with at least one of the video frames in the data store.

2. the video stream comprises a broadcast of a sporting event; the video frames constitute highlights deemed to be of particular interest to one or more users; The method of claim 1 , wherein the metadata describes a status of the sporting event in the highlight.

3. The method of claim 2 , further comprising presenting metadata during display of the highlights at an output device.

4. The method of claim 3 , wherein automatically identifying and extracting card images and analyzing the card images to obtain the metadata is performed for a highlight while the highlight is being displayed.

5. The method of claim 1 , further comprising locating and extracting the card image from the video frame region.

6. The method of claim 5 , wherein locating and extracting the card image from the video frame region comprises cropping the video frame to isolate the video frame region.

7. Locating and extracting the card image from the video frame region; segmenting the video frame region, or a processed version of the video frame region, to generate a segmented image; and modifying pixel values ​​of segments adjacent to boundaries of the segmented image.

8. The method of claim 5 , wherein locating and extracting the card image from the video frame region includes removing background from the video frame region or a processed version of the video frame region.

9. Locating and extracting the card image from the video frame region; generating an edge image based on the background-removed video frame region; Finding contours in the edge image; approximating the contour as a polygon; and extracting an area enclosed by the smallest rectangular perimeter that encompasses all of the contours to generate a perimeter rectangle image.

10. repeatedly, Counting color-corrected pixels for each edge of the surrounding rectangular image; 10. The method of claim 9, further comprising moving inward any boundary edges where the number of color-corrected pixels exceeds a threshold.

11. a first number of pixels within the video frame region; a second number of pixels in the surrounding rectangular image; and Counting a third number of pixels in the adjusted surrounding rectangle image; 11. The method of claim 10, further comprising validating the detected quadrilaterals within the region by comparing the first number, the second number, and the third number to determine whether a possible quadrilateral within the region is feasible.

12. The method of claim 5 , wherein locating and extracting the card image from the video frame region includes adjusting a left border of the card image.

13. 1. A non-transitory computer-readable medium for extracting metadata from a video stream, the medium comprising instructions stored therein, the instructions, when executed by a processor, performing: storing video frames of said video stream in a data store; automatically identifying and extracting a card image embedded in at least one of the video frames, wherein the identifying and extracting comprises: identifying a predetermined location within the video frame that defines a video frame region that includes the card image; and an identifying and extracting step performed by performing at least one of sequentially processing a plurality of regions of the video frame to identify the region of the video frame that includes the card image; analyzing the card image to obtain metadata; causing the data store to store the metadata in association with at least one of the video frames.

14. the video stream comprises a broadcast of a sporting event; the video frames constitute highlights deemed to be of particular interest to one or more users; The non-transitory computer-readable medium of claim 13 , wherein the metadata describes a status of the sporting event during the highlights.

15. 15. The non-transitory computer-readable medium of claim 14, further comprising instructions stored therein that, when executed by the processor, cause an output device to present the metadata during display of the highlight.

16. 16. The non-transitory computer-readable medium of claim 15, wherein automatically identifying and extracting card images and analyzing the card images to obtain metadata is performed for a highlight while the highlight is being displayed.

17. 14. The non-transitory computer-readable medium of claim 13, further comprising instructions stored therein that, when executed by the processor, locate and extract the card image from the video frame region.

18. 20. The non-transitory computer-readable medium of claim 17, wherein locating and extracting the card image from the video frame region comprises cropping the video frame to isolate the video frame region.

19. Locating and extracting the card image from the video frame region; segmenting the video frame region, or a processed version of the video frame region, to generate a segmented image; and modifying pixel values ​​of segments adjacent to boundaries of the segmented image.

20. 20. The non-transitory computer-readable medium of claim 17, wherein locating and extracting the card image from the video frame region includes removing background from the video frame region or a processed version of the video frame region.

21. Locating and extracting the card image from the video frame region; generating an edge image based on the background-removed video frame region; Finding contours in the edge image; approximating the contour as a polygon; and extracting an area enclosed by a smallest rectangular perimeter that encompasses all of the contours to generate a perimeter rectangle image.

22. 1. A system for extracting metadata from a video stream, the system comprising: a data store configured to store video frames of the video stream; 1. A processor, comprising: automatically identifying and extracting a card image embedded in at least one of the video frames, wherein the identifying and extracting includes: identifying a predetermined location within the video frame that defines a video frame region that includes the card image; and identifying and extracting by performing at least one of sequentially processing a plurality of regions of the video frame to identify the region of the video frame that includes the card image; analyzing the card image to obtain metadata; a processor configured to: The system, wherein the data store is further configured to store the metadata in association with at least one of the video frames.

23. the video stream comprises a broadcast of a sporting event; the video frames constitute highlights deemed to be of particular interest to one or more users; 23. The system of claim 22, wherein the metadata describes a status of the sporting event in the highlight.

24. 24. The system of claim 23, further comprising an output device configured to present the metadata during display of the highlights.

25. 25. The system of claim 24, wherein the processor is further configured to automatically identify and extract the card image and analyze the card image to obtain the metadata for the highlight during display of the highlight.

26. 23. The system of claim 22, wherein the processor is further configured to locate and extract the card image from the video frame region.

27. 27. The system of claim 26, wherein the processor is further configured to locate and extract the card image from the video frame region by cropping the video frame to isolate the video frame region.

28. the processor: segmenting the video frame region, or a processed version of the video frame region, to generate a segmented image; 27. The system of claim 26, further configured to locate and extract the card image from the video frame region by: modifying pixel values ​​of segments adjacent to boundaries of the segmented image.

29. 27. The system of claim 26, wherein the processor is further configured to locate and extract the card image from the video frame region by removing background from the video frame region or a processed version of the video frame region.

30. the processor: generating an edge image based on the background-removed video frame region; finding contours in said edge image; approximating the contour as a polygon; and 27. The system of claim 26, further configured to locate and extract the card image from the video frame area by extracting an area enclosed by a smallest rectangular perimeter that encompasses all of the contours to generate a surrounding rectangle image.