Video Processing for Embedded Information Card Position Identification and Content Extraction

The system automatically detects and interprets information cards in sports broadcasts to generate metadata for real-time highlight creation and interactive services, addressing the lack of such capabilities in existing television systems.

JP7706609B2Active Publication Date: 2025-07-11STATS LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024102182
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-14
Filing Date
2024-06-25
Publication Date
2025-07-11
Estimated Expiration
2039-05-15

AI Technical Summary

Technical Problem

Existing television systems lack the ability to automatically identify and extract metadata from embedded information cards in sports broadcasts for real-time generation of highlights and interactive services.

Method used

A method and system that utilize computer vision techniques to detect and interpret information cards within video frames, extracting text strings and generating metadata in real-time to create customized highlights and interactive services.

Benefits of technology

Enables real-time extraction of metadata from embedded information cards in sports broadcasts, allowing for the generation of customized highlights and interactive services, enhancing user engagement and interaction with sports content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706609000001
    Figure 0007706609000001
  • Figure 0007706609000002
    Figure 0007706609000002
  • Figure 0007706609000003
    Figure 0007706609000003
Patent Text Reader

Abstract

To enable metadata of one or more highlights of a video stream to be extracted from one or more card images embedded in the video stream.SOLUTION: Highlights may be segments of a video stream having specific interest, such as a broadcast of a sporting event. According to one method, video frames of the video stream are stored. One or more information cards embedded in a decoded video frame may be detected by analyzing one or more prescribed video frame areas. Next, image segmentation, edge detection, and / or identification of a closed contour may be executed to an identified video frame area. Further processing may include acquiring the smallest rectangular surrounding area surrounding all remaining segments, and after that, further, processing may be performed in order to determine an accurate boundary of the information cards. Metadata may be acquired by analyzing a card image, and the metadata may be stored in association with at least one of the video frames.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 673,412 (Attorney Docket No. THU010 - PROV) for "Machine Learning for Recognizing and Interpreting Embedded Information Card Content", filed on May 18, 2018, the entire disclosure of which is incorporated herein by reference.

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 673,411 (Attorney Docket No. THU009 - PROV) for "Video Processing for Enabling Sports Highlights Generation", filed on May 18, 2018, the entire disclosure of which is incorporated herein by reference.

[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 673,413 (Attorney Docket No. THU012 - PROV) for "Video Processing for Embedded Information Card Localization and Content Extraction", filed on May 18, 2018, the entire disclosure of which is incorporated herein by reference.

[0004] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 680,955 (Attorney Docket No. THU007 - PROV) for "Audio Processing for Detecting Occurrences of Crowd Noise in Sporting Event Television Programming", filed on June 5, 2018, the entire disclosure of which is incorporated herein by reference.

[0005] This application claims the benefit of U.S. Provisional Patent Application Ser. No. 62 / 712,041 (Attorney Docket No. THU006-PROV), filed Jul. 30, 2018, entitled “Audio Processing for Extraction of Variable Length Disjoint Segments from Television Signal”, the entire contents of which are incorporated herein by reference.

[0006] This application claims the benefit of U.S. Provisional Patent Application Ser. No. 62 / 746,454 (Attorney Docket No. THU016-PROV), filed Oct. 16, 2018, entitled “Audio Processing for Detecting Occurrences of Loud Sound Characterized by Short-Time Energy Bursts”, the entire contents of which are incorporated herein by reference.

[0007] This application claims the benefit of U.S. Utility Patent Application Ser. No. 16 / 411,710 (Attorney Docket No. THU010), filed May 14, 2019, entitled “Machine Learning for Recognizing and Interpreting Embedded Information Card Content”, the entire contents of which are incorporated herein by reference.

[0008] This application claims the benefit of U.S. Utility Patent Application Ser. No. 16 / 411,704 (Attorney Docket No. THU009), filed May 14, 2019, entitled “Video Processing for Enabling Sports Highlights Generation”, the entire contents of which are incorporated herein by reference.

[0009] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,713 (Attorney Docket No. THU012), filed on May 14, 2019, entitled "Video Processing for Embedded Information Card Localization and Content Extraction", which is hereby incorporated by reference in its entirety.

[0010] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,915, filed on August 31, 2012, entitled "Generating Excitement Levels for Live Performances", which was issued as U.S. Patent No. 9,060,210 on June 16, 2015, and is hereby incorporated by reference in its entirety.

[0011] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,927, filed on August 31, 2012, entitled "Generating Alerts for Live Performances", which was issued as U.S. Patent No. 8,842,007 on September 23, 2014, and is hereby incorporated by reference in its entirety.

[0012] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,933, filed on August 31, 2012, entitled "Generating Teasers for Live Performances", which was issued as U.S. Patent No. 8,595,763 on November 26, 2013, and is hereby incorporated by reference in its entirety.

[0013] This application is related to U.S. Utility Patent Application Serial No. 14 / 510,481 (Attorney Docket No. THU001), filed on October 9, 2014, entitled "Generating a Customized Highlight Sequence Depicting an Event", and is hereby incorporated by reference in its entirety.

[0014] This application claims the benefit of U.S. Patent Application Serial No. 14 / 710,438 (Attorney Docket No. THU002), filed May 12, 2015, entitled "Generating a Customized Highlight Sequence Depicting Multiple Events", which is hereby incorporated by reference in its entirety.

[0015] This application claims the benefit of U.S. Patent Application Serial No. 14 / 877,691 (Attorney Docket No. THU004), filed October 7, 2015, entitled "Customized Generation of Highlight Show with Narrative Component", which is hereby incorporated by reference in its entirety.

[0016] This application claims the benefit of U.S. Patent Application Serial No. 15 / 264,928 (Attorney Docket No. THU005), filed September 14, 2016, entitled "User Interface for Interaction with Customized Highlight Shows", which is hereby incorporated by reference in its entirety.

[0017] This document relates to techniques that enable an embedded software application to utilize multimedia content in order to identify multimedia content and associated information on a television device or video server that distributes multimedia content, and to provide content and services in synchronization with the multimedia content. Various embodiments relate to methods and systems for providing automated video and audio analysis used to identify and extract information within sports television video content and create metadata associated with video highlights for in-game and post-game review of sports television video content. BACKGROUND OF THE INVENTION

[0018] Extended television applications, such as interactive advertisements and enhanced program guides with interactive applications before, during, and after a game, have been envisioned for a long time. Existing cable systems originally designed for broadcast television are required to support hosting new applications and services, including interactive television services and extended (interactive) program production guides.

[0019] Several frameworks for enabling extended television applications have been standardized. Examples include OpenCable (商標) Extended TV application messaging specifications and Tru2way specifications, which refer to interactive digital cable services delivered over a cable video network and include features such as interactive program guides, interactive advertisements, and games. Further, the cable operator's "OCAP" program provides interactive services such as e-commerce shopping, online banking, electronic program guides, and digital video recording. These efforts have enabled a first-generation video synchronization application synchronized with video content distributed by program producers / broadcasters, providing interactivity with additional data in television program production.

[0020] Recent developments in video / audio content analysis technology and corresponding mobile devices have opened up a series of new possibilities in the development of advanced applications that operate in synchronization with live TV program events. These new technologies, along with advances in computer vision and video processing, and the improved computing power of the latest processors, have made it possible to generate highlights of advanced program content with metadata in real time. SUMMARY OF THE INVENTION

[0021] In a sports television broadcast program, a method and system are presented for automatically finding the position of an information card (a "card image"), such as an information scoreboard, within one video frame or multiple video frames. Also described are methods and systems for identifying text strings within various fields of a located card image and reading and interpreting text information from the various fields of the located card image.

[0022] In at least one embodiment, the detection, location, and reading of the card image are performed synchronously with respect to the presentation of sports television program content. In at least one embodiment, a digital video stream is received, one or more frames of the digital video stream are analyzed, and an automated process is provided for automatically detecting and locating a card image quadrilateral. In another embodiment, an automated process is provided for analyzing one or more located card images, recognizing and extracting text strings (e.g., within text boxes), and reading information from the extracted text boxes.

[0023] In yet another embodiment, the detected text strings associated with specific fields within the card image are interpreted, thus immediately providing information within the game related to the content of the television broadcast of the sports event. The extracted in-frame information may be used to generate metadata related to custom video content automatically created as a set of highlights of the broadcast television program content associated with visual and text data.

[0024] In at least one embodiment, a method for extracting metadata from a video stream may include storing at least one portion of the video stream in a data store. In a processor, one or more card images embedded in at least one of the video frames may be automatically identified and extracted by performing at least one of identifying a predetermined position in the video frame that defines a video frame region including the card image, and sequentially processing a plurality of regions of the video frame to identify a video frame region including the card image. In a processor, the card image may be analyzed to obtain metadata, and the metadata may be stored in the data store in association with at least one of the video frames.

[0025] In at least one embodiment, the video stream may be a broadcast of a sports event. The video frame may constitute a highlight that is considered to have a particular interest for one or more users. The metadata may describe the status of the sports event during the highlight.

[0026] In at least one embodiment, the method may further include presenting the metadata during the display of the highlight on an output device. Automatically identifying and extracting the card image and analyzing the card image to obtain metadata may be performed during the display of the highlight for the highlight.

[0027] In at least one embodiment, the method may further include locating and extracting a card image from a video frame region. Locating and extracting a card image from a video frame region may include trimming the video frame to separate the video frame region. Alternatively or additionally, locating and extracting a card image from a video frame region may include segmenting the video frame region or a processed version of the video frame region to generate a segmented image, and modifying pixel values of segments adjacent to the boundaries of the segmented image. Alternatively or additionally, locating and extracting a card image from a video frame region may include removing a background from the video frame region or a processed version of the video frame region. Alternatively or additionally, locating and extracting a card image from a video frame region may include generating an edge image based on the video frame region, finding contours within the edge image, approximating the contours as polygons, and extracting a region surrounded by a minimum rectangular perimeter that encloses all of the contours to generate a surrounding rectangular image.

[0028] In at least one embodiment, the method may further include, for each edge of the surrounding rectangular image, iteratively counting color-corrected pixels and moving any boundary edges inward by the number of color-corrected pixels that exceed a threshold.

[0029] In at least one embodiment, the method may further include verifying a quadrilateral detected within a region by counting a first number of pixels within the video frame region, a second number of pixels within the surrounding rectangular image, and a third number of pixels within the adjusted surrounding rectangular image. The first number, the second number, and the third number may be compared to determine whether a quadrilateral assumed within the region is viable.

[0030] In at least one embodiment, locating and extracting a card image from a video frame region may include adjusting a left (or other) boundary of the card image.

[0031] Further details and variations are described herein.

Brief Description of the Drawings

[0032] The accompanying drawings, together with the description, illustrate several embodiments. Those skilled in the art will recognize that the specific embodiments shown in the drawings are merely exemplary and are not intended to limit the scope.

Fig. 1A

Fig. 1B

Fig. 1C

Fig. 1D

Fig. 2

Fig. 3A

Fig. 3B

Fig. 4

Fig. 5

Fig. 6

Fig. 7

Fig. 8

Fig. 9

Fig. 10

Fig. 11

Fig. 12

DETAILED DESCRIPTION OF THE INVENTION

[0033] Definitions The following definitions are presented for illustrative purposes only and are not intended to limit the scope. · Event: For the purposes of the description in this specification, the term "event" refers to a game, session, match, series, performance, program, and / or concert, etc., or a part thereof (act, period, quarter, half, inning, scene, or chapter). An event may be a sports event, an entertainment event, or a specific performance of a single individual or a subset of multiple individuals within a larger group of event participants. Examples of events other than sports include television shows, news bulletins, socio-political events, natural disasters, movies, plays, radio programs, podcasts, audiobooks, online content, and / or musical performances, etc. An event can have any length. For illustrative purposes, this specification often describes this technology from the perspective of sports events, but those skilled in the art will recognize that this technology can also be used in other contexts, including highlight shows of any audiovisual, audio, qualification, graphics-based, interactive, non-interactive, or text-based content. Therefore, the use of the term "sports event" and any other sports-specific terms in this description is intended to illustrate one assumed embodiment, but is not intended to limit the scope of the described technology to that one embodiment. Rather, such terms should be considered to extend to any suitable non-sports context to which this technology is applicable. For ease of explanation, the term "event" is also used to refer to a report or representation of an event, such as a visual record of an event, or any other content item that includes a report, description, or depiction of an event. · Highlight: An excerpt or portion of an event, or content associated with an event, that is considered to have a particular interest to one or more users. A highlight can have any length. Generally, the techniques described herein provide a mechanism for identifying and presenting a customized set of highlights (which can be selected based on specific characteristics and / or user preferences) for any suitable event. The term "highlight" is also used to refer to a report or representation of a highlight, such as a visual recording of the highlight, or any other content item that includes a report, explanation, or depiction of the highlight. A highlight need not be limited to the rendering of the event itself, but can include other content associated with the event. For example, in the case of a sports event, highlights can include in-game audio / video, as well as other content such as interviews, analysis, and / or commentary before, during, and after the game. Such content can be recorded from a linear television (e.g., as part of a video stream depicting the event itself) or retrieved from any number of other sources. For example, various types of highlights can be provided, including occurrences (plays), strings, possessions, and sequences, all of which are defined below. A highlight need not have a fixed duration, but can incorporate start offsets and / or end offsets, as described below. · Content delimiter: One or more video frames that indicate the start or end of a highlight. · Occurrence: Something that occurs during an event. Examples include a goal, a play, a down, a hit, a save, a shot on goal, a basket, a steal, a snap or an attempt to snap, a near miss, an altercation, the start or end of a game, a quarter, a half, a period, or an inning, a pitch, a penalty, an injury, a dramatic event in an entertainment event, a song, and / or a solo, etc. An occurrence can also be an abnormal event such as a power outage and / or an incident with an uncontrollable fan. Detection of such occurrences can be used as a basis for determining whether to designate a specific portion of a video stream as a highlight. For ease of naming, an occurrence is also referred to herein as a "play", but such usage should not be construed as limiting. An occurrence may have any length, and the representation of an occurrence may have various lengths. For example, as described above, an extended representation of an occurrence may include video depicting the time periods immediately before and after the occurrence, while a simple representation may include only the occurrence itself. Any intermediate representation can also be provided. In at least one embodiment, the selection of the duration for representing an occurrence may vary depending on user preference, available time, the determined excitement level for the occurrence, the importance of the occurrence, and / or any other factor. · Offset: The amount by which the length of a highlight is adjusted. In at least one embodiment, a start offset and / or an end offset can be provided to adjust the start time and / or the end time of a highlight respectively. For example, if the highlight depicts a goal, the highlight may be extended (via the end offset) for a few seconds to include the celebration and / or the reaction of the fans following the goal. The offset can be configured to change automatically or manually based on, for example, the time available for the highlight, the importance and / or excitement level of the highlight, and / or any other suitable factor. · String: A series of occurrences that are linked or related to each other in some way. Occurrences may occur within a possession (defined below), or may span multiple possessions. Occurrences may occur within a sequence (defined below), or may span multiple sequences. Occurrences may be linked or related because they have some thematic or narrative connection to each other, or because one leads to another, or for any other reason. An example of a string is a set of paths leading to a goal or basket. This should not be confused with a "text string" which has the meaning normally assigned in the field of computer programming. · Possession: A portion of an event delimited at any time. The distinction between the start / end times of a possession may vary depending on the type of event. In the case of certain sports events (e.g., basketball or soccer) where one team can be offensive while the other is defensive, a possession can be defined as the time period during which one team has the ball. In sports where possession of the puck or ball is more fluid, such as hockey or soccer, a possession is considered to extend to the time period during which one team has substantial control of the puck or ball, ignoring momentary contact by the other team (such as a blocked shot or save). In baseball, a possession is defined as a half - inning. In soccer, a possession can include several sequences where the same team has the ball. In the case of other types of sports events and non - sports events, the term "possession" may be somewhat of a misnomer, but is still used herein for illustrative purposes. Examples in non - sports contexts include chapters, scenes, actions, or TV segments. For example, in the context of a music concert, a possession may correspond to the performance of a single song. A possession can contain any number of occurrences. · Sequence: A portion of time of an event delimited by the time of an event that includes the time period of one continuous action. For example, in a sports event, a sequence may start at the beginning of an action (such as a face-off or a tip-off) and end when a whistle is blown to indicate an interruption of the action. In sports such as baseball or soccer, a sequence may be equivalent to a play, which is a form of occurrence. A sequence can include any number of possessions or can be a part of a possession. · Highlight show: A set of highlights arranged for presentation to a user. The highlight show may be presented linearly (such as in a video stream) or in a way that allows the user to select which highlights to view in which order (for example, by clicking on links or thumbnails). The presentation of the highlight show may be non-interactive or interactive, for example, allowing the user to pause, rewind, skip, fast forward, and / or convey preferences. The highlight show can be, for example, a condensed game. The highlight show can include any number of highlights, continuous or discontinuous, from a single event or from multiple events, and can further include highlights from different types of events (for example, different sports, and / or combinations of highlights of sports and non-sports events). · User / viewer: The terms "user" or "viewer" refer to an individual, group, or other entity that sees, hears, or otherwise experiences an event, one or more highlights of an event, or a highlight show in the same sense. The terms "user" or "viewer" can also refer to an individual, group, or other entity that will see, hear, or otherwise experience any of an event, one or more highlights of an event, or a highlight show at some future point in time. The term "viewer" may be used for illustrative purposes, but since an event does not necessarily need to include a visual component, "viewer" may instead be a listener or any other consumer of the content. · Story: A consistent story that links a set of highlight segments in a specific order. · Excitement level: A measure indicating how exciting or interesting an event or highlight is for a particular user or the general user. The excitement level can also be determined for a specific occurrence or player. Various techniques for measuring or evaluating the excitement level are described in the related applications referenced above. As described, the excitement level may vary depending on occurrences within the event and other factors such as the overall context or importance of the event (e.g., playoff games, pennant implications, and / or rivalries). In at least one embodiment, the excitement level can be associated with each occurrence, string, possession, or sequence within the event. For example, the excitement level of a possession can be determined based on the occurrences that take place within that possession. The excitement level may be measured differently by different users (e.g., fans of a certain team and neutral fans), and may vary depending on the individual characteristics of each user. · Metadata: Data that is related to and stored in association with other data. The primary data may be media such as a sports program or highlight. · Card image: An image within a video frame that provides data regarding any of the things depicted in the video, such as an event, a rendering of the event, or a portion thereof. Exemplary card images include game scores, game clocks, and / or other statistics from a sports event. The card image may appear temporarily or over the entire duration of the video stream, and those that appear temporarily may be related in particular to the portion of the video stream in which they appear. A "card image" may be a modified or processed version of the actual card image that appears within the video frame. · Character image: A portion of an image that appears to be related to a single character. The character image may include the area surrounding the character. For example, the character image may include a generally rectangular bounding box surrounding the character. · Text: Words, numbers, or symbols that can be part of a word or a number expression. The text can include letters, numbers, and special characters, and can be in any language. · String: A set of characters grouped in a way that indicates they are related to a single piece of information, such as the name of a team playing in a sports event. In many cases, English strings are arranged horizontally and read from left to right. However, strings may be arranged differently in English and other languages. · Video frame area: A portion of a video frame considered to contain a card image, based on either knowledge of a predetermined position within the video frame where the card image is expected to appear or sequential analysis of multiple regions of the video frame to identify which regions are likely to contain the card image.

[0034] Summary According to various embodiments, methods and systems are provided for automatically creating time-based metadata associated with highlights of a television program of a sports event. The highlights and related in-frame time-based information may be extracted synchronously with respect to the television broadcast of the sports event, or the video content of the sports event may be extracted while it is being streamed from a backup device through a video server after the television broadcast of the sports event.

[0035] In at least one embodiment, a software application operates in synchronization with the playback and / or reception of television program content to provide information metadata associated with highlights of the content. Such software can be executed, for example, on the television device itself, or on an associated STB, or on a video server having the function of receiving and then streaming program content, or on a mobile device having the function of receiving a video feed including live programs.

[0036] In a video management and processing system, and in the context of an interactive (enhanced) program guide, a set of video clips representing television broadcast content highlights can be automatically generated and / or stored in real time, along with a database containing time-based metadata that describes in more detail the events presented within the highlights. The metadata associated with the video clips can include any information, such as text information, images, and / or any type of audiovisual data. In this way, an interactive television application can provide timely and relevant content to a user viewing program content on either a primary television display or a secondary display such as a tablet, laptop, or smartphone.

[0037] One type of metadata associated with highlights of video content during and after a game is real-time information regarding sports game parameters directly extracted from live program content by reading information cards ("card images") embedded in one or more of the video frames of the program content. In various embodiments, the systems and methods described herein enable this type of automatic metadata generation.

[0038] In at least one embodiment, the system and method automatically detect and locate card images embedded in one or more decoded video frames of a television broadcast of a sports event program or in a sports event video streamed from a playback device. A number of predetermined regions of interest within the decoded video frames are analyzed, the card image quadrilaterals are located, and processed in real time using computer vision techniques to convert the information from the identified card images into a set of metadata that describes the status of the sports event.

[0039] In another embodiment, an automated process is described in which a digital video stream is received and one or more video frames of the digital video stream are analyzed for the presence of a card image quadrilateral. Next, a text box is located within the identified card image, and the text present within this text box is interpreted to create a metadata file that associates the card image content with a video highlight of the analyzed digital video stream.

[0040] In yet another embodiment, a plurality of text strings (text boxes) are identified, and the position and size of the image of each character within the string of characters associated with this text box are detected. Next, a plurality of text strings from various fields of the card image are processed and interpreted to form corresponding metadata and provide a plurality of information related to portions of a sports event associated with the processed card image and the analyzed video frame.

[0041] The automated metadata generation video system presented herein can operate in relation to a live broadcast video stream or a digital video streamed via a computer server. In at least one embodiment, the video stream is processed in real time using computer vision techniques to extract metadata from the embedded card image.

[0042] System Architecture According to various embodiments, the system can be implemented on any electronic device or set of electronic devices equipped to receive, store, and present information. Such electronic devices can be, for example, a desktop computer, laptop computer, television, smartphone, tablet, music player, voice device, kiosk, set-top box (STB), game system, wearable device, and / or home electronic device, among others.

[0043] The system is described herein in the context of an implementation on a particular type of computing device, but those skilled in the art will recognize that the techniques described herein can be implemented in other contexts and, in fact, can be implemented on any suitable device that can receive and / or process user input and present output to a user. Accordingly, the following description is intended to illustrate various embodiments by way of example and not to be limiting.

[0044] Referring now to FIG. 1A, a block diagram depicting the hardware architecture of a system 100 for automatically extracting metadata from card images embedded in a video stream of an event, according to a client / server embodiment, is shown. Event content, such as a video stream, can be provided via a network-connected content provider 124. Examples of such client / server embodiments are web-based implementations where each of one or more client devices 106 executes a browser or an app that provides a user interface for interacting with content from various servers 102, 114, 116, including a data provider(s) server 122 and / or a content provider(s) server 124, via a communication network 104. The transmission of content and / or data in response to requests from the client device 106 can be performed using any known protocol and language, such as Hypertext Markup Language (HTML), Java, Objective C, Python, and / or JavaScript.

[0045] The client device 106 can be, for example, a desktop computer, a laptop computer, a television, a smartphone, a tablet, a music player, a voice device, a kiosk, a set-top box, a game system, a wearable device, a home electronic device, and / or any electronic device. In at least one embodiment, the client device 106 has several hardware components known to those skilled in the art. The input device(s) 151 can be any component(s) that receive input from the user 150, for example, a handheld remote control, a keyboard, a mouse, a stylus, a touch-sensitive screen (touch screen), a touch pad, a gesture receptor, a trackball, an accelerometer, a five-way switch, or a microphone, etc. The input can be provided via any suitable mode including, for example, one or more of point, tap, type, drag, gesture, tilt, shake, and / or speech. The display screen 152 can be any component that graphically displays information, video, and / or content, etc., including renderings such as events and / or highlights. Such output can also include, for example, audiovisual content, data visualization, navigation elements, graphic elements, or a query that requests information and / or parameters for the selection of content. In at least one embodiment where only some of the desired output is presented at a time, dynamic controls such as a scroll mechanism can be utilized via the input device(s) 151 to select which information is currently being displayed and / or to change the way the information is displayed.

[0046] Processor 157 can be a conventional microprocessor for performing operations on data under the instruction of software according to well-known techniques. Memory 156 can be a random access memory having a structure and architecture known in the art for use by processor 157 in the process of executing software for performing the operations described herein. Client device 106 can also include local storage (not shown), such as a hard drive, flash drive, optical or magnetic storage device, and / or web-based (cloud-based) storage, etc.

[0047] Any suitable type of communication network 104, such as the Internet, a television network, a cable network, and / or a cellular network, can be used as a mechanism for transmitting data between client device 106 and various servers 102, 114, 116 and / or content providers 124 and / or data providers 122 according to any suitable protocol and technique. In addition to the Internet, other examples include cellular phone networks, EDGE, 3G, 4G, Long Term Evolution (LTE), Session Initiation Protocol (SIP), Short Message Peer-to-Peer Protocol (SMPP), SS7, Wi-Fi, Bluetooth®, ZigBee, Hypertext Transfer Protocol (HTTP), Secure Hypertext Transfer Protocol (SHTTP), and / or Transmission Control Protocol / Internet Protocol (TCP / IP), etc., and / or any combination thereof. In at least one embodiment, client device 106 transmits requests for data and / or content via communication network 104 and receives responses from servers 102, 114, 116 that include the requested data and / or content.

[0048] In at least one embodiment, the system of FIG. 1A operates in connection with a sports event. However, it should be understood that the teachings herein are also applicable to events other than sports, and the techniques described herein are not limited to application to sports events. For example, the techniques described herein can be utilized in connection with television shows, movies, news events, game shows, political events, business shows, dramas, and / or other episodic content, or for such two or more events.

[0049] In at least one embodiment, system 100 identifies highlights of a broadcast event by analyzing a video stream of the event. This analysis can be performed in real time. In at least one embodiment, system 100 includes one or more web server(s) 102 coupled to one or more client devices 106 via a communication network 104. Communication network 104 can be a public network, a private network, or a combination of a public network and a private network such as the Internet. Communication network 104 can be a LAN, a WAN, wired, wireless, and / or a combination of the above. In at least one embodiment, client device 106 can be connected to communication network 104 via either a wired or wireless connection. In at least one embodiment, the client device can also include a recording device capable of receiving and recording an event, such as a DVR, a PVR, or other media recording device. Such a recording device can be part of client device 106 or can be external. In other embodiments, such a recording device can be omitted. Although FIG. 1A shows one client device 106, system 100 can implement any number of client device(s) 106 of a single type or multiple types.

[0050] The web server(s) 102 can include one or more physical computing devices and / or software that receive requests from the client device(s) 106, respond to those requests with data, and can send uncommitted alerts and other messages. The web server(s) 102 may employ various strategies for fault tolerance and scalability such as load balancing, caching, and clustering. In at least one embodiment, the web server(s) 102 can include caching techniques as known in the art for storing information related to client requests and events.

[0051] The web server(s) 102 can maintain or otherwise specify one or more application server(s) 114 to respond to requests received from the client device(s) 106. In at least one embodiment, the application server(s) 114 provide access to business logic for use by client application programs within the client device(s) 106. The application server(s) 114 may be located in the same location as, shared with, or co-managed with the web server(s) 102. The application server(s) 114 may also be remote from the web server(s) 102. In at least one embodiment, the application server(s) 114 interact with one or more analytics server(s) 116 and one or more data server(s) 118 to perform one or more operations of the disclosed techniques.

[0052] One or more memory devices 153 can function as a "data store" by storing data related to the operation of system 100. This data may include, for example, but is not limited to, card data 154 related to card images embedded in a video stream presenting an event such as a sports event, user data 155 related to one or more users 150, and / or highlight data 164 related to one or more highlights of the event.

[0053] Card data 154 can include any information related to a card image embedded in a video stream, such as the card image itself, a subset thereof such as a character image, text extracted from the card image such as characters and character strings, and any of the foregoing attributes useful for text and / or meaning extraction. User data 155 can include any information describing one or more users 150, such as, for example, demographics, purchasing behavior, video stream viewing behavior, interests, and / or preferences. Highlight data 164 may include highlights, highlight identifiers, time stamps, categories, excitement levels, and other data related to the highlights. Card data 154, user data 155, and highlight data 164 will be described in detail hereinafter.

[0054] In particular, many components of system 100 may be or may include computing devices. Each such computing device may have an architecture similar to that of client device 106, as shown and described above. Thus, any of communication network 104, web server 102, application server 114, analytics server 116, data provider 122, content provider 124, data server 118, and memory device 153 may optionally include one or more computing devices having an input device 151, a display screen 152, a memory 156, and / or a processor 157, as described above in connection with client device 106.

[0055] In an exemplary operation of system 100, one or more users 150 of client device 106 display content from content provider 124 in the form of a video stream. The video stream may represent an event such as a sports event. The video stream may be a digital video stream that can be easily processed with known computer vision techniques.

[0056] Once the video stream is displayed, one or more components of system 100, such as client device 106, web server 102, application server 114, and / or analytics server 116, may analyze the video stream, identify highlights within the video stream, and / or extract metadata from the video stream, for example, from an embedded card image and / or other aspects of the video stream. This analysis can be performed in response to receiving a request to identify highlights and / or metadata of the video stream. Alternatively, in another embodiment, highlights can be identified without a specific request being made by user 150. In yet another embodiment, the analysis of the video stream can be performed without the video stream being displayed.

[0057] In at least one embodiment, user 150 can specify certain parameters for the analysis of a video stream (e.g., which events / games / teams to include, how much time user 150 has available for viewing highlights, what metadata is desired, and / or any other parameters, etc.) via input device(s) 151 of client device 106. User preferences can also be extracted from storage, such as user data 155 stored in one or more storage devices 153, to customize the analysis of the video stream without necessarily requiring user 150 to specify the preferences. In at least one embodiment, user preferences can be determined based on the observed actions and activities of user 150, for example, by observing website visit patterns, TV viewing patterns, music listening patterns, online purchases, prior highlight identification parameters, and / or highlights and / or metadata actually viewed by user 150.

[0058] Additionally or alternatively, user preferences can be retrieved from pre-stored preferences explicitly provided by user 150. Such user preferences can indicate which teams, sports, players, and / or types of events are of interest to user 150, and / or they can indicate which types of metadata or other information related to highlights would be of interest to user 150. Thus, such preferences can be used to guide the analysis of the video stream, identify highlights, and / or extract metadata for the highlights.

[0059] Analysis server(s) 116, which may include one or more of the above computing devices, can analyze live and / or recorded feeds of in-play statistics related to one or more events from data provider(s) 122. Examples of data provider(s) 122 include, but are not limited to, STATSTM, Perform (available from Opta Sports, London, UK), and SportRadar, Sankt Gallen, Switzerland, providers of real-time sports information. In at least one embodiment, analysis server(s) 116 generates a set of different excitement levels for an event. Such excitement levels can then be stored in association with the highlights identified by system 100 in accordance with the techniques described herein.

[0060] Application server(s) 114 can analyze video streams to identify highlights and / or extract metadata. Additionally or alternatively, such analysis may be performed by client device(s) 106. The identified highlights and / or extracted metadata may be specific to user 150, and in such cases, it may be advantageous to identify highlights within client device 106 associated with a particular user 150. Client device 106 may receive, hold, and / or obtain applicable user preferences for highlight identification and / or metadata extraction as described above. Additionally or alternatively, highlight generation and / or metadata extraction may be performed globally (i.e., using objective criteria applicable generally to a population of users regardless of the preferences of a particular user 150). In such cases, it may be advantageous to identify highlights and / or extract metadata within application server(s) 114.

[0061] Content that facilitates highlight identification and / or metadata extraction may come from any suitable source, including content providers (plural) 124 such as websites like YouTube (registered trademark), and MLB.com, sports data providers, television stations, and / or client- or server-based DVRs. Alternatively, the content may come from a local source such as a DVR or other recording device associated with (or incorporated into) the client device 106. In at least one embodiment, the application server (plural) 114 generates a customized highlight show with highlights and metadata available to the user 150, either by downloading, or streaming content, or on-demand content, or any other means.

[0062] As described above, it may be advantageous to perform user-specific highlight identification and / or metadata extraction on a particular client device 106 associated with a particular user 150. Such an embodiment can avoid the need for video content or other high-bandwidth content to be unnecessarily transmitted over the communication network 104, especially if such content is already available on the client device 106.

[0063] For example, referring next to FIG. 1B, an example of a system 160 according to one embodiment is shown in which at least some of the card data 154 and the highlight data 164 are stored in a client-based storage device 158, and the storage device 158 may be any form of local storage device available to the client device 106. By way of example, for example, a DVR capable of recording an event such as video content of a complete sports event can be cited. Alternatively, the client-based storage device 158 can be any magnetic, optical, or electronic storage device for digital-form data. Examples include flash memory, a magnetic hard drive, a CD-ROM, a DVD-ROM, or other devices integrated with or communicably coupled to the client device 106. Based on the information provided by the application server(s) 114, the client device 106 may extract metadata from the card data 154 stored in the client-based storage device 158 without the need to retrieve other content from the content provider 124 or other remote sources, and store the metadata as the highlight data 164. Such a configuration can save bandwidth and effectively utilize existing hardware that may already be available to the client device 106.

[0064] Returning to FIG. 1A, in at least one embodiment, the application server(s) 114 can identify different highlights and / or extract different metadata for different users 150 according to individual user preferences and / or other parameters. The identified highlights and / or the extracted metadata may be presented to the user 150 via any suitable output device, such as the display screen 152 of the client device 106. Optionally, multiple highlights can be identified and assembled into a highlight show along with the associated metadata. Such a highlight show can be accessed via a menu and / or assembled into a "highlight reel" or a set of highlights that are played for the user 150 according to a predetermined sequence. In at least one embodiment, the user 150 can control the playback and / or distribution of the highlighted metadata via the input device(s) 151, for example, for the following purposes. · Select specific highlights and / or metadata for display. · Pause, rewind, and fast forward. · Skip to the next highlight. · Return to the beginning of the previous highlight within the highlight show. And / or · Perform other actions.

[0065] Additional details regarding such functionality are provided in the related U.S. patent application cited above.

[0066] In at least one embodiment, another data server(s) 118 is provided. The data server(s) 118 may respond to requests for data from any of the servers 102, 114, 116, for example, to obtain or provide card data 154, user data 155, and / or highlight data 164. In at least one embodiment, such information can be stored in any suitable storage device 153 accessible by the data server 118 and can come from any suitable source such as the client device 106 itself, the content provider(s) 124, and / or the data provider(s) 122.

[0067] Referring now to FIG. 1C, a system 180 is shown according to an alternative embodiment in which the system 180 is implemented in a stand-alone environment. Similar to the embodiment shown in FIG. 1B, at least some of the card data 154, user data 155, and highlight data 164 may be stored in a client-based storage device 158 such as a DVR. Alternatively, the client-based storage device 158 can be a flash memory or a hard drive, or other device integrated with or communicatively coupled to the client device 106.

[0068] The user data 155 may include the preferences and interests of the user 150. Based on such user data 155, the system 180 can extract metadata within the card data 154 and present it to the user 150 in the manner described herein. Additionally or alternatively, the metadata can be extracted based on objective criteria not based on information specific to the user 150.

[0069] Referring now to FIG. 1D, an overview of a system 190 having an architecture according to an alternative embodiment is shown. In FIG. 1D, the system 190 includes a broadcast service such as a content provider(s) 124, a content receiver in the form of a client device 106 such as a television set having an STB, a video server such as an analysis server(s) 116 that can capture and stream television program content, and / or other client devices 106 such as mobile devices and laptops that can receive and process television program content, all of which are connected via a network such as a communication network 104. A client-based storage device 158 such as a DVR can be connected to any of the client devices 106 and / or other components, and can store video streams, highlights, highlight identifiers, and / or metadata to facilitate the identification and presentation of highlights and / or extracted metadata via any of the client devices 106.

[0070] The particular hardware architectures depicted in FIGS. 1A, 1B, 1C, and 1D are merely exemplary. Those skilled in the art will recognize that the techniques described herein can be implemented using other architectures. Many of the components depicted herein are optional and may be omitted, integrated with other components, and / or replaced with other components.

[0071] In at least one embodiment, the system can be implemented as software written in any suitable computer programming language, whether a stand-alone or client / server architecture. Alternatively, it may be implemented in and / or embedded in hardware.

[0072] Data Structures FIG. 2 is a schematic block diagram depicting an example of a data structure that can be incorporated into card data 154, user data 155, and highlight data 164, according to one embodiment.

[0073] As shown, the card data 154 may include records of each of the plurality of broadcast networks 202. For example, for each of the distribution networks 202, the card data 154 may include a predetermined card position 203 where the distribution network displays the card image within a normal video frame. The predetermined card position can be represented, for example, as coordinates (such as Cartesian coordinates) that identify opposite corners of the position, identify the center, height, and width, and / or identify the position and / or the size of the card image.

[0074] Furthermore, the card data 154 may include one or more video frame regions 204 that have been analyzed or are to be analyzed for card image extraction and interpretation. Each video frame region 204 may be extracted from a video frame of the video stream.

[0075] For each video frame region 204, the card data 154 can also include one or more processed video frame regions 206, which may be generated by modifying the video frame region 204 in a way that facilitates the identification and / or extraction of the card image 207. For example, the processed video frame region 206 may include a version of each video frame region 204 that has been modified by one or more of trimming, recoloring, segmenting, expanding, or other methods.

[0076] Each video frame region 204 can also have a card image 207 identified within and / or extracted from the video frame region 204. Each card image 207 may include text that can be interpreted to provide metadata related to a specific time within the video stream.

[0077] Card data 154 can also include one or more interpretations 208 for each video frame region 204. Each interpretation 208 can be a specific text that is considered to be represented in the associated card image 207 after some analysis has been performed to recognize and interpret the characters appearing in the card image 207. The interpretation 208 may be used to obtain metadata from the card image 207.

[0078] As further shown, the user data 155 may include records related to the user 150, and each record may include demographic data 212, preferences 214, viewing history 216, and purchase history 218 of a specific user 150.

[0079] The demographic data 212 may include any type of demographic data, including but not limited to age, gender, location, nationality, religious affiliation, and / or education level.

[0080] The preferences 214 may include selections made by the user 150 regarding their preferences. The preferences 214 may be directly related to the collection and / or display of highlights and metadata, or may be of a more general nature. In either case, the preferences 214 can be used to facilitate the identification and / or presentation of highlights and metadata to the user 150.

[0081] The viewing history 216 can list television programs, video streams, highlights, web pages, search queries, sports events, and / or other content retrieved and / or viewed by the user 150.

[0082] The purchase history 218 can list products or services purchased or requested by the user 150.

[0083] As further shown, the highlight data 164 may include recordings of the j highlights 220, each of which may include a video stream 222, an identifier, and / or metadata 224 of a particular highlight 220.

[0084] The video stream 222 may include video depicting the highlight 220, which may be obtained from one or more video streams of one or more events (e.g., by trimming the video stream to include only the video stream 222 associated with the highlight 220). The identifier 223 may include a time code and / or other indicator indicating where in the video stream of the event from which the highlight 220 was obtained the highlight 220 is located.

[0085] In some embodiments, each recording of a highlight 220 may include only one of the video stream 222 and the identifier 223. Highlight playback may be performed by playing the user's video stream 222 or by playing only the highlighted portion of the video stream of the event from which the highlight 220 was obtained using the identifier 223.

[0086] The metadata 224 may include information about the highlight 220, such as the date of the event, the season, and the group or individual involved in the event or video stream from which the highlight 220 was obtained, such as a team, player, coach, anchor, broadcaster, and / or fan. Among other information, the metadata 224 for each highlight 220 may include a time 225, a phase 226, a clock 227, a score 228, and / or a frame number 229.

[0087] The time 225 may be the time within the video stream 222 at which the highlight 220 is acquired, or the time within the video stream 222 related to the highlight 220 for which metadata is available. In some examples, the time 225 may be the playback time within the video stream 222 related to the highlight 220 at which the card image 207 including the metadata 224 is displayed.

[0088] The phase 226 may be the phase of the event related to the highlight 220. More specifically, the phase 226 may be the stage of the sports event at which the card image 207 including the metadata 224 is displayed. For example, the phase 226 may be "the third quarter", "the second inning", or "the bottom half", etc.

[0089] The clock 227 may be the game clock related to the highlight 220. More specifically, the clock 227 may be the state of the game clock when the time card image 207 including the metadata 224 is displayed. For example, if the game clock shows 15 minutes and 47 seconds and the card image 207 is displayed, the clock 227 may be "15:47".

[0090] The score 228 may be the game score related to the highlight 220. More specifically, the score 228 may be the score when the card image 207 including the metadata 224 is displayed. For example, the score 228 may be "45-38", "7-0", or "30-love", etc.

[0091] The frame number 229 may be the number of the video frame within the video stream at which the highlight 220 is acquired, or the number of the video frame most directly related to the highlight 220 within the video stream 222 related to the highlight 220. More specifically, the frame number 229 may be the number of such a video frame at which the card image 207 including the metadata 224 is displayed.

[0092] The data structure described in FIG. 2 is merely an example. Those skilled in the art will recognize that in the implementation of highlight identification and / or metadata extraction, some of the data in FIG. 2 can be omitted or replaced with other data. Additionally or alternatively, data not shown in FIG. 2 can be used in the implementation of highlight identification and / or metadata extraction.

[0093] Card image Referring now to FIG. 3A, a screen shot illustration of an example of a video frame 300 from a video stream in which information is embedded in the form of a card image 207 is shown, as frequently appears in a television program of a sports event. FIG. 3A depicts a card image 207 in the lower right corner of the video frame 300, and a second card image 320 extending along the bottom of the video frame 300. The card images 207, 320 may include embedded information such as a game phase, a current clock, and a current score.

[0094] In at least one embodiment, the information within the card images 207, 320 is located and processed for automatic recognition and interpretation of the embedded text within the card images 207, 320. The interpreted text may then be assembled into text metadata that describes the status of the sports game at a particular point in time within the timeline of the sports event.

[0095] In particular, the card image 207 may be related to the currently shown sports event, while the second card image 320 may include information regarding a different sports event. In some embodiments, only the card images that contain information considered relevant to the currently playing sports event are processed for metadata generation. Thus, without limitation, the following exemplary description assumes that only the card image 207 is processed. However, in alternative embodiments, it may be desirable to process multiple card images within a given video frame 300, including card images related to other sports events.

[0096] As shown in FIG. 3A, the card image 207 can provide several different types of metadata 224, including team name 330, score 340, previous team performance 350, current game stage 360, game clock 370, place status 380, and / or other information 390. Each of these can be extracted from within the card image 207 and may be interpreted to provide metadata 224 corresponding to the highlight 220 including the video frame 300, more specifically, the video frame 300 in which the card image 207 is displayed.

[0097] FIG. 3B is a series of screen shot diagrams depicting additional examples of video frames 392, 394, 396, 398, each having an embedded card image 393, 395, 397, 399, respectively, to show additional examples of the position of the embedded card image in a sports television program. Different television networks may have different types, shapes, and frame positions of such card images embedded in the video frames of the television program content of the sports event.

[0098] Position Identification and Extraction of Card Images FIG. 4 is a flowchart depicting a method 400 executed by an application (e.g., executed on one of the client device 106 and / or the analysis server 116) according to one embodiment, the application receiving a video stream 222 and performing on-the-fly processing of the video frame 300 for identifying and extracting the card image 207 and related metadata such as the card image 207 and its related game status information of FIG. 3. The system 100 of FIG. 1A is referred to as a system implementing the method 400 and subsequent methods. However, alternative systems including, but not limited to, the system 160 of FIG. 1B, the system 180 of FIG. 1C, and / or the system 190 of FIG. 1D can be used in place of the system 100 of FIG. 1A.

[0099] Method 400 of FIG. 4 may include receiving video stream 222. In step 410, for example, one or more video frames 300 of video stream 222 may be read and decoded by resizing video frame 300 to a standard size. In query 420, step 430, step 440, and / or query 450, video frame 300 may be processed for identifying the position of the in-frame card image. In step 460, the detected card image 207 may be processed to extract information by reading and interpreting card image 207. Metadata 224 may be generated based on the information extracted from card image 207.

[0100] In at least one embodiment, the detection of one or more card images 207 present in the decoded video frame 300 is performed by analyzing a single predetermined frame area. Alternatively, such detection can be performed by analyzing multiple predetermined frame areas if the approximate position of the card image 207 within the decoded video frame 300 is not known in advance. Thus, query 420 can determine whether the position of the card image 207 within video frame 300 is known. For example, some broadcast networks may always show the card image 207 at the same position within video frame 300. If the broadcast network is known, the position of the card image 207 may also be known. Alternatively, the position of the card image 207 within video frame 300 may not be known and may need to be confirmed by system 100.

[0101] If, in accordance with query 420, the position of the card image 207 within the video frame 300 is known, method 400 may proceed to step 430, which can process the known portion or video frame region to isolate the quadrilateral shape normally associated with the card image 207. If the position of the card image 207 within the video frame 300 is not known, method 400 proceeds to step 440, where the video frame 300 is divided into a plurality of regions, which may be predetermined regions of the video frame 300. The regions of the video frame 300 are sequentially analyzed to determine which regions contain card images similar to the card image 207, 395, 397, and / or 399.

[0102] For example, the particular region(s) of the video frame 300 that contain the card image 207 may be known for each of the various broadcast networks. If the broadcast network is not known, system 100 may sequentially proceed through each region of the video frame 300 known to be used by the broadcast network for display of the card image 207 until one of the regions is found that contains the card image 207.

[0103] If, in accordance with query 450, the card image 207 is found, method 400 can proceed to step 460, where the card image 207 is processed and information is extracted from the card image 207 to provide the metadata 224. If, in accordance with query 450, the card image 207 is not found, method 400 can return to step 410, where a new video frame may be loaded, decoded, and then analyzed for the presence of the card image 207.

[0104] As described above, method 400 may be executed in real time in some embodiments while user 150 is viewing a program (e.g., while video stream 222 corresponding to highlight 220 is being presented). Thus, method 400 may be executed in the background for each video frame 300 while the video frame 300 is being decoded for playback for user 150. Since system 100 locates, extracts, and interprets card image 207, there may be some delay. Thus, in this application, the presentation of metadata extracted from card image 207 is considered to be "real time" even if the presentation of metadata 224 lags behind the playback of video frame 300 in which the metadata was obtained (e.g., by several video frames of video frame 300 such that it is not perceived by user 150 or causes no distraction to user 150).

[0105] FIG. 5 is a flowchart depicting in more detail step 440 for processing a predetermined region of video frame 300 to detect an executable card image 207 from FIG. 4. Each of the predetermined regions may represent an approximate location where card image 207 may be present.

[0106] In at least one embodiment, a predetermined region within decoded video frame 300 is generated based on knowledge of the approximate location of card image 207 used by various television networks engaged in the broadcast of sports event television programs, as described above. Such television networks are known to use one or more regions of video frame 300 to distribute visual and text data within the frame via card image 207.

[0107] In step 510, sequential processing of regions can be started. In step 520, one of the regions may be processed to check whether a valid card image 207 exists in that region. Query 530 can determine whether the card image 207 has been found in that region. If found, the region may be further processed to extract the card image 207. The located card image 207 may be further processed for automatic recognition and interpretation of the embedded text. Next, such interpreted text may be further assembled into text metadata that describes the status of a sports event (such as a game) at a specific point in time on the sports event timeline. In at least one embodiment, the options available for text rendering are based on the type of card image 207 detected within the video frame 300, and this type may be determined by the system 100 during the localization and / or extraction of the card image 207. Additionally or alternatively, the options available for text rendering may be based on the pre-assigned meaning of the selected fields present within the detected specific type of card image 207.

[0108] If the card image 207 is not found within the region, query 550 may check whether the region is the last region of the video frame 300. If it is not the last region, the system 100 may proceed to the next region in step 560 and then repeat the processing for the next region according to step 520. If the region is the last region of the video frame 300, the video frame 300 may not contain a valid card image 207, and the system 100 can proceed to the next video frame 300.

[0109] Automatic Detection and Localization of the Card Image Quadrilateral FIG. 6 is a flowchart depicting a method 600 for top-level processing for detecting a valid card image quadrilateral in a designated area of a decoded video frame 300 according to one embodiment. Method 600 may be performed on a video frame region at a predetermined position or on a video frame region identified through sequential processing of multiple regions of video frame 300 according to step 430 of FIG. 4 or step 440 of FIGS. 4 and 5.

[0110] First, at step 610, the decoded video frame 300 may be trimmed to a smaller area that includes the designated video frame region, and a trimmed image is provided. At step 620, the trimmed image may be segmented using any suitable segmentation algorithm such as graph-based segmentation (e.g., “Efficient Graph-Based Image Segmentation,” P. Felzenszwalb, D. Huttenlocher, Int. Journal of Computer Vision, 2004, Vol. 59), all the generated segments may be color-coded and enumerated, and a segmented image is provided. Further processing of the segmented image may include removing background material surrounding the assumed quadrilateral that defines the card image 207. In at least one embodiment, method 600 may proceed to step 630, where all pixels of segments adjacent to the boundary of the segmented trimmed image are set to the black level. At step 640, all pixels of the remaining inner segments of the segmented trimmed image are set to the white level. At step 650, the two-color trimmed image with partially removed background may be passed for further processing for accurate card image quadrilateral drawing.

[0111] FIG. 7 is a flowchart depicting a method 700 for more accurate determination of the quadrilateral of a card image according to one embodiment. First, at step 710, a trimmed image with the background partially removed (e.g., generated at step 640 of FIG. 6) is converted into a grayscale image. The grayscale image may then be blurred, and at step 720, an edge detection process may be performed to generate an edge image having the detected edges. Next, at step 730, the edge image may be processed for contour detection, and the resulting contour image may be further processed to approximate the contour having a closed polygon. Thereafter, at step 740, the contour / polygon image may be processed to determine the smallest rectangular perimeter surrounding all existing contours. The above steps can generate a rectangular enclosure potentially containing the card image 207. However, this enclosure may be larger than the card image quadrilateral due to artifacts generated during the process of segmenting the trimmed image. Thus, in at least one embodiment, further adjustment is performed to push this intermediate rectangular shape into the smallest rectangular area containing the card image 207.

[0112] FIG. 8 is a flowchart depicting an exemplary method 800 for adjusting the quadrilateral boundary of an enclosure encompassing all detected contours (e.g., generated by step 740 of FIG. 7) according to one embodiment. The enclosure can surround the internal area such that most of the surrounding pixels of the pushed-in new enclosure have the same pixel intensity (such as white in this particular example). Method 800 can remove unwanted internal area artifacts extending outward, and thus provide a new, more robust enclosure that may contain the valid card image 207.

[0113] The method of FIG. 8 can start from step 810 where the perimeter is received. The perimeter may be a rectangular image. In step 820, system 100 can "walk" around the boundary of a rectangular bounding image that encompasses all detected perimeters (or quadrilateral images), and in step 830, count the pixels that have a black level value for each boundary edge, i.e., top, bottom, left, and right. Next, in step 840, if any boundary edge contains more black-value pixels than a predetermined count, that edge is moved inward by one pixel, providing an adjusted quadrilateral area 850. This process continues until query 860 determines that the black-value pixel counts for all edges of the pushed-in quadrilateral are below a predetermined threshold. The resulting pushed-in rectangle represents the enclosure of a potential card image and may be verified in the processing steps described in relation to the flowchart of FIG. 9.

[0114] FIG. 9 is a flowchart depicting an exemplary method 900 for quadrilateral verification of a card image according to one embodiment. Method 900 may include the analysis of the following three different image regions (pixel counts): a trimmed image region (e.g., generated in step 610 of FIG. 6), a rectangular bounding image region that encompasses all detected perimeters (e.g., generated in step 740 of FIG. 7), and an image region having an adjusted (pushed-in) quadrilateral boundary (e.g., generated in one or more iterations of step 840 of FIG. 8). Three parameters (A, B, C) can be generated in steps 910, 920, and 930, respectively, as follows. · A = Total pixel count of the trimmed image area. · B = Total pixel count of the perimeter binary image. · C = Black-value pixel count of the adjusted perimeter binary image.

[0115] Next, in step 940, when a valid card image quadrilateral is to be detected, a weighted comparison of these three parameters can be performed such that the non-black value pixel area of the pushed-in quadrilateral exists at a specific ratio with respect to the other two parameters. In step 950, based on the above weighted comparison, if a valid card image 207 is detected, the flag is set to true. According to query 960, if the flag is set to true, the system 100 can proceed to step 970, where the card image 207 (and / or a processed version of the card image 207) is passed to the card image internal content processing. In step 950, if a valid card image 207 is not detected, the flag is set to false, and according to query 960, the system 100 can proceed to the next specified frame area or the next video frame 300 in step 980 to search for a valid card image 207 inside.

[0116] Figure 10 is a flowchart depicting an exemplary method 1000 for any stabilization of the left (or any other) boundary of a very elongated card image shape according to one embodiment. The process can start in step 1010 where the system 100 extends the horizontal card image 207 to the trimmed frame edge. In step 1020, the system 100 may detect the vertical lines of the straight lines in this extended image. This process may further include, in step 1030, selecting the detected vertical lines of a predetermined length and calculating the selected sparse vertical line markers. Finally, in step 1040, the marker (if any) to the left of the original position of the card image 207 that is found nearby is selected, and in step 1050, the left edge of the quadrilateral is moved to the position of the marker further to the left of the original edge position of the detected card image quadrilateral. In step 1060, the card image quadrilateral is adjusted accordingly, and the updated region of interest (ROI) of the card is returned.

[0117] Internal Processing of Card Image for Information Extraction In at least one embodiment, an automated process is performed that includes receiving a digital video stream (which may include one or more highlights of a broadcast sports event), analyzing one or more video frames of the digital video stream for the presence of a card image 207, extracting the card image 207, identifying text boxes within the card image 207, and interpreting text present within the text boxes to create metadata 224 that associates the content from the card image 207 with video highlights of the analyzed digital video stream.

[0118] FIG. 11 is a flowchart depicting a method 1100 for performing text extraction from a card image 207, according to one embodiment. In step 1110, the extracted card image 207 may be resized to a standard size. Next, in step 1120, the resized card image 207 may be preprocessed using a series of filters including, for example, contrast enhancement, bilateral and median filtering for noise reduction, and gamma correction followed by illumination compensation. In at least one embodiment, in step 1130, an "extreme region filter" with a two-stage classifier is created (e.g., L. Neumann, J. Matas, "Real-Time Scene Text Localization and Recognition", 5th IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, June 2012), and in step 1140, a cascade classifier is applied to each image channel of the card image 207. Next, in step 1150, character groups are detected and a group of word boxes is extracted.

[0119] In at least one embodiment, a plurality of text strings (text boxes) are identified within the card image 207, and the position and size of each character within the string of characters associated with this text box are detected. Next, the text strings from various fields of the card image 207 are processed and interpreted, and corresponding metadata 224 is generated, thus providing real-time information related to the current sports event television program and the current timeline associated with the processed embedded card image 207.

[0120] FIG. 12 is a flowchart depicting a method 1200 for performing processing and interpretation of text strings, according to one embodiment. In step 1210, the detected and extracted card image 207 may be processed, and the text to be interpreted may be selected from a group of character bounding boxes within the card image 207. Next, in step 1220, text may be extracted, and the extracted text may be read and interpreted, for example, via optical character recognition (e.g., “An Overview of the Tesseract OCR Engine”, R. Smith, Proceedings ICDAR’07, Vol. 02, Sept. 2007.). In step 1230, metadata 224 is generated and may be structured. Next, the in-frame information from the card image 207 is combined with video highlight text and visual metadata.

[0121] The system and method have been described in particular detail with respect to the assumed embodiments. Those skilled in the art will understand that the system and method may be implemented in other embodiments. First, the specific naming of components, the use of capital letters in terms, attributes, data structures, or any other programming or structural aspects are neither essential nor important, and the mechanisms and / or functions may have different names, formats, and protocols. Further, the system may be implemented via a combination of hardware and software, or entirely within a hardware element, or entirely within a software element. Also, the specific division of functions among the various system components described herein is merely exemplary and not essential. The functions performed by a single system component may instead be performed by multiple components, and the functions performed by multiple components may instead be performed by a single component.

[0122] References to "one embodiment" or "an embodiment" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrases "in one embodiment" or "in at least one embodiment" in various places in this specification are not necessarily all referring to the same embodiment.

[0123] The various embodiments may include any number of systems and / or methods for implementing the above-described techniques, either alone or in any combination. Another embodiment includes a non-transitory computer-readable storage medium for causing a processor in a computing device or other electronic device to implement the above-described techniques, and a computer program product including computer program code encoded on the medium.

[0124] Some of the foregoing has been presented from the perspective of algorithms and symbolic representations of operations on data bits within the memory of a computing device. These descriptions and representations of algorithms are the means used by those skilled in the data processing arts to most effectively convey the essence of their work to other such skilled artisans. An algorithm is here generally considered to be a self-consistent series of steps (instructions) leading to a desired result. The steps are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical, magnetic, or optical signals capable of being manipulated in memory, transferred, combined, compared, and otherwise. For mainly reasons of common usage, it may be convenient to refer to these signals as bits, values, elements, symbols, characters, terms, or numerical values, etc. Further, without loss of generality, it may be convenient to refer to a particular arrangement of steps requiring physical manipulation of physical quantities as a module or a coded device.

[0125] However, it should be borne in mind that all of these and similar terms are merely associated with appropriate physical quantities and are nothing more than convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the following description, throughout this specification, descriptions using terms such as "processing" or "computing" or "calculating" or "displaying" or "determining" refer to the operation and processes of a computer system, or similar electronic computing modules and / or devices, and mean the manipulation and transformation of data represented as physical (electronic) quantities within the memory or registers or other such storage, transmission device, or display device of the computer system.

[0126] Certain aspects include the process steps and instructions described herein in the form of an algorithm. The process steps and instructions can be embodied in software, firmware, and / or hardware, and when embodied in software, can be downloaded to and exist on various platforms used by various operating systems, and it should be noted that they can be operated from various platforms.

[0127] This document also relates to an apparatus for performing the operations herein. This apparatus can be specially constructed for the required purposes or can include a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computing device. Such a computer program may be stored in a computer-readable storage medium such as a floppy disk, optical disk, CD-ROM, DVD-ROM, magneto-optical disk, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, flash memory, solid state drive, magnetic card or optical card, application specific integrated circuit (ASIC), or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus. The program and its associated data may also be hosted and executed remotely, for example, on a server. Furthermore, the computing devices referred to herein can include a single processor or can be an architecture that employs multiple processor designs to enhance computing power.

[0128] The algorithms and displays presented in this specification are not inherently related to any particular computing device, virtualization system, or other apparatus. Various general-purpose systems may also be used with programs in accordance with the teachings of this specification, or it may prove convenient to construct more specialized apparatus for performing the required method steps. The required structure for these various systems will become apparent from the description provided herein. Further, the systems and methods are not described with reference to any particular programming language. It will be understood that various programming languages may be used to implement the teachings described herein, and any such reference to a particular language is provided for the purpose of enabling and best mode disclosure.

[0129] Accordingly, various embodiments include software, hardware, and / or other elements for controlling a computer system, computing device, or other electronic device, or any combination or plurality of these. Such electronic devices may include, for example, input devices such as a processor, keyboard, mouse, touchpad, trackpad, joystick, trackball, microphone, and / or any combination thereof, output devices such as a screen and / or speaker, long-term storage devices such as memory, magnetic storage, and / or optical storage, and / or network connectivity, using techniques well known in the art. Such electronic devices may be portable or non-portable. Examples of electronic devices that can be used to implement the described systems and methods include desktop computers, laptop computers, televisions, smartphones, tablets, music players, voice devices, kiosks, set-top boxes, game systems, wearable devices, home electronics, and / or server computers. The electronic devices can use any operating system, such as, for example, Linux (registered trademark), Microsoft Windows available from Microsoft Corporation, Redmond, Washington, Mac OS X available from Apple Inc., Cupertino, California, iOS available from Apple Inc., Cupertino, California, Android available from Google Inc., Mountain View, California, and / or any other operating system adapted for use on the device, but are not limited thereto.

[0130] Although a limited number of embodiments have been described herein, those of ordinary skill in the art having the benefit of the above description will appreciate that other embodiments may be devised. Further, note that the language used herein has been principally selected for readability and for the purpose of education and may not have been selected to delineate or circumscribe the subject matter. Accordingly, this disclosure is intended to be illustrative, but not limiting, of the scope.

Claims

1. A method for extracting a card image from a video frame, comprising: a video frame area selection step of selecting a video frame area that forms a part of the video frame; a pixel value correction step of correcting pixel values of a set of pixels adjacent to a boundary of the video frame area; a step of removing a background from the video frame area based on the corrected pixel values; a step of generating an edge image based on the video frame area from which the background has been removed; a step of identifying a contour in the edge image; a card image extraction step of extracting the card image as an area surrounded by an area including the contour.

2. The method according to claim 1, wherein the video frame area selection step includes trimming the video frame to separate the video frame area.

3. The video frame area selection step includes: determining whether the position of the card image is known; selecting, in response to determining that the position of the card image is known, a part of the video frame at the known position as the video frame area; identifying the video frame area by sequentially processing a plurality of areas of the video frame in response to determining that the position of the card image is not known. The method according to claim 1.

4. The method according to claim 1, wherein the pixel value correction step includes setting the pixel values to a black level.

5. further comprising a step of approximating the identified contour as a polygon, wherein the card image extraction step includes extracting the card image as an area surrounded by an area as a minimum rectangular perimeter including all of the polygon. The method according to claim 1.

6. counting the corrected pixels for one or more boundaries of the area; moving the one or more boundaries inward to generate an adjusted area in response to determining that the number of the corrected pixels exceeds a threshold value. The method according to claim 5.

7. further comprising a verification step of verifying the adjusted area, wherein in the verification step: counting a first number of pixels in the video frame area; counting a second number of pixels in the area. Counting the number of third pixels in the adjusted area; Verifying is performed by: determining that the adjusted area is verified based on comparing the number of first pixels, the number of second pixels, and the number of third pixels. The method according to claim 6.

8. The method according to claim 1, further comprising the step of generating metadata associated with the video frame based on analyzing the extracted card image.

9. The video frame is within a highlight of a video stream of a sports event broadcast, The metadata describes the status of the sports event during the highlight. The method according to claim 8.

10. The method according to claim 9, further comprising the step of presenting the metadata during viewing of the highlight on an output device.

11. A system for extracting a card image from a video frame, comprising: A processor; A non-transitory storage medium storing computer program instructions, wherein when the processor executes the computer program instructions, the system: A video frame area selection step of selecting a video frame area forming a part of the video frame; A pixel value correction step of correcting pixel values of a set of pixels adjacent to the boundary of the video frame area; A step of removing a background from the video frame area based on the corrected pixel values; A step of generating an edge image based on the video frame area from which the background has been removed; A step of identifying a contour in the edge image; A card image extraction step of extracting the card image as an area surrounded by an area including the contour. A system that performs operations including.

12. The operation of the video frame area selection step includes separating the video frame area by trimming the video frame. The system according to claim 11.

13. The operation of the video frame area selection step includes: Determining whether the position of the card image is known; In response to determining that the position of the card image is known, selecting a part of the video frame at the known position as the video frame area; In response to determining that the position of the card image is not known, sequentially processing a plurality of regions of the video frame to identify the video frame region, the system according to claim 11, comprising.

14. The operation of the pixel value correction step includes setting the pixel value to a black level, the system according to claim 11.

15. The operation further includes a step of approximating the identified contour as a polygon, The operation of the card image extraction step includes extracting the card image as a region surrounded by the region as the smallest rectangular perimeter including all of the polygon, the system according to claim 11.

16. The operation is Counting the corrected pixels for one or more boundaries of the region; In response to determining that the number of the corrected pixels exceeds a threshold value, further including a step of moving the one or more boundaries inward to generate an adjusted region, the system according to claim 15.

17. The operation further includes a verification step of verifying the adjusted region, and in the verification step, Counting a first number of pixels in the video frame region; Counting a second number of pixels in the region; Counting a third number of pixels in the adjusted region; Verification is performed by determining that the adjusted region is verified based on comparing the first number of pixels, the second number of pixels, and the third number of pixels, the system according to claim 16.

18. The operation is Based on analyzing the extracted card image, further comprising a step of generating metadata associated with the video frame, the system according to claim 11.

19. The video frame is within a highlight of a video stream of a sports event broadcast, The metadata describes the status of the sports event during the highlight, the system according to claim 18.

20. A non-transitory storage medium storing computer program instructions, the computer program instructions A video frame region selection step of selecting a video frame region forming a part of the video frame, A pixel value correction step of correcting pixel values of a set of pixels adjacent to a boundary of the video frame area; A step of removing a background from the video frame area based on the corrected pixel values; A step of generating an edge image based on the video frame area from which the background has been removed; A step of identifying a contour in the edge image; A non-transitory storage medium executed by a processor to perform operations including a card image extraction step of extracting a card image from the video frame as an area surrounded by an area including the contour.

Citation Information

Patent Citations

  • Video attribute information output apparatus, video summarizing device, program, and method for outputting video attribute information

    JP2008176538A

  • Image reproducing device

    JP2011009816A

  • How to adapt video images to smaller screen sizes.

    JP2011514789A

  • Extraction program, method, device, and baseball video meta information generating device, method and program

    JP2015139016A

  • Determination program, method, and device

    JP2016048852A