Machine learning for recognizing and interpreting embedded information card content

A machine learning and computer vision-based system processes live sporting event broadcasts to extract and interpret embedded information cards, addressing the lack of real-time metadata generation in existing systems and enhancing user interaction.

JP2025172828AActive Publication Date: 2025-11-26STATS LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025140260
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-05-14
Filing Date
2025-08-26
Publication Date
2025-11-26
Estimated Expiration
2039-05-15

AI Technical Summary

Technical Problem

Existing television systems lack the ability to automatically and efficiently extract and interpret embedded information cards from live sporting event broadcasts for real-time metadata generation, limiting the interactivity and enhancement of television viewing experiences.

Method used

A method and system utilizing machine learning and computer vision to process sporting event television content in real-time, extracting and interpreting embedded information cards to generate metadata, which is then synchronized with video highlights.

Benefits of technology

Enables real-time extraction and synchronization of metadata with video highlights, enhancing user interaction and providing timely, relevant information during television broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025172828000001_ABST
    Figure 2025172828000001_ABST
Patent Text Reader

Abstract

To extract metadata for highlights of a video stream from a card image embedded in the video stream.SOLUTION: Highlights may be segments of a video stream, such as a broadcast of a sporting event, that are of particular interest to one or more users. Card images embedded in video frames of the video stream are identified and processed to extract text. The text characters may be recognized by applying a machine-learned model trained with a set of characters extracted from card images embedded in sports television programming contents. A training set of character vectors may be pre-processed to maximize metric distance between the training set members. The text may be interpreted to be able to obtain the metadata. The metadata may be stored in association with a portion of the video stream. The metadata may provide information regarding the highlights, and may be presented concurrently with playback of the highlights.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 673,412 (Attorney Docket No. THU010-PROV), filed May 18, 2018, for "Machine Learning for Recognizing and Interpreting Embedded Information Card Content," which is hereby incorporated by reference in its entirety.

[0002] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,710 (Attorney Docket No. THU010), filed May 14, 2019, for "Machine Learning for Recognizing and Interpreting Embedded Information Card Content," which is hereby incorporated by reference in its entirety.

[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 673,411 (Attorney Docket No. THU009-PROV), filed May 18, 2018, for "Video Processing for Enabling Sports Highlights Generation," which is incorporated herein by reference in its entirety.

[0004] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,704 (Attorney Docket No. THU009), filed May 14, 2019, for "Video Processing for Enabling Sports Highlights Generation," which is incorporated herein by reference in its entirety.

[0005] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 673,413 (Attorney Docket No. THU012-PROV), filed May 18, 2018, for "Video Processing for Embedded Information Card Localization and Content Extraction," which is hereby incorporated by reference in its entirety.

[0006] This application claims the benefit of U.S. Utility Patent Application Serial No. 16 / 411,713 (Attorney Docket No. THU012), filed May 14, 2019, for "Video Processing for Embedded Information Card Localization and Content Extraction," which is hereby incorporated by reference in its entirety.

[0007] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 680,955 (Attorney Docket No. THU007-PROV), filed June 5, 2018, for "Audio Processing for Detecting Occurrences of Crowd Noise in Sporting Event Television Programming," which is incorporated herein by reference in its entirety.

[0008] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 712,041 (Attorney Docket No. THU006-PROV), filed July 30, 2018, for "Audio Processing for Extraction of Variable Length Disjoint Segments from Television Signal," which is hereby incorporated by reference in its entirety.

[0009] This application is filed on October 16, 2018, for Detecting Occurrences of Loud Sound This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 746,454 (Attorney Docket No. THU016-PROV) for "Suitable for Applications Characterized by Short-Time Energy Bursts," which is incorporated herein by reference in its entirety.

[0010] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,915, filed August 31, 2012, and issued June 16, 2015, as U.S. Patent No. 9,060,210, for "Generating Excitement Levels for Live Performances," which is hereby incorporated by reference in its entirety.

[0011] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,927, filed August 31, 2012, and issued September 23, 2014, as U.S. Patent No. 8,842,007, for "Generating Alerts for Live Performances," which is hereby incorporated by reference in its entirety.

[0012] This application is related to U.S. Utility Patent Application Serial No. 13 / 601,933, filed August 31, 2012, and issued November 26, 2013, as U.S. Patent No. 8,595,763, for "Generating Teasers for Live Performances," which is hereby incorporated by reference in its entirety.

[0013] This application is related to U.S. Utility Patent Application Serial No. 14 / 510,481 (Attorney Docket No. THU001), filed October 9, 2014, for "Generating a Customized Highlight Sequence Depicting an Event," which is hereby incorporated by reference in its entirety.

[0014] This application is related to U.S. Utility Patent Application Serial No. 14 / 710,438 (Attorney Docket No. THU002), filed May 12, 2015, for "Generating a Customized Highlight Sequence Depicting Multiple Events," which is hereby incorporated by reference in its entirety.

[0015] This application is related to U.S. Utility Patent Application Serial No. 14 / 877,691 (Attorney Docket No. THU004), filed October 7, 2015, for "Customized Generation of Highlight Show with Narrative Component," which is hereby incorporated by reference in its entirety.

[0016] This application is related to U.S. Utility Patent Application Serial No. 15 / 264,928 (Attorney Docket No. THU005), filed September 14, 2016, for "User Interface for Interaction with Customized Highlight Shows," which is hereby incorporated by reference in its entirety.

[0017] This document relates to techniques for identifying multimedia content and associated information on television devices or video servers that deliver the multimedia content, and for enabling embedded software applications to utilize the multimedia content to provide content and services synchronized with the delivery of the multimedia content. Various embodiments relate to methods and systems for providing automated video and audio analysis used to identify and extract significant event-based video segments within sports television video content, identify video highlights, and associate metadata with such highlights for pre-game, in-game, and post-game review. [Background technology]

[0018] Enhanced television applications such as interactive advertising and enhanced program guides with pre-game, in-game, and post-game interactive applications have long been envisioned. Existing cable systems, originally designed for broadcast television, are being called upon to support a host of new applications and services, including interactive television services and enhanced (interactive) programming guides.

[0019] Several frameworks have been standardized to enable enhanced television applications, for example OpenCable (商標) These include the Enhanced TV Application Messaging specification and the Tru2way specification, which refer to interactive digital cable services delivered over cable video networks and include features such as interactive program guides, interactive advertising, and games. Additionally, cable operators' "OCAP" programs offer interactive services such as e-commerce shopping, online banking, electronic program guides, and digital video recording. These efforts enable the first generation of video-synchronized applications that synchronize with video content delivered by program producers / broadcasters, providing additional data and interactivity for television programming.

[0020] Recent developments in video and audio content analysis technologies and compatible mobile devices have opened up a range of new possibilities for developing advanced applications that operate in sync with live TV programming events. These new technologies, along with advances in computer vision and video processing, and the improved computing power of modern processors, make it possible to generate high-quality program content highlights accompanied by metadata in real time. Summary of the Invention

[0021] A method and system for automated real-time processing of sporting event television program content for embedded information card location and embedded text string recognition and interpretation is presented. In at least one embodiment, a machine-learned character classification model is generated based on a training set of characters extracted from multiple information cards (card images) embedded in the sporting event television program content. The extracted character images are processed to generate a standardized training set of multidimensional character vectors in a multidimensional vector space. Principal component analysis (PCA) is then performed on this training set to derive orthogonal basis vectors spanning the vector space of the training set.

[0022] In at least one embodiment, the dimensionality of the training set vector space is reduced by selecting a limited number of representative orthogonal vectors from the orthogonal basis. A classification model is generated for this particular projected set of alphanumeric characters that appear in the embedded information cards by utilizing a machine learning algorithm structure, which may be a known machine learning algorithm, such as a multi-class support vector machine (SVM) or convolutional neural network (CNN) algorithm.

[0023] In at least one embodiment, sporting event television program content is processed in real time to extract queries (embedded characters from strings in information cards) and set up a query infrastructure using individual character images extracted from the embedded strings. In another embodiment, individual query images are normalized to generate a query vector for each query character. These query vectors are then projected onto an orthogonal basis spanning the training vector space to generate projected query vectors. In yet another embodiment, the projected query vectors are recognized (predicted) by applying a pre-trained character classification model to each projected query vector. Finally, the predicted query characters (forming predicted strings) are interpreted through semantic extraction. In at least one embodiment, semantic extraction is performed based on known string locations in various television program card image types and based on knowledge of the positions of individual characters within the strings. In at least one embodiment, the extracted information is automatically appended to sporting event metadata associated with sporting event video highlights.

[0024] In at least one embodiment, a method for extracting metadata from a video stream includes storing at least one portion of the video stream, identifying one or more card images embedded in one or more video frames of the portion of the video stream, and then processing the one or more information card images to extract text. In yet another embodiment, the text extracted from the information card images is interpreted to generate and store metadata associated with the portion of the video stream.

[0025] In at least one embodiment, the video stream may be a broadcast of a sporting event. Portions of the video stream may be highlights deemed to be of particular interest to one or more users. Metadata may describe the highlights.

[0026] In at least one embodiment, the method may further include playing a video stream to the user during at least one of identifying the one or more card images, processing the one or more card images, and interpreting the text.

[0027] In at least one embodiment, the method may further include playing the highlights to a user and presenting metadata to the user during playback of the highlights. The metadata may provide real-time information related to the highlights and a timeline of the card image from which the metadata was obtained.

[0028] In at least one embodiment, extracting the text may include identifying one or more character strings in one or more card images and recording the position and / or size of the character images of the card images with one or more card images corresponding to each character of the one or more character strings.

[0029] In at least one embodiment, extracting text may further include disambiguating character boundaries of characters of the one or more strings of characters by performing multiple comparisons of the detected character boundaries, and purging character boundaries that occur too close to each other.

[0030] In at least one embodiment, extracting the text may further include performing image verification on the characters of the one or more character strings by establishing a contrast ratio between the low intensity pixel count and the high intensity pixel count.

[0031] In at least one embodiment, interpreting the text may include generating a query based on the text, generating an n-dimensional query feature vector, projecting the n-dimensional query feature vector onto a training set orthogonal basis, applying the projected n-dimensional query feature vector to a classification model to produce a predicted query, and extracting text meaning from the predicted query.

[0032] In at least one embodiment, the method may further include generating a training set feature vector and deriving a training set orthogonal basis using the training set feature vector.

[0033] In at least one embodiment, the method may further include generating a training set feature vector and generating a classification model using the training set feature vector and the derived training set orthogonal basis vectors.

[0034] In at least one embodiment, interpreting the text may further include using at least two selections from the group consisting of: string lengths of one or more strings of characters within the text, character boundaries and / or positions of characters within the text, and horizontal positions of character boundaries and / or characters within the text.

[0035] In at least one embodiment, storing the metadata in association with the portion of the video stream may include storing video frame numbers of one or more video frames associated with the query.

[0036] In at least one embodiment, interpreting the text may include determining field positions of characters in one or more strings of text, determining alphanumeric values ​​of the characters, and sequentially interpreting the one or more strings using the field positions and alphanumeric values.

[0037] In at least one embodiment, interpreting the text may further include obtaining position and other information regarding one or more card fields of each of the card images and using the position and other information to compensate for one or more possibly missing leading characters of one or more character strings.

[0038] In at least one embodiment, a method for generating a character recognition and classification model is described in the context of automatic video highlight generation. The method includes extracting and storing at least one portion of a video stream for which automatic highlight metadata is to be generated, identifying one or more information card images embedded in one or more video frames of the portion of the video stream, and processing the one or more information card images to extract a plurality of character images. The method further includes generating training feature vectors associated with the plurality of character images, processing the training feature vectors, training a character recognition and classification model using at least some of the training feature vectors, and then storing the processed training set and classification model. The training feature vectors may be processed in a manner that increases the uniqueness of the training feature vectors by increasing their mutual metric distance and / or by reducing the dimensionality of the overall vector space that includes the training feature vectors.

[0039] In at least one embodiment, the method may further include normalizing the character images to a standard size and / or standard lighting before generating the training feature vectors.

[0040] In at least one embodiment, generating the training feature vector may include formatting a set of n pixels extracted from the character image into an n-dimensional vector.

[0041] In at least one embodiment, the method may further include performing principal component analysis on the training feature vectors. Training a classification model using at least some of the training feature vectors may include selecting a subset of training feature orthogonal basis vectors and training a character recognition and classification model using the subset of orthogonal basis vectors.

[0042] In at least one embodiment, the orthogonal basis vectors may span an overall training feature vector space. Reducing the dimensionality of the overall training feature vector space may include selecting a limited number of orthogonal basis vectors that represent the training feature vector space sufficiently accurately. Reducing the dimensionality of the overall training vector space may include selecting only orthogonal basis vectors that correspond to a maximal set of singular values ​​derived from a matrix of orthogonal basis vectors. Storing the classification model may include storing the limited number of orthogonal basis vectors for subsequent use in classification model generation and / or query processing. Generating the classification model may include using the limited number of training set orthogonal basis vectors in combination with a machine learning algorithm selected from the group consisting of SVM and CNN.

[0043] In at least one embodiment, the method may further include processing one or more information card images to extract text, interpreting the text to obtain metadata, and storing the metadata in association with a portion of the video stream. The method further includes playing the portion of the video stream to a user and presenting the metadata to the user during playback of the portion of the video stream. The video stream may be a broadcast of a sporting event. The portion of the video stream may include highlights deemed to be of particular interest to one or more users. The metadata may describe the highlights.

[0044] In at least one embodiment, extracting the text may include extracting a text string of text as the query.

[0045] In at least one embodiment, extracting the text may include extracting at least one of a current time within the sporting event, a current phase of the sporting event, a game clock associated with the sporting event, and a game score associated with the sporting event.

[0046] Further details and variations are described herein. [Brief explanation of the drawings]

[0047] The accompanying drawings, together with the description, illustrate several embodiments. Those skilled in the art will recognize that the specific embodiments shown in the drawings are merely exemplary and are not intended to be limiting in scope. [Figure 1A] FIG. 1 is a block diagram depicting a hardware architecture according to a client / server embodiment, where event content is provided via networked content providers. [Figure 1B] 1 is a block diagram depicting a hardware architecture according to another client / server embodiment, where event content is stored on a client-based storage device. [Figure 1C] FIG. 2 is a block diagram illustrating a hardware architecture according to a stand-alone embodiment. [Figure 1D] FIG. 1 is a block diagram illustrating an overview of a system architecture, according to one embodiment. [Figure 2] FIG. 1 is a schematic block diagram illustrating examples of card images, user data, highlight data, and data structures that can be incorporated into a classification model, according to one embodiment. [Figure 3] FIG. 1 is a screenshot view of an example video frame from a video stream showing an information card image embedded within the frame, such as might be found in sporting event television program content. [Figure 4] 1 is a flowchart depicting an overall application process for real-time reception and processing of television program content for in-frame information card location and content extraction and rendering, according to one embodiment. [Figure 5] 10 is a flowchart illustrating the internal processing of detected and extracted information card images for string bounding box extraction, according to one embodiment. [Figure 6] 1 is a flowchart depicting a method for processing text boxes for verification of the final bounded character image and associated position parameter extraction, according to one embodiment. [Figure 7] 1 is a flowchart illustrating a method for query generation from text images of embedded information cards, according to one embodiment. [Figure 8] 1 is a flowchart depicting a method for generating predicted alphanumeric characters of an extracted query string based on a machine-learned classification model, according to one embodiment. [Figure 9] 1 is a flowchart illustrating a method for predicted query alphanumeric string interpretation, according to one embodiment. [Figure 10] 1 is a flowchart illustrating pre-processing of training set vectors and subsequent classification model generation based on a multi-class SVM classifier or a CNN classifier, according to one embodiment. [Figure 11] 1 is a flowchart depicting the overall process of reading and interpreting text fields in an information card and updating video highlight metadata with real-time information within a frame, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0048] definition The following definitions are provided for illustrative purposes only and are not intended to limit the scope. Event: For purposes of this description, the term "event" refers to a game, session, matchup, series, performance, program, and / or concert, etc., or portions thereof (such as an act, period, quarter, half, inning, scene, or chapter). An event may be a sporting event, an entertainment event, or a specific performance of a single individual or a subset of individuals within a larger group of event participants. Examples of non-sporting events include television shows, breaking news, sociopolitical events, natural disasters, movies, plays, radio programs, podcasts, audiobooks, online content, and / or musical performances. An event can be of any length. For illustrative purposes, the technology is often described herein in terms of sporting events, but those skilled in the art will recognize that the technology can be used in other contexts, including highlight shows of any audiovisual, audio, qualification, graphics-based, interactive, non-interactive, or text-based content. Thus, the use of the term “sporting event” and any other sports-specific terminology in this description is intended to illustrate one contemplated embodiment, but is not intended to limit the scope of the described technology to that one embodiment. Rather, such terms should be considered to extend to any suitable non-sporting context appropriate to this technology. For ease of explanation, the term “event” is also used to refer to a report or representation of an event, such as an audiovisual recording of the event, or any other content item that includes a report, description, or depiction of the event. Highlights: An excerpt or portion of an event, or content associated with an event, that is deemed to be of particular interest to one or more users. Highlights can be of any length. Generally, the technology described herein provides a mechanism for identifying and presenting a customized set of highlights (which may be selected based on specific characteristics and / or user preferences) for any suitable event.The term "highlight" is also used to refer to a report or representation of a highlight, such as an audiovisual recording of a highlight, or any other content item that includes a report, description, or depiction of a highlight. Highlights need not be limited to depictions of the event itself, but can include other content associated with the event. For example, in the case of a sporting event, highlights can include audio / video from the game as well as other content, including pre-game, in-game, and post-game interviews, analysis, and / or commentary. Such content can be recorded from linear television (e.g., as part of a video stream depicting the event itself) or can be derived from any number of other sources. Various types of highlights can be provided, including, for example, occurrences (plays), strings, possessions, and sequences, all of which are defined below. Highlights need not be of a fixed duration, but can incorporate a start offset and / or end offset, as described below. Content delineator: One or more video frames that indicate the beginning or end of a highlight. Occurrence: Something that occurs during an event. Examples include a goal, a play, a down, a hit, a save, a shot on goal, a basket, a steal, a snap or snap attempt, a near miss, a fight, the start or end of a game, a quarter, half, period, or inning, a pitch, a penalty, an injury, a dramatic event at an entertainment event, a song, and / or a solo. An occurrence can also be an unusual incident, such as a power outage and / or an incident with an unruly fan. Detection of such an occurrence can be used as the basis for determining whether to designate a particular portion of a video stream as a highlight. An occurrence is also referred to herein as a "play" for ease of naming, although such usage should not be construed as limiting in scope. An occurrence may have any length, and representations of occurrences may have varying lengths.For example, as described above, an extended representation of an occurrence may include footage depicting the time period immediately before and after the occurrence, while a simple representation may include only the occurrence itself. Optional intermediate representations may also be provided. In at least one embodiment, the selection of a duration for representing an occurrence may vary depending on user preferences, available time, a determined excitement level for the occurrence, the importance of the occurrence, and / or any other factors. Offset: An amount by which the length of a highlight is adjusted. In at least one embodiment, a start offset and / or end offset may be provided to adjust the start time and / or end time of the highlight, respectively. For example, if a highlight depicts a goal, the highlight may be extended by several seconds (via an end offset) to include the celebration and / or fan reaction following the goal. The offset may be configured to vary automatically or manually based, for example, on the time available for the highlight, the importance and / or excitement level of the highlight, and / or any other suitable factors. String: A series of occurrences that are linked or related to each other in some way. An occurrence may occur within a possession (defined below) or across multiple possessions. An occurrence may occur within a sequence (defined below) or across multiple sequences. Occurrences may be linked or related because they have some thematic or narrative connection to one another, or because one leads to another, or for any other reason. An example of a string is a set of passes leading to a goal or basket. This should not be confused with "string of text," which has the meaning typically assigned to it in computer programming. Possession: Any time-bound portion of an event. The distinction between the start and end times of a possession may vary depending on the type of event. For certain sporting events (e.g., basketball or soccer) where one team can be offensive while the other team is defensive, a possession can be defined as the period of time one team has the ball.In sports where possession of the puck or ball is more fluid, such as hockey or soccer, possession is considered to extend to the period of time when one of the teams has substantial control of the puck or ball, ignoring momentary contact by the other team (such as a blocked shot or save). In baseball, a possession is defined as a half-inning. In soccer, a possession can include several sequences in which the same team has the ball. For other types of sporting and non-sporting events, the term "possession" may be somewhat misleading, but is still used herein for illustrative purposes. Examples in non-sporting contexts include a chapter, scene, act, or television segment. For example, in the context of a music concert, a possession might correspond to the performance of a single song. A possession can include any number of occurrences. Sequence: A time-delimited portion of an event that encompasses the time period of one continuous action. For example, in a sporting event, a sequence may begin at the start of an action (such as a face-off or tip-off) and end when the whistle is blown to indicate a stoppage of action. In sports such as baseball or soccer, a sequence may equate to a play, which is a form of occurrence. A sequence can include any number of possessions or may be a portion of a possession. Highlight show: A set of highlights arranged for presentation to a user. A highlight show may be presented linearly (e.g., a video stream) or in a way that allows the user to choose which highlights to watch and in what order (e.g., by clicking links or thumbnails). The presentation of a highlight show may be non-interactive or interactive, allowing, for example, the user to pause, rewind, skip, fast-forward, and / or communicate preferences. A highlight show may be, for example, a condensed game.A highlight show can include any number of consecutive or non-consecutive highlights from a single event or from multiple events, and can even include highlights from different types of events (e.g., a combination of highlights from different sports and / or sporting and non-sporting events). User / Viewer: The terms "user" or "viewer" refer interchangeably to an individual, group, or other entity that watches, listens to, or otherwise experiences an event, one or more highlights from an event, or a highlight show. User or viewer can also refer to an individual, group, or other entity that watches, listens to, or otherwise experiences an event, one or more highlights from an event, or a highlight show at some future point in time. While the term "viewer" is sometimes used for descriptive purposes, an event need not include a visual component, and therefore a "viewer" may instead be a listener or any other consumer of content. Narrative: A coherent story that links a set of highlight segments in a particular order. Excitement Level: A measure of how exciting or interesting an event or highlight will be to a particular user or users in general. Excitement level can also be determined with respect to a particular occurrence or player. Various techniques for measuring or assessing excitement levels are described in the related applications referenced above. As described, excitement levels may vary depending on the occurrence within an event and other factors, such as the overall context or importance of the event (e.g., playoff matches, pennant influence, and / or rivalries). In at least one embodiment, an excitement level may be associated with each occurrence, string, possession, or sequence within an event. For example, the excitement level of a possession may be determined based on the occurrences occurring within that possession. Excitement levels may be measured differently by different users (e.g., fans of a team and neutral fans) and may vary depending on each user's personal characteristics.Metadata: Data that relates to and is stored in association with other data. Primary data may be media such as a sports program or highlights. Card Image: An image within a video frame that provides data about something depicted in the video, such as an event, a depiction of an event, or a portion thereof. Exemplary card images include the match score, match clock, and / or other statistics from a sporting event. Card images may appear momentarily or for the entire duration of the video stream, and those that appear momentarily may be specifically related to the portion of the video stream in which they appear. Character Image: A portion of an image that appears to relate to a single character. A character image may include the area surrounding the character. For example, a character image may include a roughly rectangular bounding box that surrounds the character. Character: A word, a number, or a symbol that can be part of the representation of a word or number. Characters can include letters, numbers, and special characters and may be in any language. String: A set of characters grouped in a way that indicates they relate to a single piece of information, such as the names of teams playing in a sporting event. Often, strings in English are arranged horizontally and read from left to right. However, strings may be arranged differently in English than in other languages.

[0049] overview According to various embodiments, methods and systems are provided for automatically creating time-based metadata associated with highlights of a television program of a sporting event. The highlights and associated intra-frame time-based information may be extracted synchronously with respect to the television broadcast of the sporting event, or may be extracted while the video content of the sporting event is being streamed from a backup device via a video server after the television broadcast of the sporting event.

[0050] In at least one embodiment, a software application operates synchronously with the playback and / or reception of television program content to provide informational metadata associated with highlights of the content. Such software may run, for example, on the television device itself, on an associated set-top box (STB), on a video server capable of receiving and subsequently streaming program content, or on a mobile device capable of receiving a video feed containing live programming. In at least one embodiment, the highlight and associated metadata application operates synchronously with the presentation of television program content.

[0051] Interactive television applications can enable timely and relevant presentation of highlighted television program content to users watching television programs on either a primary television display or a secondary display such as a tablet, laptop, or smartphone. A set of video clips representing highlights of the television broadcast content can be generated and / or stored in real time, along with a database including time-based metadata that more fully describes the events presented by the highlight video clips.

[0052] The metadata accompanying a video clip can be any information, such as text information, a set of images, and / or any type of audiovisual data. One type of metadata associated with highlights of in-game and post-game video content conveys real-time information about sports match parameters extracted directly from the live program content by reading information cards (“card images”) embedded in one or more of the video frames of the program content. In at least one embodiment, the described systems and methods enable this type of automatic metadata generation, thus associating card image content with video highlights of the analyzed digital video stream.

[0053] In various embodiments, an automated process is described that includes receiving a digital video stream, analyzing one or more video frames of the digital video stream to present and extract a card image, locating a text box within the card image, and recognizing and interpreting a string of characters present within the text box.

[0054] The automated metadata generation video system presented herein can receive a live broadcast video stream or digital video streamed via a computer server and can process the video stream in real time using computer vision and machine learning techniques to extract metadata from embedded information cards.

[0055] In at least one embodiment, character strings associated with the extracted information card text fields are identified, and the position and size of the image of each character within the string of characters is recorded. Any number of characters within the string of text from various fields of the information card are then recognized, and the text string with the recognized characters is interpreted to provide real-time information related to a television broadcast of the sporting event, such as the current time and phase of the match, the match score, and / or play information.

[0056] In another embodiment, individual character images are extracted from the embedded string and then used to generate normalized query vectors. These normalized query vectors are then projected onto an orthogonal basis spanning a training vector space, which has been pre-assembled and used to train a machine learning classifier, such as a multi-class support vector machine (SVM) classifier (e.g., C. Burges, “A Tutorial on Support Vector Machines for Pattern Recognition,” Kluwer Academic Publishers, 1998). The projected query is then used to generate query predictions as the output of a pre-trained classification model produced by an exemplary SVM training mechanism. Note that the classification model is not limited to SVM-based models. Classification models can also be implemented using other techniques, such as convolutional neural networks (CNNs), and numerous variations of the CNN algorithm mechanism (e.g., Y. LeCun at al., "Efficient NN Back Propagation", Springer 1998), and this variant is preferred for the training dataset presented here.

[0057] In yet another embodiment, query character predictions are generated by applying the projected query character vector to a pre-developed machine-learned classification model. In this step, a string of predicted characters is generated according to pre-established classification labels, and the predicted string of alphanumeric characters is passed to a recognition and interpretation process. The query recognition and interpretation process applies prior knowledge and positional understanding of characters present in multiple information card fields. The meaning of each predicted alphanumeric character located in a specific character group is further interpreted, and the derived information is added to the video highlight metadata handled by the video highlight generation application.

[0058] In yet another embodiment, character classification model generation is considered, where the model is based on a training set of characters extracted from any number of information cards embedded in sporting event television program content. Character bounding boxes are detected to extract characters from the multiple information cards. These character images are then normalized to a standard size and lighting to form descriptors associated with each specific character from the set of alphanumeric characters appearing on the embedded information cards. In this manner, each extracted character image represents an n-dimensional vector in a multidimensional vector space comprising a training set of vectors. The n-dimensional training vectors representing the set of character images are further processed to increase uniqueness and mutual metric distance and to reduce the dimensionality of the overall vector space of training vectors.

[0059] In at least one embodiment, principal component analysis (see, e.g., G. Golub and F. Loan, "Matrix Computations," Johns Hopkins Univ. Press, Baltimore, 1989) is performed on the training vector set. An orthogonal basis of vectors is then devised from the training set such that the orthogonal basis vectors span the vector space of the training set. Furthermore, the dimensionality of the training set vector space is reduced by selecting a limited number of orthogonal basis vectors such that only the most significant orthogonal vectors associated with the largest set of singular values, generated by singular value decomposition of the training set matrix of basis vectors, are retained. The selected training set basis vectors are then saved for later use in generating a classification model using one or more of the algorithmic structures available for dataset classification, such as a multiclass SVM-based classifier or a CNN-based classifier.

[0060] System Architecture According to various embodiments, the system can be implemented in any electronic device or set of electronic devices equipped to receive, store, and present information, such as, for example, a desktop computer, a laptop computer, a television, a smartphone, a tablet, a music player, an audio device, a kiosk, a set-top box (STB), a gaming system, a wearable device, and / or a consumer electronic device.

[0061] Although the system is described herein with reference to implementation on a particular type of computing device, those skilled in the art will recognize that the techniques described herein may be implemented in other contexts, and indeed on any suitable device capable of receiving and / or processing user input and presenting output to a user. Accordingly, the following description is intended to illustrate various embodiments by way of example, rather than to limit the scope.

[0062] 1A, a block diagram is shown depicting the hardware architecture of a system 100 for automatically extracting metadata from card images embedded in a video stream of an event, according to a client / server embodiment. Event content, such as a video stream, may be provided via a network-connected content provider 124. An example of such a client / server embodiment is a web-based implementation, in which one or more client devices 106 each run a browser or app that provides a user interface for interacting with content from various servers 102, 114, 116, including data provider(s) server(s) 122 and / or content provider(s) server(s) 124, via a communications network 104. Transmission of content and / or data in response to requests from the client devices 106 may be performed using any known protocols and languages, such as Hypertext Markup Language (HTML), Java, Objective C, Python, and / or JavaScript.

[0063] The client device 106 may be a desktop computer, a laptop computer, a television, a smartphone, a tablet, a music player, an audio device, a kiosk, a set-top box, a gaming system, a wearable device, a consumer electronic device, and / or any electronic device, etc. In at least one embodiment, the client device 106 has several hardware components known to those skilled in the art. The input device(s) 151 may be any component(s) that receive input from the user 150, including, for example, a handheld remote control, a keyboard, a mouse, a stylus, a touch-sensitive screen (touch screen), a touchpad, a gesture receptor, a trackball, an accelerometer, a five-way switch, or a microphone. The input may be provided via any suitable mode, including, for example, one or more of pointing, tapping, typing, dragging, gesturing, tilting, shaking, and / or speech. The display screen 152 may be any component that graphically displays information, video, and / or content, including depictions of events and / or highlights, etc. Such output may also include, for example, audiovisual content, data visualization, navigational elements, graphical elements, or queries requesting information and / or parameters for content selection, etc. In at least one embodiment in which only some of the desired outputs are presented at a time, dynamic control, such as a scrolling mechanism, may be available via input device(s) 151 to select which information is currently displayed and / or to change how the information is displayed.

[0064] Processor 157 may be a conventional microprocessor for performing operations on data under the direction of software in accordance with well-known techniques. Memory 156 may be random access memory having a structure and architecture known in the art for use by processor 157 in the course of executing software to perform the operations described herein. Client device 106 may also include local storage (not shown), which may be a hard drive, flash drive, optical or magnetic storage device, and / or web-based (cloud-based) storage, etc.

[0065] Any suitable type of communications network 104, such as the Internet, a television network, a cable network, and / or a cellular network, may be used as a mechanism for transmitting data between the client device 106 and the various server(s) 102, 114, 116 and / or content provider(s) 124 and / or data provider(s) 122 according to any suitable protocols and techniques. In addition to the Internet, other examples include cellular networks, EDGE, 3G, 4G, Long Term Evolution (LTE), Session Initiation Protocol (SIP), Short Message Peer-to-Peer Protocol (SMPP), SS7, Wi-Fi, Bluetooth, ZigBee, Hypertext Transfer Protocol (HTTP), Secure Hypertext Transfer Protocol (SHTTP), and / or Transmission Control Protocol / Internet Protocol (TCP / IP), etc., and / or any combination thereof. In at least one embodiment, the client device 106 transmits requests for data and / or content over the communications network 104 and receives responses from the servers 102, 114, 116 that include the requested data and / or content.

[0066] 1A operates in connection with a sporting event. However, it should be understood that the teachings herein apply to events other than sporting events, and the techniques described herein are not limited to application to sporting events. For example, the techniques described herein can be utilized to operate in connection with or for television shows, movies, news events, game shows, political campaigns, business shows, dramas, and / or other episodic content.

[0067] In at least one embodiment, system 100 identifies highlights of a broadcast event by analyzing a video stream of the event. This analysis can be performed in real time. In at least one embodiment, system 100 includes one or more web server(s) 102 coupled to one or more client devices 106 via a communications network 104. Communications network 104 may be a public network, a private network, or a combination of public and private networks, such as the Internet. Communications network 104 may be a LAN, a WAN, wired, wireless, and / or a combination of the above. Client device 106, in at least one embodiment, can connect to communications network 104 via either a wired or wireless connection. In at least one embodiment, client device 106 may also include a recording device, such as a DVR, PVR, or other media recording device, capable of receiving and recording the event. Such a recording device may be part of client device 106 or may be external. In other embodiments, such a recording device may be omitted. Although FIG. 1A shows one client device 106, the system 100 may be implemented with any number of client device(s) 106 of a single type or multiple types.

[0068] Web server(s) 102 may include one or more physical computing devices and / or software capable of receiving requests from client device(s) 106, responding to those requests with data, as well as sending unsolicited alerts and other messages. Web server(s) 102 may employ various strategies for fault tolerance and scalability, such as load balancing, caching, and clustering. In at least one embodiment, web server(s) 102 may include caching techniques, as known in the art, for storing information related to client requests and events.

[0069] The web server(s) 102 may maintain or otherwise designate one or more application server(s) 114 to respond to requests received from the client device(s) 106. In at least one embodiment, the application server(s) 114 provide access to business logic for use by client application programs in the client device(s) 106. The application server(s) 114 may be co-located, shared, or co-managed with the web server(s) 102. The application server(s) 114 may also be remote from the web server(s) 102. In at least one embodiment, the application server(s) 114 interact with one or more analytics server(s) 116 and one or more data server(s) 118 to perform one or more operations of the disclosed techniques.

[0070] One or more storage devices 153 may function as a "data store" by storing data related to the operation of system 100. This data may include, but is not limited to, card data 154 associated with card images embedded in a video stream presenting an event, such as a sporting event, user data 155 associated with one or more users 150, highlight data 164 associated with one or more highlights of an event, and / or classification models 165 that can be used to predict and / or extract text from card data 154.

[0071] Card data 154 may include any information related to card images embedded in the video stream, such as the card images themselves, subsets thereof, such as character images, text extracted from card images, such as characters and strings of characters, and any of the aforementioned attributes useful for extracting text and / or meaning. User data 155 may include any information describing one or more users 150, including, for example, demographics, purchasing behavior, video stream viewing behavior, interests, and / or preferences. Highlight data 164 may include highlights, highlight identifiers, time indexes, categories, excitement levels, and other data related to the highlights. Classification model 165 may include machine-trained classification models, queries, query feature vectors, training set orthogonal bases, predicted queries, extracted text meanings, and / or other information that facilitates extraction of text and / or meaning from card data 154. Card data 154, user data 155, highlight data 164, and classification model 165 are described in more detail below.

[0072] Notably, many components of system 100 may be or include computing devices. Each of such computing devices may have an architecture similar to that of client device 106, as shown and described above. Accordingly, any of communication network 104, web server 102, application server 114, analytics server 116, data provider 122, content provider 124, data server 118, and storage device 153 may include one or more computing devices, which may optionally have input device 151, display screen 152, memory 156, and / or processor 157, as described above in connection with client device 106.

[0073] In an exemplary operation of the system 100, one or more users 150 of the client devices 106 view content from the content provider 124 in the form of a video stream. The video stream may depict an event, such as a sporting event. The video stream may be a digital video stream that can be easily processed with known computer vision techniques.

[0074] Once the video stream is displayed, one or more components of the system 100, such as the client device 106, the web server 102, the application server 114, and / or the analytics server 116, may analyze the video stream to identify highlights within the video stream and / or extract metadata from the video stream, for example, from embedded card images and / or other aspects of the video stream. This analysis may be performed in response to receiving a request to identify highlights and / or metadata from the video stream. Alternatively, in another embodiment, highlights may be identified without a specific request by the user 150. In yet another embodiment, analysis of the video stream may occur without the video stream being displayed.

[0075] In at least one embodiment, user 150 can specify certain parameters for the analysis of the video stream (e.g., which events / matches / teams to include, how much time user 150 has available to watch highlights, what metadata is desired, and / or any other parameters, etc.) via input device(s) 151 of client device 106. User preferences can also be retrieved from storage, such as from user data 155 stored on one or more storage devices 153, to customize the analysis of the video stream without necessarily requiring user 150 to specify preferences. In at least one embodiment, user preferences can be determined based on observed behavior and actions of user 150, for example, by observing website visiting patterns, television viewing patterns, music listening patterns, online purchases, prior highlight identification parameters, and / or highlights and / or metadata actually viewed by user 150, etc.

[0076] Additionally or alternatively, user preferences may be retrieved from pre-stored preferences explicitly provided by user 150. Such user preferences may indicate which teams, sports, players, and / or types of events are of interest to user 150, and / or they may indicate what type of metadata or other information associated with highlights will be of interest to user 150. Such preferences may therefore be used to guide analysis of the video stream to identify highlights and / or extract metadata for highlights.

[0077] Analysis server(s) 116, which may include one or more computing devices described above, can analyze live and / or recorded feeds of play-by-play statistics related to one or more events from data provider(s) 122. Examples of data provider(s) 122 include, but are not limited to, providers of real-time sports information such as STATS™, Perform (available from Opta Sports, London, UK), and SportRadar, St. Gallen, Switzerland. In at least one embodiment, analysis server(s) 116 generate a set of different excitement levels for the event. Such excitement levels can then be stored in association with highlights identified by system 100 according to the techniques described herein.

[0078] The application server(s) 114 may analyze the video stream to identify highlights and / or extract metadata. Additionally or alternatively, such analysis may be performed by the client device(s) 106. The identified highlights and / or extracted metadata may be specific to a user 150; in such cases, it may be advantageous to identify highlights within the client device 106 that are associated with the particular user 150. The client device 106 may receive, maintain, and / or acquire applicable user preferences for highlight identification and / or metadata extraction, as described above. Additionally or alternatively, highlight generation and / or metadata extraction may be performed globally (i.e., using objective criteria generally applicable to a user population, regardless of the preferences of a particular user 150). In such cases, it may be advantageous to identify highlights and / or extract metadata within the application server(s) 114.

[0079] The content facilitating highlight identification and / or metadata extraction may come from any suitable source, including content provider(s) 124, including websites such as YouTube® and MLB.com, sports data providers, television stations, and / or client- or server-based DVRs. Alternatively, the content may come from a local source, such as a DVR or other recording device associated with (or embedded in) the client device 106. In at least one embodiment, the application server(s) 114 generate a customized highlight show with highlights and metadata available to the user 150, either as download, or streaming content, or on-demand content, or in some other manner.

[0080] As noted above, it may be advantageous for user-specific highlight identification and / or metadata extraction to be performed at a particular client device 106 associated with a particular user 150. Such an embodiment may avoid the need for video content or other high-bandwidth content to be unnecessarily transmitted over the communications network 104, especially if such content is already available at the client device 106.

[0081] For example, referring now to FIG. 1B , an example of a system 160 according to one embodiment is shown in which at least some of the card data 154, highlight data 164, and classification model 165 are stored on a client-based storage device 158, which may be any type of local storage device available to the client device 106. An example includes a DVR capable of recording events, such as video content of a complete sporting event. Alternatively, the client-based storage device 158 may be any magnetic, optical, or electronic storage device for data in digital format. Examples include flash memory, a magnetic hard drive, a CD-ROM, a DVD-ROM, or other devices integrated with or communicatively coupled to the client device 106. Based on information provided by the application server(s) 114, the client device 106 may extract metadata from the card data 154 stored on the client-based storage device 158 and store the metadata as highlight data 164, without having to retrieve other content from the content provider 124 or other remote sources. Such a configuration can conserve bandwidth and can make effective use of existing hardware that may already be available for the client device 106 .

[0082] Returning to FIG. 1A , in at least one embodiment, application server(s) 114 can identify different highlights and / or extract different metadata for different users 150 depending on individual user preferences and / or other parameters. The identified highlights and / or extracted metadata may be presented to the user 150 via any suitable output device, such as a display screen 152 of the client device 106. If desired, multiple highlights can be identified and organized into a highlight show along with associated metadata. Such highlight shows may be assembled into a “highlight reel” or set of highlights that are accessed via a menu and / or played for the user 150 according to a predetermined sequence. The user 150, in at least one embodiment, can control the highlight playback and / or delivery of associated metadata via input device(s) 151, for example, to: select particular highlights and / or metadata for display; pause, rewind, and fast-forward; skip to the next highlight; return to the beginning of the previous highlight within the highlight show; and / or perform other actions.

[0083] Additional details regarding such functionality are provided in the related US patent applications cited above.

[0084] In at least one embodiment, another data server(s) 118 is provided. The data server(s) 118 may respond to requests for data from any of the server(s) 102, 114, 116, for example, to obtain or provide card data 154, user data 155, highlight data 164, and / or classification models 165. In at least one embodiment, such information may be stored in any suitable storage device 153 accessible by the data server 118 and may come from any suitable source, such as the client device 106 itself, the content provider(s) 124, and / or the data provider(s) 122.

[0085] 1C , an alternative embodiment of system 180 is shown in which system 180 is implemented in a standalone environment. Similar to the embodiment shown in FIG. 1B , at least some of card data 154, user data 155, highlight data 164, and classification model 165 may be stored on a client-based storage device 158, such as a DVR. Alternatively, client-based storage device 158 may be a flash memory or hard drive, or other device integrated with or communicatively coupled to client device 106.

[0086] User data 155 may include preferences and interests of user 150. Based on such user data 155, system 180 can extract metadata in card data 154 and present it to user 150 in the manner described herein. Additionally or alternatively, metadata can be extracted based on objective criteria that are not based on information specific to user 150.

[0087] 1D , an overview of a system 190 having an architecture according to an alternative embodiment is shown. In FIG. 1D , the system 190 includes broadcast services, such as content provider(s) 124, content receivers in the form of client devices 106, such as television sets with STBs, video servers, such as analysis server(s) 116, that can ingest and stream television program content, and / or other client devices 106, such as mobile devices and laptops, that can receive and process television program content, all connected via a network, such as the communications network 104. A client-based storage device 158, such as a DVR, can be connected to any of the client devices 106 and / or other components and can store video streams, highlights, highlight identifiers, and / or metadata to facilitate identification and presentation of highlights and / or extracted metadata via any of the client devices 106.

[0088] The particular hardware architectures depicted in Figures 1A, 1B, 1C, and 1D are merely exemplary. Those skilled in the art will recognize that the techniques described herein can be implemented using other architectures. Many of the components depicted herein are optional and may be omitted, combined with, and / or replaced by other components.

[0089] In at least one embodiment, the system can be implemented as software written in any suitable computer programming language, whether in a stand-alone or client / server architecture, or it may be implemented and / or embedded in hardware.

[0090] Data Structure FIG. 2 is a schematic block diagram illustrating example data structures that may be incorporated into card data 154, user data 155, highlight data 164, and classification model 165, according to one embodiment.

[0091] As shown, card data 154 may include a record for each of multiple card images embedded in one or more video streams. Each of the card images may include one or more character strings 200. Each of the character strings 200 may have a record for n characters. Each such record may include a character image 202, a processed character image 203, a character boundary 204, a size 205, a position 206, a contrast ratio 207, and / or an interpretation 208. Each of the character strings 200 may further include a string length 209 indicating the length of the character string 200 (e.g., the length of characters, pixels, etc.).

[0092] Character image 202 may be a particular portion of a card image containing a single character. Processed character image 203 may be character image 202 after application of one or more processing steps, such as normalization for size and / or brightness.

[0093] Character boundaries 204 may indicate boundaries of character images 202 , processed character images 203 , and / or characters represented by character images 202 and processed character images 203 .

[0094] Size 205 may be the size, eg, in pixels, of character image 202, processed character image 203, and / or the character represented by character image 202 and processed character image 203.

[0095] Location 206 may be a location within the card image of character image 202, processed character image 203, and / or a character represented by character image 202 and processed character image 203. In some examples, location 206 may indicate a two-dimensional location (e.g., x and y coordinates of a corner or center of character image 202, processed character image 203, and / or a character represented by character image 202 and processed character image 203).

[0096] Contrast ratio 207 may be a measure of the contrast of character image 202, processed character image 203, and / or a character represented by character image 202 and processed character image 203. In some examples, contrast ratio 207 may be a ratio of the luminance value of one or more lightest pixels to the luminance value of one or more darkest pixels in character image 202, processed character image 203, and / or a character represented by character image 202 and processed character image 203.

[0097] The interpretation 208 may be a particular character, such as a, b, c, 1, 2, 3, #, &, etc., that is believed to be represented in the character image 202 after some analysis has been performed to interpret the character string 200.

[0098] The structure of card data 154 shown in FIG. 2 is merely exemplary, and in some embodiments, data associated with card images embedded in a video stream may be organized differently. For example, in other embodiments, each character string may not necessarily be broken down into individual character images. Rather, the character string may be interpreted as a whole, and data useful for interpreting the character string may be stored for the entire character string. Additionally, in alternative embodiments, data not specifically described above may be incorporated into card data 154. The structures of user data 155, highlight data 164, and classification model 165 in FIG. 2 are likewise merely exemplary, and many alternatives may be envisioned by those skilled in the art.

[0099] As further shown, user data 155 may include records associated with users 150, each of which may include demographic data 212, preferences 214, viewing history 216, and purchasing history 218 for a particular user 150.

[0100] Demographic data 212 may include any type of demographic data, including, but not limited to, age, gender, location, nationality, religious affiliation, and / or education level.

[0101] Preferences 214 may include selections made by user 150 regarding their preferences. Preferences 214 may relate directly to the collection and / or display of highlights and metadata, or may be more general in nature. In either case, preferences 214 may be used to facilitate the identification and / or presentation of highlights and metadata to user 150.

[0102] The viewing history 216 may list television programs, video streams, highlights, web pages, search queries, sporting events, and / or other content retrieved and / or viewed by the user 150 .

[0103] The purchase history 218 may list products or services purchased or requested by the user 150 .

[0104] As further shown, the highlight data 164 may include records of highlights 220 , each of which may include a video stream 222 , an identifier, and / or metadata 224 for a particular highlight 220 .

[0105] Video stream 222 may include video depicting highlight 220, which may be obtained from one or more video streams of one or more events (e.g., by trimming the video streams to include only the video stream 222 associated with highlight 220). Identifier 223 may include a time code and / or other indicator that indicates where highlight 220 resides within the video stream of the event from which it was obtained.

[0106] In some embodiments, each recording of highlight 220 may include only one of video stream 222 and identifier 223. Highlight playback may be performed by playing user 150's video stream 222 or by using identifier 223 to play only the highlighted portion of the video stream of the event from which highlight 220 was captured.

[0107] The metadata 224 may include information about the highlight 220, such as the date of the event, the season, and the groups or individuals involved in the event or video stream from which the highlight 220 was captured, such as the team, players, coaches, anchors, broadcasters, and / or fans. Among other information, the metadata 224 for each highlight 220 may include the time 225, the phase 226, the clock 227, the score 228, and / or the frame number 229.

[0108] The time 225 may be the time within the video stream 222 at which the highlight 220 is captured or the time within the video stream 222 associated with the highlight 220 for which metadata is available. In some examples, the time 225 may be the play time within the video stream 222 associated with the highlight 220 at which the card image including the metadata 224 is displayed.

[0109] Phase 226 may be a phase of an event associated with highlight 220. More specifically, phase 226 may be a stage of a sporting event during which card images containing metadata 224 are displayed. For example, phase 226 may be the "third quarter," "second inning," or "bottom half," etc.

[0110] Clock 227 may be the match clock associated with highlight 220. More specifically, clock 227 may be the state of the match clock when a card image is displayed that includes metadata 224. For example, clock 227 may read "15:47" for a card image that is displayed with 15 minutes and 47 seconds displayed on the match clock.

[0111] The score 228 may be the match score associated with the highlight 220. More specifically, the score 228 may be the score at the time the card image containing the metadata 224 is displayed. For example, the score 228 may be "45-38," "7-0," or "30-love," etc.

[0112] The frame number 229 may be the number of the video frame in the video stream from which the highlight 220 is captured, or the number of the video frame in the video stream 222 associated with the highlight 220 that is most directly associated with the highlight 220. More specifically, the frame number 229 may be the number of such a video frame in which the card image containing the metadata 224 is displayed.

[0113] As further shown, classification model 165 may include various information that facilitates extraction and interpretation of string 200. Classification model 165 may then enable automatic generation of metadata 224 for highlight 220. Specifically, classification model 165 may include query 230, query feature vector 232, orthogonal basis 234, predicted query 236, and / or text meaning 238.

[0114] The operation of query 230, query feature vector 232, orthogonal basis 234, and predicted query 236 are described in more detail herein. Text meaning 238 may be an interpretation of string 200 rendered in a way that can be easily copied into metadata 224.

[0115] The data structures depicted in Figure 2 are merely exemplary. Those skilled in the art will recognize that in performing highlight identification and / or metadata extraction, some of the data in Figure 2 can be omitted or replaced with other data. Additionally or alternatively, data not shown in Figure 2 can be used in performing highlight identification and / or metadata extraction.

[0116] Card Image Referring now to Figure 3, there is shown a screenshot illustration of an example video frame 300 from a video stream with embedded information in the form of card images, as often occurs in television programs of sporting events. Figure 3 depicts a card image 310 at the bottom right of the video frame 300 and a second card image 320 extending along the bottom of the video frame 300. The card images 310, 320 may include embedded information such as the game phase, the current clock, and the current score.

[0117] In at least one embodiment, information within the card images 310, 320 is located and processed for automatic recognition and interpretation of embedded text within the card images 310, 320. The interpreted text may then be assembled into text metadata that describes the status of a sports match at a particular point in time within the sporting event timeline.

[0118] In particular, card image 310 may relate to the currently shown sporting event, while second card image 320 may contain information related to a different sporting event. In some embodiments, only card images containing information deemed relevant to the currently playing sporting event are processed for metadata generation. Accordingly, without limiting scope, the following exemplary description assumes that only card image 310 is processed. However, in alternative embodiments, it may be desirable to process multiple card images in a given video frame 300, even including card images related to other sporting events.

[0119] 3, the card image 310 may provide several different types of metadata 224, including team name 330, score 340, prior team performance 350, current match stage 360, match clock 370, playing status 380, and / or other information 390. Each of these may be extracted from within the card image 310 and interpreted to provide metadata 224 corresponding to the highlights 220, including video frames 300, and more specifically, the video frames 300 in which the card image 310 is displayed.

[0120] Metadata Extraction 4 is a flowchart depicting a method 400 performed by an application executing, for example, on one of the client device 106 and / or the analytics server 116, to receive a video stream 222 and perform on-the-fly processing of video frames 300 to extract metadata from a card image, such as card image 310, according to one embodiment. System 100 of FIG. 1A is referred to as the system that performs method 400 and the systems that follow. However, alternative systems, including, but not limited to, system 160 of FIG. 1B, system 180 of FIG. 1C, and / or system 190 of FIG. 1D, may be used in place of system 100 of FIG. 1A.

[0121] Method 400 of FIG. 4 depicts the process outlined above in more detail. A video stream, such as video stream 222 corresponding to pre-identified highlight 220, may be received and decoded. At step 410, one or more video frames 300 of video stream 222 may be received, resized to a standard size, and decoded. At step 420, video frame 300 may be processed to detect and, if applicable, extract one or more card images, such as card image 310 of FIG. 3, from video frame 300. If a valid card image 310 is not found in video frame 300 according to query 430, method 400 may return to step 410 to decode and analyze a different video frame 300.

[0122] If a valid card image 310 is found, then in step 440, the video frame 300 may be further processed to locate, extract, and process the detected card image 310, and to extract and process any text boxes and / or strings of characters embedded in the card image 310. If a valid string of characters 200 is not found in the card image 310 according to query 450, then the method 400 may return to step 410 to process a new video frame 300.

[0123] If a valid string 200 is found in the card image 310, the method 400 may proceed to step 460, where the extracted string(s) 200 are recognized and interpreted, and corresponding metadata 224 is generated based on the interpretation of information from the card image 310. In various embodiments, the available choices for text interpretation are based on determining the card image type of the card image 310 detected in the video frame 300 and / or prior knowledge of detected fields present within the particular type of card image applicable to the card image 310 detected in the video frame 300.

[0124] As previously indicated, the detection, location, and interpretation of embedded text in card images present in television program content may be performed entirely locally on the TV, STB, or mobile device, or remotely on a remote video server with broadcast video capture and streaming capabilities, or any combination of local and remote processing may be used.

[0125] Information Card String Processing: Location and Extraction An "Extreme Region" (ER) is an image region whose outer boundary pixels have strictly higher values ​​than the region itself (e.g., Neumann, J. Matas, "Real-Time Scene Text Localization and Recognition", 5th IEEE Conference on Computer Vision and (Pattern Recognition, Providence, RI, June 2012). One well-known method used to detect ERs in images uses the so-called maximally stable ER detector, or MSER detector. Additional detection methods allow for the examination of a wider range of ERs while keeping computational complexity relatively low. When a wider range of ERs is included in the examination, a sequential classifier based on specific features related to character regions can be introduced. This classifier can be pre-trained to generate the probability that a character is present, resulting in multiple possible detected boundaries of the character (i.e., character boundaries 204). The first stage of ER classification estimates the probability that a character is present, and the second stage selects the ER with the locally highest probability. Classification can be further improved by using several more computationally intensive features. Furthermore, in at least one embodiment, an iterative and exhaustive search is applied to detect character combinations and group ERs into words. Such methods can also include region edges in the ER considerations to improve character detection. The final result is the ER selected with the highest probability representing the character boundary 204.

[0126] Because the character detector described above generates multiple regions for the same character, the next step is to perform disambiguation on the detected regions. In at least one embodiment, this disambiguation involves performing multiple comparisons of the detected character boundaries 204 and then purging character boundaries 204, which may be in the form of character bounding boxes that appear too close to each other. As a result, only one character bounding box is accepted within a particular perimeter, thus enabling the correct formation of the character string 200 representing the appropriate text field of the card image 310.

[0127] 5 is a flowchart depicting a method 500 for carrying out the process outlined above in more detail. A video frame 300 is selected for processing, or an option is selected to process each video frame 300 in succession. In step 510, if a card image 310 is detected within the video frame 300, it is extracted and resized to a standardized size. Next, in step 520, the resized card image is preprocessed with a series of filters including, for example, contrast enhancement, bilateral and median filtering for noise reduction, gamma correction, and / or illumination compensation.

[0128] In step 530, an ER filter with a two-stage classifier is created, and in step 540, this cascade classifier is applied to each image channel of the card image 310. Character groups are detected, and groups of one or more word boxes are extracted for further processing. In step 550, the string 200 with individual character boundaries 204 is analyzed for character boundary disambiguation. Finally, a clean string 200 is generated, accepting only one character within each perimeter of the character position 206.

[0129] 6 is a flowchart depicting a method 600 of further processing for verification of character boundaries 204. Method 600 may begin in step 610 with extraction of character string 200, removal of duplicate characters, and final processing and acceptance of character string 200. As depicted, each character in the disambiguated character string may be further processed for character image verification.

[0130] Thus, in step 620, a ratio of pixel counts can be obtained in low-intensity and high-intensity regions of each character image 202 (or processed character image 203) for comparison with a predetermined contrast ratio between low-intensity and high-intensity pixel counts. In step 620, for each character image 202 or processed character image 203, pixels at high-intensity and low-intensity levels are grouped and counted.

[0131] Next, in step 630, the ratio of these two counts is calculated and thresholded so that only those character images 202 or processed character images 203 with a sufficiently high contrast ratio are retained. Thereafter, in step 640, the location bounding box coordinates (i.e., location 206) of the verified character are recorded and saved for further use in interpreting the character string 200.

[0132] In alternative embodiments, the character bounding box verification described above may precede character boundary disambiguation, or verification may be used in combination with character boundary disambiguation for final character verification.

[0133] Information Card Processing for Query Extraction and Recognition In at least one embodiment, an automated process is implemented that includes the following steps: Receive a digital video stream, such as video stream 222 associated with highlight 220. Analyze one or more video frames 300 of the digital video stream for the presence of card images 310. Extract card images 310. Locate character boundaries 204 of characters of character string 200 within card images 310. Extract text that is within text boxes to create a query string of characters.

[0134] 7 is a flowchart depicting a method 700 of information card query generation according to one embodiment. In step 710, a card image 310 is extracted from a decoded video frame 300. In step 720, the card image 310 is processed to identify and extract character strings 200 as described above. In step 730, character images 202 are extracted from the card image 310, and a normalized query image (e.g., query 230) is generated. In step 740, the query infrastructure is input with the normalized query character image (query feature vector 232).

[0135] In another embodiment, a query prediction is generated by first projecting the query feature vector onto a pre-developed training set orthogonal basis (e.g., orthogonal basis 234) and then applying the resulting projected query feature vector to a machine-learned classification model, such as classification model 165. A predicted string of alphanumeric characters may be generated according to pre-established classification labels, and this predicted alphanumeric string may be passed to an interpretation process for ultimate extraction of the meaning 238 of the text.

[0136] 8 is a flowchart depicting a method 800 including processing steps for query recognition leading to query alphanumeric string generation and query interpretation and understanding. In step 810, orthogonal basis vectors of an orthogonal basis 234 are loaded across the training set vector space. In step 820, the normalized query may be projected onto the orthogonal basis 234. In step 830, a classification model 165, such as one previously developed, may be loaded. The classification model 165 may be applied to the projected query. Finally, in step 840, a predicted alphanumeric character string may be generated, which is then used for interpretation and semantic extraction to generate the text meaning 238.

[0137] Query interpretation and semantic extraction In at least one embodiment, one or more character strings 200 present in the card image 310 are identified. Subsequent steps may include locating, sizing, and extracting each character image 202 in the identified character strings 200. The detected and extracted character images 202 are converted into query feature vectors 232 and projected onto a training set orthogonal basis 234. The projected query is then applied against a classification model 165 to produce a predicted string of alphanumeric characters.

[0138] In at least one embodiment, the predicted query alphanumeric characters are sent to an interpretation process that applies prior knowledge and positional understanding of characters present in multiple card images 310. Meaning is then derived for each predicted alphanumeric character located in a particular string 200, and the extracted information is added to metadata 224 stored in association with the highlight 220.

[0139] 9 is a flowchart illustrating in more detail a method 900 for predicted query string interpretation according to one embodiment. The method 900 includes combining consideration of string length, character box position and horizontal distance, and alphanumeric readings for semantic extraction.

[0140] The method 900 begins at step 910, where the character count of each processed query for the string 200 is loaded along with the size 205 and position 206 of the character within the string 200. The video frame number and / or time associated with the extracted query 230 being processed may also be made available for reference relative to absolute time. At step 920, the string length 209, the size 205 of the character, and / or the position 206 of the character may be considered in the analysis.

[0141] Next, in step 930, system 100 steps through string 200, which may be interpreted by applying knowledge of the characters' field positions as well as knowledge of the characters' alphanumeric values. In step 930, knowledge and understanding of the particular card image 310 may also be used to compensate for any leading characters that may be missing. Finally, in step 940, the derived meaning is recorded (e.g., text meaning 238) and corresponding metadata 224 is formed, providing real-time information related to the current sporting event television program and the current timeline associated with the processed embedded card image 310.

[0142] Generating machine-learned classification models with application to recognizing query characters extracted from embedded information cards In at least one embodiment, the generation of the classification model is performed using a convolutional neural network. Typically, a neural network develops its information classification capabilities using known (desired) classification results through a supervised learning process applied to a training set of character vectors. During the training process, the neural network's algorithmic structure adjusts its weights and biases to perform accurate classification. One example of a known architecture used to learn the neural network's internal weights and biases during the training process is the backpropagation neural network architecture, or the feedforward backpropagation neural network architecture. When such a network is presented with a set of training data, the backpropagation algorithm calculates the difference between the actual output and the desired output and feeds back the error to correct the internal network weights and biases that cause the error. During the classification / inference phase, the neural network structure is first loaded with pre-trained model parameters, weights, and biases, and then a query is fed forward through the network, resulting in the network output of one or more identified label(s) representing the query's prediction.

[0143] Another exemplary system for generating classification models uses multiclass SVMs. Such SVM classification systems are fundamentally different from comparable approaches, such as neural network learning systems, which rely heavily on heuristics to construct various network architectures, and the training process does not always end at a global minimum. In contrast, SVMs are mathematically very well-defined, with a training process that consistently finds a global minimum. Furthermore, SVMs offer a relatively simple and clear geometric interpretation of the training process and classification objective, which improves intuitive insight into the process of generating classification models. SVMs can be efficiently used to classify linearly non-separable datasets and can be extended to multi-label classification tasks. SVMs for the classification of linearly non-separable datasets are characterized by the selection of a kernel function that helps project the dataset into a high-dimensional vector space in which the original dataset becomes linearly separable. However, the selection of the kernel function is important and involves a degree of heuristics and data dependence.

[0144] In at least one embodiment, character classification model generation is based on a training set of characters extracted from one or more exemplary card images 310 embedded in television program content of a sporting event. Character boundaries 204 are detected and characters are extracted from a number of card images 310. Such character boundaries 204 include small character images 202 that can then be normalized to a standard size and lighting to provide processed character images 203. Feature vectors (or query feature vectors 232) are formed for the character images 202 and / or the processed character images 203, and these feature vectors are then associated with each particular character from the set of character images that appear in the embedded card images 310.

[0145] In a structural approach to character image feature formation, a character feature vector, or query feature vector 232, is associated with a set of n pixels extracted from preprocessed character images 202. These n pixels are formatted into an n-dimensional vector that represents a single point in the n-dimensional feature vector space of training vectors. The primary goal of feature selection is to construct a decision boundary in the feature space that properly separates different classes of character images 202. Thus, in at least one embodiment, the set of extracted character images 202, representing the training vector, is further processed to increase the uniqueness and mutual metric distance of the training vectors, as well as to reduce the dimensionality of the overall vector space of the training vectors.

[0146] In accordance with the above considerations, in another embodiment, principal component analysis (PCA) is performed on the training vector set. Thus, orthogonal basis vectors of the orthogonal basis 234 are derived from the training set such that the orthogonal basis vectors span the training vector space. Furthermore, the dimensionality of the training vector space is reduced by selecting a limited number of orthogonal basis vectors such that only the most significant orthogonal vectors associated with the largest set of singular values ​​(generated by singular value decomposition of the matrix of basis vectors) are retained. The selected training set basis vectors are saved for later use in generating a classification model using one or more algorithmic structures available for dataset classification, such as an SVM classifier or a CNN classifier.

[0147] In various embodiments, the systems and methods described herein provide techniques for extracting individual character images 202 from character strings 200 embedded in a card image 310 and then utilizing the character images 202 to generate query feature vectors 232. In a next processing step, these query feature vectors are projected onto an orthogonal basis 234 spanning the training vector space to generate projected queries. The projected queries are then applied to generate query predictions or predicted queries 236 as the output of a pre-trained classification model produced by an exemplary SVM (or CNN) classifier. These predicted queries 236 form predicted character strings, which are then interpreted to generate text meaning 238 and ultimately used to generate metadata 224 for highlights 220 enriched with real-time information read directly from the card image 310.

[0148] FIG. 10 is a flowchart depicting a method 1000 of classification model generation in more detail. In at least one embodiment, method 1000 begins at step 1010, where an exemplary training set of character images 202 is extracted from a number of exemplary card image types. The character images 202 are normalized to a standard size and illumination to form processed character images 203. Feature vectors are derived and a labeled training set is generated. In at least one embodiment, in step 1020, PCA analysis is performed on the training set by computing an orthogonal basis 234 spanning the training vector space. In step 1030, a subset of the orthogonal training vectors is selected. The selected training set basis vectors may be saved for query processing in step 1040. In step 1050, a classification model 165 may be trained on the subset of orthogonal training vectors. The classification model and orthogonal basis vectors may be saved for future generation of predicted queries 236 in step 1060.

[0149] 11 is a flowchart depicting an overall method 1100 for reading and interpreting text fields in a card image 310 and updating the metadata 224 of a highlight 220 with in-frame real-time information. In step 1110, the fields to be processed are selected from the character boundaries 204 of characters present in the card image 310. In step 1120, groups of characters are extracted from the line fields and text strings are recognized and interpreted as described above. Finally, in step 1130, the reading of the card image performed at the decoded video frame boundaries is embedded in the metadata 224 generated for the highlight 220.

[0150] The present system and method have been described in particular detail with respect to the envisioned embodiment. Those skilled in the art will appreciate that the system and method may be implemented in other embodiments. First, the specific naming of components, terminology capitalization, attributes, data structures, or any other programming or structural aspect is not required or important, and mechanisms and / or functions may differ in name, format, or protocol. Furthermore, the system may be implemented via a combination of hardware and software, entirely within hardware elements, or entirely within software elements. Also, the specific division of functionality among various system components described herein is merely exemplary and not required. Functionality performed by a single system component may instead be performed by multiple components, and functionality performed by multiple components may instead be performed by a single component.

[0151] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. The appearances of the phrases "in one embodiment" or "in at least one embodiment" in various places in the specification do not necessarily all refer to the same embodiment.

[0152] Various embodiments may include any number of systems and / or methods for implementing the above-described techniques, either alone or in any combination. Another embodiment includes a computer program product including a non-transitory computer-readable storage medium and computer program code encoded on the medium for causing a processor within a computing device or other electronic device to implement the above-described techniques.

[0153] Some portions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computing device's memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is herein generally conceived to be a self-consistent sequence of steps (instructions) leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It is sometimes convenient, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. Further, without loss of generality, it is also convenient to refer to specific arrangements of steps requiring physical manipulations of physical quantities as modules or code devices.

[0154] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, and as will be apparent from the description that follows, throughout this specification, descriptions utilizing terms such as "processing" or "computing" or "calculating" or "displaying" or "determining" will be understood to refer to operations and processes of a computer system or similar electronic computing module and / or device, and to mean manipulating and transforming data that are represented as physical (electronic) quantities in the computer system's memory or registers or other such storage, transmission, or display devices.

[0155] Certain aspects include process steps and instructions described herein in the form of an algorithm. It should be noted that the process steps and instructions may be embodied in software, firmware, and / or hardware, and that if embodied in software, may be downloaded to reside on and operate from a variety of platforms for use by a variety of operating systems.

[0156] This document also relates to apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, or may include a general-purpose computing device selectively activated or reconfigured by a computer program stored on the computing device. Such a computer program may be stored on a computer-readable storage medium, such as a floppy disk, optical disk, CD-ROM, DVD-ROM, magneto-optical disk, read-only memory (ROM), random-access memory (RAM), EPROM, EEPROM, flash memory, solid-state drive, magnetic or optical card, application-specific integrated circuit (ASIC), or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus. The program and its associated data may also be hosted and executed remotely, such as on a server. Furthermore, the computing devices referred to herein may include a single processor or may be architectures employing multiple processor designs to increase computing power.

[0157] The algorithms and displays presented herein are not inherently related to any particular computing device, virtualization system, or other apparatus. Various general-purpose systems may also be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the description provided herein. Moreover, the systems and methods are not described with reference to any particular programming language. It will be understood that a variety of programming languages ​​can be used to implement the teachings described herein, and any references above to specific languages ​​are provided for the purpose of enabling and best mode disclosure.

[0158] Accordingly, various embodiments include software, hardware, and / or other elements, or any combination or plurality of elements, for controlling a computer system, computing device, or other electronic device. Such electronic devices may include, for example, a processor, input devices such as a keyboard, mouse, touchpad, trackpad, joystick, trackball, microphone, and / or any combination thereof, output devices such as a screen and / or speaker, long-term storage such as memory, magnetic storage, and / or optical storage, and / or network connectivity. Such electronic devices may be portable or non-portable. Examples of electronic devices that can be used to implement the described systems and methods include desktop computers, laptop computers, televisions, smartphones, tablets, music players, audio devices, kiosks, set-top boxes, gaming systems, wearable devices, home electronic devices, and / or server computers, etc. The electronic device may use any operating system such as, but not limited to, Linux, Microsoft Windows available from Microsoft Corporation, Redmond, Washington, Mac OS X available from Apple Inc., Cupertino, California, iOS available from Apple Inc., Cupertino, California, Android available from Google Inc., Mountain View, California, and / or any other operating system adapted for use on the device.

[0159] While a limited number of embodiments have been described herein, those skilled in the art, having the benefit of the above description, will appreciate that other embodiments may be devised. Furthermore, it should be noted that the language used herein has been chosen primarily for ease of reading and educational purposes, and may not have been chosen to delineate or limit the subject matter. Accordingly, the present disclosure is intended to be illustrative, but not limiting, in scope.

Claims

1. 1. A method for extracting metadata from a video stream, said method comprising: receiving, in a processor, at least one portion of a video stream; identifying, in the processor, one or more card images embedded in one or more video frames of the portion of the video stream; processing, in the processor, the one or more card images to extract text; interpreting the text in the processor to obtain metadata; storing the metadata in association with the portion of the video stream in a data store.

2. The method of claim 1 , further comprising storing the received portion of the video stream in the data store.

3. the video stream comprises a television broadcast of a sporting event; the portion of the video stream includes highlights deemed to be of particular interest to one or more users; The method of claim 1 , wherein the metadata describes the highlight.

4. 4. The method of claim 3, further comprising outputting the video stream at an output device simultaneously with at least one of identifying the one or more card images, processing the one or more card images, and interpreting the text.

5. outputting the highlights at an output device; outputting the metadata simultaneously with outputting the highlights; The metadata is real-time information related to said highlights; and The method of claim 3 , wherein the metadata includes at least one selected from the group consisting of: a timeline of when the card image was acquired;

6. Extracting the text identifying one or more characters of text within the one or more card images; and recording the position and / or size of character images of the card images having the one or more card images corresponding to each character of the one or more character strings.

7. Extracting the text disambiguating character boundaries of characters of the one or more character strings by performing multiple comparisons of the detected character boundaries; 7. The method of claim 6, further comprising: purging any character boundaries that occur too close together.

8. 7. The method of claim 6, wherein extracting the text further comprises performing image verification on characters of one or more character strings by establishing a contrast ratio between low intensity pixel counts and high intensity pixel counts.

9. Interpreting said text generating a query based on the text; generating a plurality of n-dimensional query feature vectors; projecting the n-dimensional query feature vector onto a training set orthogonal basis; applying the projected n-dimensional query feature vector to a classification model to generate at least one predicted query; and extracting the textual meaning from the at least one predicted query.

10. generating a plurality of training set feature vectors; The method of claim 9 , further comprising: using the training set feature vectors to derive the training set orthogonal basis.

11. generating a plurality of training set feature vectors; The method of claim 9 , further comprising: generating the classification model using the training set feature vectors.

12. Interpreting said text the sequence length of one or more strings within said text; character boundaries and / or character positions within said text; The method of claim 9 , further comprising using at least two selections from the group consisting of character boundaries and / or horizontal positions of characters within the text.

13. 10. The method of claim 9, wherein storing the metadata in association with the portion of the video stream comprises storing a video frame number of the one or more video frames associated with a query.

14. Interpreting said text determining field positions of characters of one or more strings of said text; determining the alphanumeric value of said character; and sequentially interpreting the one or more character strings using the field position and alphanumeric value.

15. Interpreting said text obtaining positional and other information regarding one or more card fields of each of said card images; 15. The method of claim 14, further comprising: using the positional information and other information to compensate for one or more possibly missing leading characters of the one or more character strings.

16. 1. A method for generating a classification model for extracting metadata from a video stream, the method comprising: receiving, in a processor, at least one portion of the video stream; identifying, in a processor, one or more card images embedded in one or more video frames of the portion of the video stream; processing, in the processor, the one or more card images to extract, for each instance where the card image includes a character, a plurality of character images; generating, in the processor, a training feature vector associated with the character image; processing the training feature vectors in the processor, wherein said processing comprises: Increasing the uniqueness of the training feature vectors; Increasing the mutual numerical distance of the training feature vectors; and / or processing, performed in a manner that reduces the dimensionality of the overall vector space containing the training feature vectors; training, in the processor, a classification model using at least some of the training feature vectors; storing the classification model in a data store.

17. The method of claim 16 , further comprising storing the received portion of the video stream in the data store.

18. The method of claim 16 , further comprising: in the processor, normalizing the character images to a standard size and / or standard lighting before generating the training feature vectors.

19. 17. The method of claim 16, wherein generating the training feature vector comprises formatting a set of n pixels extracted from the character image into an n-dimensional vector.

20. further comprising performing, in the processor, principal component analysis on the training feature vectors; training the classification model using at least some of the training feature vectors; selecting a subset of the training feature vectors that are orthogonal basis vectors; and training the classification model using the orthogonal basis vectors.

21. the orthogonal basis vectors span the global vector space; reducing the dimensionality of the overall vector space includes selecting a limited number of the orthogonal basis vectors; reducing the dimensionality of the overall vector space further comprises selecting only orthogonal basis vectors corresponding to a maximal set of singular values ​​derived from the matrix of orthogonal basis vectors; storing the classification model includes storing a limited number of orthogonal basis vectors for subsequent use in classification model generation and / or query processing; and / or 21. The method of claim 20, wherein generating the classification model comprises using the limited number of orthogonal basis vectors in combination with a machine learning algorithm selected from the group consisting of an SVM and a CNN.

22. The method comprises: processing, in the processor, the one or more card images to extract text; interpreting the text in the processor to obtain metadata; storing, in the data store, the metadata associated with the portion of the video stream; outputting the portion of the video stream at an output device; In the output device, outputting the metadata simultaneously with outputting the portion; the video stream comprises a broadcast of a sporting event; the portion of the video stream includes highlights deemed to be of particular interest to one or more users; The method of claim 16 , wherein the metadata describes the highlight.

23. 23. The method of claim 22, wherein extracting the text comprises extracting a text string of the text as a query.

24. Extracting the text the current time within said sporting event; the current phase of said sporting event; a game clock associated with said sporting event; and 23. The method of claim 22, comprising extracting at least one of a match score associated with the sporting event.

25. 1. A non-transitory computer-readable medium for extracting metadata from a video stream, the medium comprising instructions stored therein, the instructions, when executed by a processor, performing: receiving at least a portion of the video stream; identifying one or more card images embedded in one or more video frames of the portion of the video stream; processing the one or more card images to extract text; interpreting the text to obtain metadata; A non-transitory computer-readable medium performing the step of storing the metadata in association with the portion of the video stream in a data store.

26. the video stream comprises a television broadcast of a sporting event; the portion of the video stream includes highlights deemed to be of particular interest to one or more users; 26. The non-transitory computer-readable medium of claim 25, wherein the metadata describes the highlight.

27. 27. The non-transitory computer-readable medium of claim 26, further comprising instructions stored therein that, when executed by a processor, cause an output device to output the video stream simultaneously with at least one of identifying the one or more card images, processing the one or more card images, and interpreting the text.

28. and further including instructions stored therein, said instructions, when executed by the processor, causing an output device to output said highlight; outputting the metadata simultaneously with outputting the highlights; The metadata is real-time information related to said highlights; and 27. The non-transitory computer-readable medium of claim 26, wherein the metadata includes at least one selected from the group consisting of: a timeline of when the card image was acquired.

29. Extracting the text identifying one or more characters of text within the one or more card images; and recording the position and / or size of character images of the card images having the one or more card images corresponding to each character of the one or more character strings.

30. Interpreting said text generating a query based on the text; generating a plurality of n-dimensional query feature vectors; projecting the n-dimensional query feature vector onto a training set orthogonal basis; applying the projected n-dimensional query feature vector to a classification model to generate at least one predicted query; and extracting the textual meaning from the at least one predicted query.

31. and further including instructions stored therein, said instructions, when executed by the processor, generating a plurality of training set feature vectors; and 31. The non-transitory computer-readable medium of claim 30, wherein the training set feature vectors are used to derive the training set orthogonal basis and / or generate the classification model.

32. Interpreting said text determining field positions of characters of one or more strings of said text; determining the alphanumeric value of said character; and sequentially interpreting the one or more character strings using the field position and alphanumeric value.

33. 1. A non-transitory computer-readable medium for generating a classification model to extract metadata from a video stream, the medium comprising instructions stored therein, the instructions, when executed by a processor, performing: receiving at least a portion of the video stream; identifying one or more card images embedded in one or more video frames of the portion of the video stream; processing the one or more card images to extract, for each instance where the card image includes text, a plurality of text images; generating a training feature vector associated with the character image; processing the training feature vectors, said processing comprising: Increasing the uniqueness of the training feature vectors; Increasing the mutual numerical distance of the training feature vectors; and / or processing, performed in a manner that reduces the dimensionality of the overall vector space containing the training feature vectors; training a classification model using at least some of the training feature vectors; and storing the classification model in a data store.

34. further comprising instructions stored therein that, when executed by the processor, perform principal component analysis on the training feature vectors; training the classification model using at least some of the training feature vectors; selecting a subset of the training feature vectors that are orthogonal basis vectors; and training the classification model using the orthogonal basis vectors.

35. the orthogonal basis vectors span the global vector space; reducing the dimensionality of the overall vector space includes selecting a limited number of the orthogonal basis vectors; reducing the dimensionality of the overall vector space further comprises selecting only orthogonal basis vectors corresponding to a maximal set of singular values ​​derived from the matrix of orthogonal basis vectors; storing the classification model includes storing a limited number of orthogonal basis vectors for subsequent use in classification model generation and / or query processing; and / or 35. The non-transitory computer-readable medium of claim 34, wherein generating the classification model comprises using the limited number of orthogonal basis vectors in combination with a machine learning algorithm selected from the group consisting of an SVM and a CNN.

36. and further including instructions stored therein, said instructions, when executed by the processor, processing the one or more card images to extract text; interpreting the text to obtain metadata; storing the metadata in association with the portion of the video stream in the data store; causing an output device to output said portion of said video stream; and causing the output device to output the metadata simultaneously with outputting the portion of the video stream; the video stream comprises a broadcast of a sporting event; the portion of the video stream includes highlights deemed to be of particular interest to one or more users; 34. The non-transitory computer-readable medium of claim 33, wherein the metadata describes the highlight.

37. 1. A system for extracting metadata from a video stream, the system comprising:

1. A processor, comprising: receiving at least a portion of the video stream; identifying one or more card images embedded in one or more video frames of the portion of the video stream; processing the one or more card images to extract text; a processor configured to interpret the text to obtain metadata; a data store configured to store the metadata in association with the portion of the video stream.

38. the video stream comprises a television broadcast of a sporting event; the portion of the video stream includes highlights deemed to be of particular interest to one or more users; 38. The system of claim 37, wherein the metadata describes the highlight.

39. 39. The system of claim 38, further comprising an output device configured to output the video stream simultaneously with at least one of identifying the one or more card images, processing the one or more card images, and interpreting the text.

40. further comprising an output device configured to output the highlights; the processor is further configured to output the metadata simultaneously with outputting the highlights; The metadata is real-time information related to said highlights; and 40. The system of claim 38, wherein the metadata includes at least one selected from the group consisting of: a timeline of the card image captured.

41. the processor: identifying one or more characters of text within the one or more card images; 38. The system of claim 37, further configured to extract the text by: recording the position and / or size of character images of card images having the one or more card images corresponding to each character of the one or more character strings.

42. the processor: generating a query based on the text; generating a plurality of n-dimensional query feature vectors; projecting the n-dimensional query feature vector onto a training set orthogonal basis; applying the projected n-dimensional query feature vector to a classification model to generate at least one predicted query; 38. The system of claim 37, further configured to interpret the text by: extracting the text meaning from the at least one predicted query.

43. the processor: generating a plurality of training set feature vectors; and 43. The system of claim 42, further configured to use the training set feature vectors to derive the training set orthogonal basis and / or generate a classification model.

44. the processor: determining field positions of characters of one or more strings of said text; determining the alphanumeric value of said character; 38. The system of claim 37, further configured to interpret the text by: sequentially interpreting the one or more character strings using the field position and alphanumeric value.

45. 1. A system for generating a classification model for extracting metadata from a video stream, the system comprising:

1. A processor, comprising: receiving at least a portion of the video stream; identifying one or more card images embedded in one or more video frames of the portion of the video stream; processing the one or more card images to extract, for each instance where the card image includes text, a plurality of text images; generating a training feature vector associated with the character image; processing the training feature vectors, said processing comprising: Increasing the uniqueness of the training feature vectors; Increasing the mutual numerical distance of the training feature vectors; and / or processing, performed in a manner that reduces the dimensionality of the overall vector space containing the training feature vectors; training a classification model using at least some of the training feature vectors; and a data store configured to store the classification model.

46. the processor: performing principal component analysis on the training feature vectors; Using at least some of the training feature vectors, the classification model is selecting a subset of the training feature vectors that are orthogonal basis vectors; and training the classification model using the orthogonal basis vectors.

47. the orthogonal basis vectors span the global vector space; the processor is further configured to reduce the dimensionality of the overall vector space by selecting a limited number of the orthogonal basis vectors; the processor is further configured to reduce the dimensionality of the overall vector space by selecting only orthogonal basis vectors corresponding to a maximal set of singular values ​​derived from the matrix of orthogonal basis vectors; the data store is further configured to store the classification model by storing a limited number of orthogonal basis vectors for subsequent use in classification model generation and / or query processing; and / or 47. The system of claim 46, wherein the processor is further configured to generate the classification model by using the limited number of orthogonal basis vectors in combination with a machine learning algorithm selected from the group consisting of SVM and CNN.

48. the processor: processing the one or more card images to extract text; further configured to interpret the text to obtain metadata; the data store is further configured to store the metadata in association with the portion of the video stream; The system is an output device, outputting the portion of the video stream; and an output device configured to output the metadata simultaneously with outputting the portion of the video stream; the video stream comprises a broadcast of a sporting event; the portion of the video stream includes highlights deemed to be of particular interest to one or more users; 46. ​​The system of claim 45, wherein the metadata describes the highlight.

Citation Information

Patent Citations

  • Video attribute information output apparatus, video summarizing device, program, and method for outputting video attribute information

    JP2008176538A

  • Term meaning code determination device, method and program

    JP2017021523A

  • Text detection in video

    JP2017522648A

  • Customized generation of highlight shows with narrative elements

    JP2017538989A

  • Metadata generation system

    JP2018033048A