Identifying and Retrieving Video Metadata Using Perceptual Frame Hashing
By employing perceptual hashing to generate hash vectors for frames of shopable videos, the technology efficiently matches viewer requests with database information across various video formats, enhancing the accessibility of product information in shopable videos.
Patent Information
- Application Number
- JP2021576733
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-03
- Filing Date
- 2020-07-02
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2040-07-02
AI Technical Summary
The challenge in shopable video technology is to efficiently match viewer requests for product information in video frames with corresponding database information, particularly when the same video is displayed in various formats, making it impractical to tag and store information for each format.
The use of perceptual hashing to identify frames in the source video by generating hash vectors for each frame of different versions of the video, which are then associated with video information in a database, allowing for efficient matching and retrieval of metadata.
This approach reduces the time and friction between identifying elements in a video and displaying related information to the viewer, enabling immediate access to products and other information shown in the video.
Smart Images

Figure 0007672348000001 
Figure 0007672348000002 
Figure 0007672348000003
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 U.S.C. §119(e) of U.S. patent application Ser. No. 62 / 870,127, filed July 3, 2019, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Shoppable video allows viewers watching a video to shop for items such as fashion, accessories, household goods, technology devices, and even menus and recipes that appear in the video. As the viewer watches the video, they find the item they want to purchase displayed in the video. The viewer can obtain information about that item, such as price and availability information, by pressing a button on the remote control or speaking into the microphone on the remote control. A processor in or coupled to the television receives this request and transmits it to a server, which retrieves information about the item from a database and returns it to the processor. The television displays the information about the item to the viewer, who can then purchase the item or request information about similar products.
[0003] Shoppable videos are typically tagged before being displayed on a television, either manually or using machine learning techniques to recognize products in each video frame. Product metadata for tagged products is matched to corresponding video frames and stored in a database. When a viewer requests product metadata, a processor identifies the corresponding video frames and then retrieves the product metadata for those video frames. Summary of the Invention [Problem to be solved by the invention]
[0004] One of the challenges of shoppable video is matching a viewer's request for information about a product in a video frame with information in a database. The same shoppable video may be displayed in one of many different formats, complicating the ability to match a displayed video frame with a corresponding video frame. One reason is that the number of possible formats grows over time, making it impractical to tag each possible format and store information for each corresponding frame. [Means for solving the problem]
[0005] The present technique addresses this challenge by identifying frames of a source video using perceptual hashing. In one embodiment of the method, a processor generates a hash vector for each frame of different versions of the source video. These hash vectors are associated with information about the source video in a database. When a playback device, such as a smart TV, a TV with a set-top box, a computer, or a mobile device, plays a first version of the source video, a first hash vector is generated for a first frame of the first version of the source video. This first hash vector is matched to a matching hash vector among the hash vectors in the database, for example, using an application programming interface (API) to query the database. In response to matching the first hash vector to the matching hash vector, information about the source video is retrieved from the database.
[0006] Determining that the first hash vector matches the matching hash vector can include determining that the first hash vector is within a threshold distance of the matching hash vector. The matching hash vector can be for a frame of a second version of the source video that is different from the first version of the source video.
[0007] The hash vector and the first hash vector may be generated using a perceptual hashing process, such as perception hashing (pHash), difference hashing (dHash), average hashing (aHash), and wavelet hashing (wHash). The generation of the first hash vector may take approximately 100 milliseconds or less. The first hash vector may have a size of 4096 bits or less. The generation of the first hash vector may occur automatically at regular intervals and / or in response to a command from the viewer.
[0008] If desired, hash vectors can be sharded and stored in a sharded database. In other words, the hash vectors can be separated or divided into subsets, with each subset stored in a different shard of the database. The hash vectors can be divided into subsets randomly or based on how frequently or recently the hash vectors are accessed, the distance between the hash vectors, and / or characteristics of the hash vectors.
[0009] The techniques can also be used to identify and retrieve metadata associated with a source video. Again, a processor generates a hash vector for each frame of at least one version of the source video and stores the hash vector in a first database. A second database stores metadata corresponding to each frame. (This metadata can be updated without changing the hash vector in the first database.) A playback device plays the first version of the source video. The playback device or an associated processor generates a first hash vector for a first frame of the first version of the source video. The API server matches the first hash vector to a matching hash vector among the hash vectors in the first database. In response to matching the first hash vector to the matching hash vector, the API retrieves metadata corresponding to the matching hash vector from the second database, and the playback device displays the metadata to a viewer.
[0010] The metadata may represent at least one of a location within the source video, clothing worn by an actor in the source video, a product appearing in the source video, or music playing the source video. The hash vectors may be associated with the metadata by their respective timestamps.
[0011] Matching the first hash vector to the matching hash vector can include sending the first hash vector to an API server. The API server determines that the first hash vector matches the matching hash vector and then identifies a timestamp associated with the matching hash vector in the first database. In this case, retrieving the metadata further includes querying a second database based on the timestamp and retrieving the metadata associated with the timestamp from the second database. Determining that the first hash vector matches the matching hash vector can include the first hash vector being within a threshold distance of the matching hash vector.
[0012] The method for identifying, retrieving, and displaying metadata associated with a video can also include playing the video via a display, generating a first hash vector for a first frame of the video, transmitting the first hash vector to an API server, and retrieving, via the API server, metadata associated with the first frame from a metadata database. The metadata is retrieved from the first database in response to matching the first hash vector to a second hash vector stored in the hash vector database. The display displays the metadata associated with the first frame to a user.
[0013] From another perspective, a database receives a first hash vector generated for a first frame of the video. The database stores the first hash vector and receives a query from a playback device based on the second hash vector. The database performs a query on the second hash vector and, in response to matching the second hash vector to the first hash vector, sends a timestamp associated with the first hash vector to the API. The timestamp associates metadata with the first frame of the video.
[0014] In another embodiment, a processor generates a first hash vector for a first frame of a source video. A first database stores the first hash vector in the first database. A playback device plays one version of the source video. The same processor or another processor generates a second hash vector for a second frame of the version of the source video. The second hash vector is matched to the first hash vector in the first database. In response to matching the second hash vector to the first hash vector, a timestamp corresponding to the second hash vector may be retrieved and sent to the playback device.
[0015] All combinations of the above concepts and additional concepts (provided that such concepts are not mutually inconsistent) are discussed in more detail below and are part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter presented in the appended claims are part of the inventive subject matter disclosed herein. The terms used herein, which may also appear in any disclosure that is incorporated herein by reference, should be given the meaning most consistent with the specific concepts disclosed herein.
[0016] Those skilled in the art will appreciate that the drawings are presented primarily for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale, and in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally and / or structurally similar elements). [Brief description of the drawings]
[0017] [Figure 1] FIG. 1 shows a system that allows instant access to elements in a video. [Diagram 2] FIG. 2 is a flow diagram illustrating a process for generating and storing hash vectors for different versions of a source video. [Diagram 3] FIG. 3 is a flow diagram showing a process for generating and storing metadata for a source video. [Figure 4] FIG. 4 is a flow diagram illustrating a process for retrieving metadata about objects in a video using perceptual frame hashing. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] The techniques disclosed herein help provide television viewers with instant access to products, places, and other information shown in a video. More specifically, the techniques disclosed herein reduce the time and / or friction between identifying elements (e.g., products, places, etc.) in a displayed video and displaying information about those elements to the viewer. The viewer can then save the information and / or purchase the products they like in the video.
[0019] The technique uses perceptual hashing to identify video frames shown to a viewer on a playback device. To obtain information about an item in a video frame, the playback device generates a perceptual hash of the frame image and sends the hash to a server, which queries a hash database that contains perceptual hashes and timestamps, or other identifiers, for frames from a variety of different videos in a variety of different formats. The query returns an identifier that can be used to query another database for information or metadata about the item in the frame, or for viewer surveys or other data collection operations. Alternatively, metadata can be stored in the same database as the hash and returned with the identifier. This metadata can include information about locations, clothing, products, music, or sports scores from the source video, or detailed information about the video itself (e.g., runtime, synopsis, cast, etc.). The server returns this information or metadata to the playback device, which can then be displayed to the user.
[0020] The use of perceptual hashes offers several advantages over other techniques for identifying video frames and objects in those video frames. First, transmitting perceptual hashes consumes less upstream bandwidth than transmitting video frames or other identifying information. Generating and matching perceptual hashes does not require specialized hardware. Perceptual hashes are highly robust to degradation of video quality, increasing the likelihood of a correct match across a wider range of video viewing and transmission conditions. It also protects the privacy of the viewer. In the absence of a hash database, someone eavesdropping on a perceptual hash would never know what content is being displayed on the playback device. And for any content that is owned only by the viewer (e.g., home movies), there is virtually no way (even with a hash database as described above) that anyone can identify what is being displayed on the playback device based on the hash value. This is because there is an almost infinite number of images that can be generated that result in the same hash value, making it virtually impossible to infer or reverse engineer the source image from the hash. (Technically, multiple frames can generate the same hash vector, but for a 512-bit hash vector, 512 Because there are 1 possible hash vectors (a 1 followed by 154 zeros), the compressed hash space is large enough to encode a nearly infinite number of frames without encoding different frames using the same hash vector.
[0021] Additionally, bifurcate the hash database and the metadata database (as opposed to item information associated with hashes in the same database), allowing metadata associated with a given frame to be updated without affecting the hash database. For example, metadata about a product displayed in a given piece of content can be updated frequently as its availability / price changes without requiring a change to the hash database.
[0022] 1 illustrates a system 100 that allows immediate access to elements in a source video 121 that can be displayed in one of several formats 125 on a playback device 110, such as a smart TV, a TV with a separate set-top box, a computer, a tablet, or a smartphone coupled to a content provider 120. This source video 121 can be provided by a content partner (e.g., the Hallmark Channel can provide the source video 121 for episodes prior to broadcast), a distribution partner (e.g., Comcast), downloaded from the Internet (e.g., from YouTube®), or captured from a live video feed (e.g., live capture of a feed of an NBA game). The system 100 includes a metadata database 150 communicatively coupled to a hash database 140 and an application programming interface (API) server 130, which is also communicatively coupled to the playback device 110. For example, the playback device 110, the content provider 120, the API server 130, the hash database 140, and the metadata database 150 may be located in the same or different geographic locations, operated by the same or different parties, and communicate with each other via the Internet or one or more other suitable communications networks.
[0023] (Depending on the content source, system 100 may perform steps in addition to those shown in FIGS. 1-4. For example, in the case of a set-top box, the content may be distributed from a content provider 120 (e.g., Disney) to a cable company (e.g., Comcast) that plays the content on a playback device.)
[0024] 2-4 illustrate a process for populating and retrieving information from the hash database 140 and the metadata database 150. FIG. 2 illustrates how hash vectors 145, also referred to as hash value vectors, hash values, or hashes, are populated into the hash database 140. First, the source video 121 is split (block 202) into a number of individual frames 123a-123c (collectively, frames 123). This splitting can be done at a constant frame rate (e.g., 12 frames per second (fps)) or can be guided by a metric that indicates how significantly the content of the video changes from frame to frame, with splitting occurring whenever the content of the video changes beyond a threshold, which can be based on a hash-matching threshold or can be selected empirically based on a desired level of hash-matching accuracy. This reduces the number of hashes to store in the database. Multiple versions 125 of each source frame 123 can be generated (block 204) with modifications made to aspect ratios (e.g., 21x9 or 4x3 instead of 16x9), color values, or other parameters, with the goal of replicating the various ways in which the source video 121 may be displayed on the playback device 110 after going through a broadcast transcoding system, etc.
[0025] Each version 125 of each source frame 123 is run through a perceptual hashing process by a hash generation processor 142 (block 206). This hash generation processor 130 converts each frame version 125 into a corresponding perceptually meaningful hash vector 145. The hash generation processor 142 may use one or more perceptual hashing processes to generate the hash vector 145, such as perceptual hashing (pHash), differential hashing (dHash), average hashing (aHash), or wavelet hashing (wHash). The hash vector 145 may be a fixed-size binary vector (an N×1 vector, where each element of the vector contains either a 1 or a 0) or a floating-point vector. Hash vectors 145 may be any of a variety of sizes, including, but not limited to, 128 bits, 256 bits, 512 bits, 1024 bits, 2048 bits, 4096 bits, or larger. Hash vectors 145 for different versions 125 of the same source frame 123 may be close or far apart depending on their visual similarity. For example, versions 125 that differ slightly in color may have hash vectors 145 that are close enough to match each other, whereas versions 125 with different aspect ratios (e.g., 4:3 vs. 16:9) may have hash vectors 145 that are so far apart that they do not match each other.
[0026] Considerations for selecting the perceptual hashing process and the size of the hash vector 145 include: (1) How quickly the hash can be calculated on inexpensive hardware of the playback device 110 (e.g., a processor in a smart TV or set-top box (STB)) (an example target time to calculate a hash vector is 100 milliseconds or less). (2) The size of the hash vector 145. A smaller hash vector 145 allows for less bandwidth consumption between the playback device 110, the API server 120, and the hash database 140, less memory requirements for the hash database 140, and faster search times. A dHash of size 16x16 has a 512-bit output. A larger hash vector 145 allows for more accurate matches, but consumes more bandwidth and has longer search times. (3) Probability of collisions (the possibility that two different images will generate the same hash vector). The computation speed and size of the hash vector should be weighed against the ability of the hashing process to generate precisely different hashes for two similar but different inputs. For example, running dHash with a size of 32x32 (as opposed to 16x16) will result in a hash vector that is 2048 bits in size, allowing for more accurate discrimination between frames (i.e., greater precision) at the cost of four times the memory storage space. In some use cases, this may be a worthwhile trade-off, while in others it may not.
[0027] The hash vectors 145 are stored (block 208) in a hash database 140, which is configured to enable rapid (e.g., <100 milliseconds) and high-throughput (e.g., thousands of searches per second) approximate nearest-neighbor searches of the hash vectors 145. Although FIG. 1 depicts this hash database 140 as a single entity, in practice the hash database 140 may include multiple shards, each of which contains a subset of the hash vectors 145. The hash vectors 145 may be randomly distributed across the shards, or may be purposefully distributed according to a particular scheme. For example, the hash database 140 may store similar (close) vectors 145 on the same shard. Or, the hash database 140 may store the most frequently or most recently accessed vectors 145 on the same shard. Alternatively, the hash database 140 can use characteristics of the hash to determine which shard to place it on (e.g., using a learning process such as Locality Sensitive Hashing or a neural network trained on a subset of the hashes).
[0028] Sharding allows each subset of vectors 145 to be searched simultaneously and the results aggregated, keeping search times low even when many hash vectors 145 are being searched. Shards can also be searched sequentially until a match is found, for example, if the shards are organized by access frequency, the shard storing the most frequently accessed vectors 145 is searched first, then the shard storing the second most frequently accessed vectors 145 is searched second, and so on. Additionally, all hashes from a given video can be treated as a group in this or other schemes, in that if search volume for one or more hashes from a given video increases, all of the hashes for that video can be promoted simultaneously to a first shard in anticipation of more searches for the remaining hashes for that video. (Testing has shown that using commercially available hardware and software, each database shard can process at least hundreds of millions of hash vectors 145, and the total system 100 can process billions of hash vectors 145.) If this system 100 were used for a live event (i.e., a live basketball game), the time from inserting a new hash 145 into the database 140 to when that hash vector 145 is indexed and available for search should be short (e.g., less than 5 seconds).
[0029] Each hash vector 145 is associated in the hash database 140 with information identifying the corresponding source video 101 as well as a timestamp and / or frame number / identifier of the corresponding frame 123 / 125. In some cases, different versions 125 of the same source frame 123 may have different absolute timestamps due to content or length edits. In these cases, each hash vector 145 may also be associated with a timestamp offset indicating the difference between the timestamp of the associated frame version 125 and the timestamp of the corresponding frame 123 of the source video 101. The hash database 140 may return the timestamp, timestamp offset, and source video information in response to a query of the hash vector to query the metadata database 150.
[0030] FIG. 3 illustrates how metadata for a source video 121 is generated. The source video 121 is split into a separate set of frames 127a-127c (collectively, frames 127) for metadata generation and tagging (block 302). Because the frame rate can be lower for metadata generation than for perceptual hashing, the set of frames 127 for metadata generation can be smaller than the set of frames 123 for hashing. The split among frames 127 for metadata generation is selected to generate relevant metadata for the source video 121 (e.g., information about on-screen actors / characters, filming locations, clothing worn by on-screen characters, and / or the like), while the split among frames 123 for perceptual hashing is selected to perform automatic content recognition (ACR) to identify the source video. For example, frames with high motion blur may not be valid for metadata generation if the motion blur is severe enough to make it difficult or impossible to identify items appearing on the screen, and can be excluded from metadata generation, but can still be useful for ACR because they are visually unique images. Because the segmentation makes different selections according to different criteria, the frames 127 selected for metadata generation may not directly match or align with the frames 123 used to generate the hash vector 125. As a result, the same source video 121 can produce frames 123, 127 with different numbers and / or different timestamps for hash generation and metadata generation.
[0031] The metadata generation processor 152 operates on frames 127 selected for metadata generation and stores metadata associated with timestamps or other identifiers of the corresponding frames 127 in the metadata database 150 (blocks 304 and 306). Metadata generation may be accomplished automatically, with optional user intervention, by a user, for example, using techniques disclosed in U.S. Patent Application Publication No. 2020 / 0134320 A1, entitled "Machine-Based Object Recognition of Video Content," which is incorporated herein by reference in its entirety. Automated frame ingestion and metadata generation makes metadata available for search in just a few seconds (e.g., 5 seconds or less), making the process suitable for tagging and searching live videos such as sports, performances, and news. Examples of metadata that may be generated by the metadata processor 152 may include information about which actors / characters are on screen, which clothing items are worn by those actors / characters, or the filming locations depicted on screen. This metadata generation can be done independently or in conjunction with the frame hashing shown in FIG.
[0032] If desired, some or all of the metadata for source video 101 can be updated after it has been entered into the metadata database. For example, if a product tagged in source video 101 is sold, is no longer available, or is available from another vendor, the corresponding entries in metadata database 140 can be updated to reflect the change. Entries in metadata database 140 can also be updated to include references to similar products available from other vendors. These updates can be made without modifying any of the entries in hash database 140. And, as long as the entries in metadata database 150 contain timestamps or other identifying information for frames 127, they can be matched to corresponding hashes in hash database 140.
[0033] In some cases, metadata database 150 stores metadata that is linked or associated only with the source video 101 and not with different versions of the source video. This metadata can be searched for different versions of the video using the timestamps and timestamp offsets described above. In other cases, metadata database 150 stores metadata that is linked or associated with different versions of the source video 101 (e.g., a theatrical release and a shorter version edited for television). In these cases, a query of the metadata database can identify and return the metadata associated with the corresponding version of the source video 101.
[0034] 4 illustrates how the playback device 110 and API server 130 query the hash database 140 and metadata database 150. The playback device 110 (e.g., a smart TV, set-top box, or other Internet-connected display) displays a potentially modified version of the source video 121 to the viewer. This modified version may be edited (e.g., in terms of content or length) or reformatted (e.g., letterboxed or cropped) and may also include commercial and other breaks.
[0035] When a viewer views an altered version of the source video, the video playback device captures the image displayed on its screen and generates a hash vector 115 from that image (block 404) using the same perceptual hashing process (e.g., pHash, dHash, aHash, or wHash) used to generate the hash vectors 145 stored in the hash database 140. The playback device 110 generates the hash vector 115 quickly to keep latency as low as possible, e.g., 100 milliseconds or less.
[0036] The playback device 110 can capture and hash images in response to a viewer request or command made by pressing a button on a remote control or speaking into a microphone on a remote control or other device. The playback device 110 can also, or alternatively, capture and hash frames at regular intervals (e.g., every N frames or one frame every 1-300 seconds) and use the most recently derived hash vector 115 to perform a search in response to a viewer request. (If the playback device 110 can sense commercials or other program interruptions, it can stop generating hash vectors 115 during commercials to reduce processing load and / or bandwidth consumption.) Or, the playback device 110 can instead use the automatically generated hashes 115 to automatically perform a search in the background and display the search results in response to a subsequent viewer request. The playback device 110 can also use the automatically retrieved results to prompt the viewer by displaying an on-screen notice that metadata is available for the currently displayed video.
[0037] The playback device 110 transmits one or more of these hash vectors 115, as well as, optionally, frame timestamps and information identifying the video content to identify people, items, and / or places in the images, to the API server 130 (block 406). The playback device 110 can transmit each hash vector 115 to the API server 130, or can transmit only a subset of the hash vectors 115 to the API server 130. For example, if the playback device 110 periodically calculates the hash vectors 115, it can transmit each hash vector 115 to the API server 130. This consumes more bandwidth, but may reduce latency by sending requests for information from the API server 130 and receiving responses to those requests without waiting for commands from the viewer. As a result, the playback device 110 can display responses to the viewer's requests for information about people, objects, or places displayed by the playback device without waiting for database queries, because those queries have already been performed.
[0038] Alternatively, or in addition, the playback device 110 can send the hash vector 115 to the API server 130 in response to a command from the viewer, whether the hash vector 115 was generated periodically or in response to a command from the viewer. This consumes less bandwidth and reduces the processing load on the API server 130, hash database 140, and metadata database 150 by reducing the number of database queries. However, waiting to query the API server 130 until a viewer requests the information can increase latency.
[0039] In some cases, the identity of the source video 121 indicated by the content of the playback device 110 may already be known (e.g., a smart TV or set-top box may know the identity of the program shown on the TV), and the system 100 may only be used to identify the exact timestamp of the corresponding frame of the source video 121. In these cases, the playback device 110 may transmit an identifier of the content in addition to the hash value 115. For example, the content identifier may be generated by an auxiliary ACR system (e.g., Gracenote) or may be pulled from electronic programming guide (EPG) information using the set-top box. The content identifier may then be used to limit the search space or filter out false-positive matches from the hash database 140 based on the specified content.
[0040] When the API server 130 receives the hash vector 115 from the playback device 110, it queries the hash database 140 for a matching stored hash vector 145 (block 408). Because hashing is one-way, it may not be possible to determine the exact source value (video frame) from which the hash was generated. However, because hash values 115 and 145 were generated using perceptual hashing (e.g., dHash), the location relationship / distance between the hash values is meaningful, since hashing of similar source images will result in similar hash values. This is in contrast to standard cryptographic hash algorithms such as SHA or MD5, which are designed to produce dramatically different hash values even with slight perturbations of the input.
[0041] If the search yields similar hash vectors 145 within a predefined tight threshold distance (block 410), the hash database 140 returns the timestamps of the matching frames to the API server 130. The distance threshold can be determined based on experimental data and an acceptable false positive rate for a given use case (higher thresholds tend to give higher true positive rates, but also higher false positive rates). For example, the system 100 can be tested and tuned using different thresholds to return known ground-truth timestamps. The distance between hash vectors can be calculated using one of a variety of distance metrics, such as, for example, L2 (Euclidean) distance or Hamming distance (if the vectors are binary). Alternatively, other notions of similarity, such as cosine similarity or cross-correlation, can be used to compare hashes. Additionally, the thresholds can be set differently for different videos or for different shards of the hash database.
[0042] If no match is found within the strict threshold distance, a less strict (loose) threshold can be used (block 412) and a consensus method is used to maintain accuracy for false positives. For example, if the three closest matches using the looser threshold are all from the same source video 101 and have timestamps within a few seconds of each other, this provides greater confidence that the closest match is correct even if it is outside the strict threshold. However, if the three closest matches are from different source videos 101 and / or have timestamps more than a few seconds apart, it may be safer to assume that there is no match. If no match is found, the API server 130 returns a null result to the playback device 110 (block 420).
[0043] If there is a matching hash in the hash database 140, the query of the hash database returns a timestamp and associated information about the source video 101 for the matching hash to the API server 130 (block 414). The API server 130 can send this timestamp to the playback device 110 and / or use this timestamp and associated source video information to query the metadata database 150 for metadata of the matching frame (block 416). The metadata database 150 returns the requested metadata to the API server 130, which then sends the requested metadata to the playback device 110 for display to the viewer (block 418). The playback device 110 displays the requested information to the viewer in an overlay that appears on or is integrated with the video. The displayed information can include links or other information that allow the viewer to purchase the product via the playback device 110 or another device, such as a smartphone or tablet. For more information regarding displaying metadata to viewers, see, for example, U.S. Patent No. _____, entitled "Dynamic Media-Product Searching Platform Apparatuses, Methods and Systems" (issuing U.S. patent application Ser. No. 14 / 527,854, which is incorporated herein by reference in its entirety).
[0044] conclusion While various inventive embodiments have been described and illustrated herein, those skilled in the art will readily envision a variety of other means and / or structures for performing the functions described herein and / or obtaining the results and / or one or more advantages, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the particular application(s) for which the teachings of the present invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. Thus, the foregoing embodiments are presented by way of example only, and it will be understood that, within the scope of the appended claims and their equivalents, the inventive embodiments may be practiced otherwise than as specifically described and claimed. The inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the inventive scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0045] Also, various inventive concepts can be embodied as one or more methods, examples of which have been provided. The acts performed as part of a method can be ordered in any suitable manner. As a result, embodiments can be constructed in which acts are performed in an order different from that illustrated, including performing some acts simultaneously, even when shown as sequential acts in the example embodiments.
[0046] All definitions and uses herein should be understood to control over dictionary definitions, definitions in documents incorporated herein by reference, and / or ordinary meanings of the defined terms.
[0047] The indefinite articles "a" and "an," as used in the specification and claims, unless expressly indicated otherwise, should be understood to mean "at least one."
[0048] The term "and / or" as used herein and in the claims should be understood to mean "either or both" of the conjoined elements, i.e., elements that are conjunctive in some cases and disjunctive in other cases. Multiple elements listed with "and / or" should be interpreted in the same manner, i.e., "one or more" of the conjunctive elements. Other elements may optionally be present other than the elements specifically identified by the "and / or" clause, whether related or unrelated to the elements specifically identified. Thus, as a non-limiting example, a reference to "A and / or B," when used in conjunction with open-ended language such as "comprising," can refer to A only (optionally including elements other than B) in one embodiment, B only (optionally including elements other than A) in another embodiment, both A and B (optionally including other elements) in yet another embodiment, and so forth.
[0049] As used herein and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as inclusive, i.e., including at least one, but also more than one, of a number or list of elements, and optionally additional items not listed. Only terms clearly indicated to the contrary, such as "only one of" or "exactly one of," or "consisting of" when used in the claims, shall refer to the inclusion of exactly one element of a number or list of elements. In general, as used herein, the term "or" shall only be interpreted to indicate exclusive alternatives (i.e., "one or the other but not both") when preceded by an exclusive term, such as "either," "one of," "only one of," or "exactly one of." "Consisting essentially of," when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0050] As used herein and in the claims, the phrase "at least one" in connection with a list of one or more elements should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, and not excluding any combination of elements in the list of elements. This definition also allows that elements other than those specifically identified in the list of elements to which the phrase "at least one" refers may optionally be present, whether related or unrelated to the elements specifically identified. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently, "at least one of A and / or B") can refer to at least one A (optionally including elements other than B), in one embodiment where B is absent, and optionally including two or more As; in another embodiment, at least one B (optionally including elements other than A), in which A is absent, and optionally including two or more Bs; in yet another embodiment, at least one A, optionally including two or more As, and at least one B (optionally including other elements), optionally including two or more Bs;
[0051] In the claims, as well as in the above specification, all transitional phrases, such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like, are to be understood as open-ended, i.e., meaning including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively, as defined in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03. The claims as originally filed are as follows: Claim 1: 1. A method for identifying frames of a source video, comprising the steps of: generating a hash vector for each frame of the different versions of the source video; associating the hash vector with information about the source video in a database; playing a first version of the source video on a playback device; generating a first hash vector for a first frame of the first version of the source video; matching the first hash vector to a matching hash vector among the hash vectors in the database; retrieving information about the source video from the database in response to matching the first hash vector to the matching hash vector; The method includes: Claim 2: The method of claim 1 , wherein the playback device comprises at least one of a television, a set-top box, a computer, or a mobile device. Claim 3: 2. The method of claim 1, wherein determining that the first hash vector matches the matching hash vector comprises determining that the first hash vector is within a threshold distance of the matching hash vector. Claim 4: The method of claim 1 , wherein the matching hash vector is for a frame of a second version of the source video that is different from the first version of the source video. Claim 5: The method of claim 1 , wherein the hash vector and the first hash vector are generated using a perceptual hashing process. Claim 6: 6. The method of claim 5, wherein the perceptual hashing process is one member of the group consisting of perception hashing (pHash), differential hashing (dHash), average hashing (aHash), and wavelet hashing (wHash). Claim 7: 7. The method of claim 6, wherein generating the first hash vector occurs within about 100 milliseconds. Claim 8: The method of claim 6 , wherein the first hash vector has a size of 4096 bits or less. Claim 9: separating the hash vectors into subsets based on at least one of how frequently the hash vectors have been accessed or how recently the hash vectors have been accessed; storing each subset in a different shard of said database; The method of claim 1 further comprising: Claim 10: Separating the hash vectors into subsets based on a distance between the hash vectors; storing each subset in a different shard of said database; The method of claim 1 further comprising: Claim 11: Separating the hash vectors into subsets based on characteristics of the hash vectors; storing each subset in a different shard of said database; The method of claim 1 further comprising: Claim 12: randomly separating the hash vectors into subsets; storing each subset in a different shard of said database; The method of claim 1 further comprising: Claim 13: The method of claim 1 , wherein generating the first hash vector occurs automatically at regular intervals. Claim 14: The method of claim 1 , wherein generating the first hash vector occurs in response to a command from a viewer. Claim 15: a database storing hash vectors for each frame of different versions of a source video, the hash vectors being associated in the database with information about the source video; and an application programming interface (API) communicatively coupled to the database for querying the database having a first hash vector for a first frame of a first version of the source video played on a playback device, and for returning information about the source video in response to a match between the first hash vector and a matching hash vector among the hash vectors in the database; A system comprising: Claim 16: 1. A method for identifying and obtaining metadata associated with a source video, comprising: generating a hash vector for each frame of at least one version of the source video; storing the hash vector in a first database; storing metadata corresponding to each of the frames in a second database; playing a first version of the source video on a playback device; generating a first hash vector for a first frame of the first version of the source video; matching the first hash vector to a matching hash vector among the hash vectors in the first database; retrieving the metadata corresponding to the matching hash vector from the second database in response to matching the first hash vector to the matching hash vector; displaying the metadata to the viewer via the playback device; The method includes: Claim 17: The method of claim 16 , wherein the playback device comprises at least one of a television, a set-top box, a computer, or a mobile device. Claim 18: 17. The method of claim 16, wherein the metadata represents at least one of locations in the source video, clothing worn by actors in the source video, products that appear in the source video, or music playing the source video. Claim 19: The method of claim 16 , wherein the hash vectors are associated with the metadata by their respective timestamps. Claim 20: Matching the first hash vector to the matching hash vector comprises: sending the first hash vector to an application programming interface (API) server; determining, via the API server, that the first hash vector matches the matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, identifying the timestamp associated with the matching hash vector in the first database; and searching the metadata further comprises: Querying the second database based on the timestamp; retrieving the metadata associated with the timestamp from the second database; 20. The method of claim 19, comprising: Claim 21: 17. The method of claim 16, wherein determining that the first hash vector matches the matching hash vector comprises determining that the first hash vector is within a threshold distance of the matching hash vector. Claim 22: 17. The method of claim 16, wherein the matching hash vector is for a frame of a second version of the source video that differs from the first version of the source video. Claim 23: The method of claim 16 , wherein the hash vector and the first hash vector are generated using a perceptual hashing process. Claim 24: 24. The method of claim 23, wherein the perceptual hashing process is one member of the group consisting of perception hashing (pHash), differential hashing (dHash), average hashing (aHash), and wavelet hashing (wHash). Claim 25: 25. The method of claim 24, wherein generating the first hash vector occurs within about 100 milliseconds. Claim 26: 25. The method of claim 24, wherein the first hash vector has a size of 4096 bits or less. Claim 27: Storing the hash vector comprises: separating the hash vectors into subsets based on at least one of how frequently the hash vectors have been accessed or how recently the hash vectors have been accessed; storing each subset in a different shard of the first database; 17. The method of claim 16, comprising: Claim 28: Storing the hash vector comprises: Separating the hash vectors into subsets based on a distance between the hash vectors; storing each subset in a different shard of the first database; 17. The method of claim 16, comprising: Claim 29: Storing the hash vector comprises: Separating the hash vectors into subsets based on characteristics of the hash vectors; storing each subset in a different shard of the first database; 17. The method of claim 16, comprising: Claim 30: storing the hash vector, randomly separating the hash vectors into equal subsets; storing each subset in a different shard of the first database; 17. The method of claim 16, comprising: Claim 31: 17. The method of claim 16, further comprising updating the metadata without changing the hash vector in the first database. Claim 32: 20. The method of claim 16, wherein generating the first hash vector occurs automatically at regular intervals. Claim 33: The method of claim 16 , wherein generating the first hash vector occurs in response to a command from a viewer. Claim 34: a first database for storing hash vectors for each frame of different versions of a source video, the hash vectors being associated in the first database with information about the source video; a second database for storing metadata relating to the source video; an application programming interface (API) communicatively coupled to the first database and the second database for querying the first database having a first hash vector for a first frame of a first version of the source video played on a playback device, and for querying the second database for at least a portion of the metadata related to the source video based on the information about the source video returned from the first database in response to a match between the first hash vector and a matching hash vector among the hash vectors in the database; A system comprising: Claim 35: 1. A method for identifying, obtaining, and displaying metadata associated with a video, comprising: playing said video via a display; generating a first hash vector for a first frame of the video; sending the first hash vector to an application programming interface (API) server; retrieving, via the API server, the metadata associated with the first frame from a metadata database, the metadata being retrieved from the first database in response to matching the first hash vector to a second hash vector stored in a hash vector database; displaying the metadata associated with the first frame to a user via the display; and The method includes: Claim 36: receiving, at a database, a first hash vector generated for a first frame of the video; storing the first hash vector in the database; receiving, at the database, a query from a playback device based on the second hash vector; executing the query of the database against the second hash vector; in response to matching the second hash vector to the first hash vector, sending a timestamp associated with the first hash vector to an application programming interface (API) server, the timestamp associating metadata with the first frame of the video; The method includes: Claim 37: generating a first hash vector for a first frame of the source video; storing the first hash vector in a first database; playing a version of the source video on a playback device; generating a second hash vector for a second frame of the version of the source video; matching the second hash vector to the first hash vector in the first database; The method includes: Claim 38: retrieving a timestamp corresponding to the second hash vector in response to matching the second hash vector to the first hash vector; transmitting said time stamp to said playback device; 38. The method of claim 37, comprising:
Claims
1. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video within about 100 milliseconds using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; The method includes:
2. The method of claim 1 , wherein the playback device comprises at least one of a television, a set-top box, a computer, or a mobile device.
3. 2. The method of claim 1, wherein determining that the first hash vector matches the matching hash vector comprises the system determining that the first hash vector is within a threshold distance of the matching hash vector.
4. The method of claim 1 , wherein the matching hash vector is for a frame of a second version of the source video that is different from the first version of the source video.
5. 2. The method of claim 1, wherein the perceptual hashing process is one member of the group consisting of perception hashing (pHash), differential hashing (dHash), average hashing (aHash), and wavelet hashing (wHash).
6. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; Including, The method of claim 1, wherein the first hash vector has a size of 4096 bits or less.
7. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system separating the hash vectors into subsets based on at least one of how frequently the hash vectors have been accessed or how recently the hash vectors have been accessed; the system storing each subset in a different shard of the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector using the perceptual hashing process for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; A method comprising:
8. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system separating the hash vectors into subsets based on the distance between the hash vectors; the system storing each subset in a different shard of the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; A method comprising:
9. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system separating the hash vectors into subsets based on characteristics of the hash vectors; the system storing each subset in a different shard of the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; A method comprising:
10. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system randomly separating the hash vectors into subsets; the system storing each subset in a different shard of the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; A method comprising:
11. The method of claim 1 , wherein generating the first hash vector occurs automatically at regular intervals.
12. 1. A method for identifying frames of a source video by a system including a playback device for playing a first version of the source video, a database storing hash vectors for each frame of different versions of the source video and metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the database, comprising: the system generating a hash vector for each frame of the different versions of the source video using a perceptual hashing process; the system associating the hash vector with information about the source video in the database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the database; in response to matching the first hash vector to the matching hash vector, the system retrieves information about the source video from the database; Including, A method wherein generating the first hash vector occurs in response to a command from a viewer.
13. a database storing hash vectors for each frame of different versions of a source video, the hash vectors being generated using a perceptual hashing process and associated in the database with information about the source video; and an application programming interface (API) server communicatively coupled to the database for querying the database having a first hash vector generated within about 100 milliseconds using the perceptual hashing process for a first frame of a first version of the source video played on a playback device, and for returning information about the source video in response to a match between the first hash vector and a matching hash vector among the hash vectors in the database; A system comprising:
14. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, the hash vectors are associated with the metadata by their respective timestamps; Matching the first hash vector to the matching hash vector includes: the system sending the first hash vector to the API server; the API server determining that the first hash vector matches the matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system identifying the timestamp associated with the matching hash vector in the first database; Including, Retrieving the metadata includes: the system querying the second database based on the timestamp; the system retrieving the metadata associated with the timestamp from the second database; The method further comprising:
15. The method of claim 14 , wherein the playback device comprises at least one of a television, a set-top box, a computer, or a mobile device.
16. 15. The method of claim 14, wherein the metadata represents at least one of locations in the source video, clothing worn by actors in the source video, products appearing in the source video, or music playing the source video.
17. 15. The method of claim 14, wherein determining that the first hash vector matches the matching hash vector comprises the system determining that the first hash vector is within a threshold distance of the matching hash vector.
18. The method of claim 14 , wherein the matching hash vector is for a frame of a second version of the source video that is different from the first version of the source video.
19. The method of claim 14 , wherein the hash vector and the first hash vector are generated using a perceptual hashing process.
20. 20. The method of claim 19, wherein the perceptual hashing process is one member of the group consisting of perception hashing (pHash), differential hashing (dHash), average hashing (aHash), and wavelet hashing (wHash).
21. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: the system generating a hash vector for each frame of at least one version of the source video using a perceptual hashing process; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, A method wherein generating the first hash vector occurs within about 100 milliseconds.
22. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: the system generating a hash vector for each frame of at least one version of the source video using a perceptual hashing process; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video using the perceptual hashing process; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, The method of claim 1, wherein the first hash vector has a size of 4096 bits or less.
23. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, Storing the hash vector comprises: the system separating the hash vectors into subsets based on at least one of how frequently the hash vectors have been accessed or how recently the hash vectors have been accessed; storing each subset in a different shard of the first database; A method comprising:
24. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, Storing the hash vector comprises: the system separating the hash vectors into subsets based on the distance between the hash vectors; storing each subset in a different shard of the first database; A method comprising:
25. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, Storing the hash vector comprises: the system separating the hash vectors into subsets based on characteristics of the hash vectors; storing each subset in a different shard of the first database; A method comprising:
26. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; Including, storing the hash vector, the system randomly separating the hash vectors into equal subsets; storing each subset in a different shard of the first database; A method comprising:
27. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; the system displaying the metadata to a viewer via the playback device; updating the metadata without changing the hash vector in the first database; A method comprising:
28. The method of claim 14 , wherein the system automatically generates the first hash vector at regular intervals.
29. 1. A method for identifying and obtaining metadata associated with a source video by a system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata about the source video, and an application programming interface (API) server for retrieving information about the source video from the second database, the method comprising: generating a hash vector for each frame of at least one version of the source video; the system storing the hash vector in the first database; the system storing metadata corresponding to each of the frames in the second database; the system playing the first version of the source video on the playback device; in response to a command from a viewer, the system generating a first hash vector for a first frame of the first version of the source video; the system matching the first hash vector to a matching hash vector among the hash vectors in the first database; in response to matching the first hash vector to the matching hash vector, the system retrieves the metadata corresponding to the matching hash vector from the second database; said system displaying said metadata to a viewer via said playback device; A method comprising:
30. a first database for storing hash vectors for each frame of different versions of a source video, the hash vectors being associated in the first database with information about the source video, the hash vectors separating shades based on distance between the hash vectors, or based on characteristics of the hash vectors, or based on at least one of how frequently the hash vectors have been accessed or how recently the hash vectors have been accessed; a second database for storing metadata about the source video; an application programming interface (API) server communicatively coupled to the first database and the second database for querying the first database having a first hash vector for a first frame of a first version of the source video played on a playback device, and for querying the second database for at least a portion of the metadata related to the source video based on the information about the source video returned from the first database in response to a match between the first hash vector and a matching hash vector among the hash vectors in the database; A system comprising:
31. 1. A method for identifying, obtaining, and displaying metadata associated with a video by a system including a playback device for playing a first version of a source video, a hash vector database for storing hash vectors for each frame of different versions of the source video, a metadata database for storing metadata, which is information about the source video, and an application programming interface (API) server for retrieving information about the source video from the metadata database, comprising: said playback device playing said video to a viewer; in response to a command from the viewer, the system generating a first hash vector for a first frame of the video; the system sending the first hash vector to the API server; the API server retrieving the metadata associated with the first frame from the metadata database, the metadata being retrieved from the metadata database in response to matching the first hash vector to a second hash vector stored in the hash vector database; the playback device displays the metadata associated with the first frame to a user; and The method includes:
32. A system including a playback device for playing a first version of a source video, a first database for storing hash vectors for each frame of different versions of the source video, a second database for storing metadata related to the source video, and an application programming interface (API) server for retrieving the metadata related to the source video, comprising: receiving, in the first database, a first hash vector generated for a first frame of a video; the system storing the first hash vector in the first database; the system receiving a query at the first database from the playback device based on the second hash vector; the system executing the query of the first database against the second hash vector; in response to matching the second hash vector to the first hash vector, the system sending a timestamp associated with the first hash vector to the API server, the timestamp causing the system to associate metadata with the first frame of the video; the system querying the second database based on the timestamp; the system retrieving the metadata associated with the timestamp from the second database; The method includes:
33. A system including a playback device for playing a first version of a source video and a database storing hash vectors for each frame of different versions of the source video, the system comprising: generating a first hash vector for a first frame of the source video using a perceptual hashing process; the system storing the first hash vector in the database; the system playing a version of the source video on the playback device; the system generating a second hash vector for a second frame of the version of the source video within 100 milliseconds using the perceptual hashing process; the system matching the second hash vector to the first hash vector in the database; The method includes:
34. in response to matching the second hash vector to the first hash vector, the system retrieving a timestamp corresponding to the second hash vector; said system transmitting said time stamp to said playback device; 34. The method of claim 33, comprising:
Citation Information
Patent Citations
Server system and method for authenticating document image
JP2008243209A
Media fingerprinting for determining and searching content
JP2013529325A
Moving image information acquisition system
JP2015233182A
Video hashing system and method
US8494234B1
Video reception device, added-information display method, and added-information display system
WO2015015712A1