Entity card containing descriptive content related to entities from video
By employing machine learning to identify key entities in video transcripts and generating entity cards displayed in sync with the video, the method addresses user challenges in understanding unfamiliar content, enhancing the viewing experience.
Patent Information
- Application Number
- JP2024566821
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-30
- Filing Date
- 2023-05-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-05-11
AI Technical Summary
Users watching videos on unfamiliar topics may struggle to understand content due to unfamiliar keywords or concepts, leading to difficulties in searching for information and potentially causing users to stop watching the video.
A computer-implemented method that uses machine learning resources to identify entities likely to be searched by users based on video transcripts, generates entity cards with descriptive content, and displays these cards in a user interface while the video plays, synchronizing with mentions of the entities in the video.
This solution enhances user understanding of video content by providing immediate access to descriptive information about mentioned entities, reducing the need for users to pause the video or perform separate searches, thereby improving user experience and retention.
Smart Images

Figure 2025517205000001_ABST
Abstract
Description
Technical Field
[0001] Claims of Priority This application claims priority to U.S. Application No. 18 / 091,837, filed Dec. 30, 2022, which claims priority to U.S. Provisional Application No. 63 / 341,674, filed May 13, 2022. Each of these is hereby incorporated by reference in its entirety for all purposes.
[0002] The present disclosure generally relates to providing entity cards to a user interface in connection with video displayed on a user computing device's display. More particularly, the present disclosure relates to providing entity cards that assist in understanding the content of a video and that include descriptive content related to entities (e.g., concepts, terms, topics, etc.) mentioned in the video.
Background Art
[0003] When a user watches a video, e.g., a video on a challenging or new topic, there may be keywords or concepts that the user is not familiar with but that would be helpful in understanding the content of the video. For example, in a video about the pyramids of Egypt, the term "sarcophagus" may be an important concept that is widely discussed. However, a user who is not familiar with the term "sarcophagus" may not be able to fully understand the content of the video. The user can pause the video and move to a search page to perform a search for the term "sarcophagus". In some cases, the user may have difficulty spelling the term they want to search, may not obtain accurate search results, or may experience inconvenience when searching. In other examples, the user may stop watching the video after determining that the content of the video is difficult to understand.
Summary of the Invention
[0004] Aspects and advantages of embodiments of the present disclosure will be described in part in the following description, or can be learned from the description, or can be learned through the practice of exemplary embodiments.
[0005] In one or more exemplary embodiments, a computer-implemented method for a server system includes obtaining a transcript of content from a video, applying machine learning resources to identify one or more entities that are most likely to be searched for by a user viewing the video based on the transcript of the content, generating one or more entity cards for each of the one or more entities, wherein each of the one or more entity cards includes descriptive content regarding the respective entity among the one or more entities, providing a user interface that is displayed on a respective display of one or more user computing devices, playing the video in a first portion of the user interface, and when the video is played and a first entity among the one or more entities is mentioned in the video, displaying the first entity card in a second portion of the user interface, wherein the first entity card includes descriptive content related to the first entity, and providing the user interface.
[0006] In some embodiments, applying machine learning resources to identify one or more entities includes obtaining training data to train the machine learning resources based on observational data of users who perform searches in response to viewing only the video.
[0007] In some embodiments, applying machine learning resources to identify one or more entities includes associating text from a transcript with a knowledge graph to identify multiple candidate entities from a video, and for each of the candidate entities, its relevance to the topic of the video, its relevance to one or more other candidate entities among the multiple candidate entities, the number of mentions of the candidate entity within the video, and the number of videos in which the candidate entity appears across a corpus of videos stored in one or more databases, and ranking the candidate entities to obtain one or more entities based on one or more of the foregoing.
[0008] In some embodiments, applying machine learning resources to identify one or more entities includes evaluating user interactions with a user interface and determining at least one adjustment to the machine learning resources based on the evaluation of the user interactions with the user interface.
[0009] In some embodiments, a first entity is mentioned within a video at a first point in time within the video, and a first entity card is displayed in a second portion of the user interface at the first point in time.
[0010] In some embodiments, one or more entities include a second entity, one or more entity cards include a second entity card, and a method includes providing a user interface that is displayed on a respective display of one or more user computing devices, the method including, while continuously playing a video, displaying, in a third portion of the user interface, the second entity card in a collapsed form, the collapsed second entity card referring to the second entity mentioned within the video at a second timepoint after a first timepoint; and, when the second entity is mentioned within the video at the second timepoint, while continuously playing the video, displaying, in the third portion of the user interface, the second entity card in a fully expanded form, the fully expanded second entity card including descriptive content regarding the second entity. The method further includes providing the user interface for these purposes.
[0011] In some embodiments, one or more entities include a second entity, one or more entity cards include a second entity card, and a method includes providing a user interface that is displayed on a respective display of one or more user computing devices, the method including, when the second entity is mentioned within the video while the video is being played, while continuously playing the video, displaying, in a second portion of the user interface, the second entity card, the second entity card including descriptive content regarding the second entity, the second entity card being displayed in the second portion of the display by replacing a first entity card at the timepoint at which the second entity is mentioned within the video. The method further includes providing the user interface for these purposes.
[0012] In some embodiments, the method is to provide a user interface that is displayed on each display of one or more user computing devices, and when a first entity is mentioned in the video while the video is being played, while continuing to play the video, display a notification user interface element in a third portion of the user interface, where the notification user interface element indicates that additional information related to the first entity is available, and in response to the first entity being mentioned in the video while the video is being played and in response to receiving a selection of the notification user interface element, while continuing to play the video, display a first entity card in a second portion of the user interface, and further includes providing the user interface for these purposes.
[0013] In some embodiments, the first entity card includes at least one of a text summary providing information related to the first entity or an image related to the first entity.
[0014] In one or more exemplary embodiments, a computer-implemented method for a user computing device includes receiving a video for playback in a user interface, providing the video for display in a first portion of the user interface that is displayed on a display of the user computing device, and when a first entity is mentioned in the video while the video is being played, providing a first entity card for display in a second portion of the user interface while continuing to play the video, where the first entity card includes descriptive content related to the first entity and the first entity is generated in response to an automatic recognition of the first entity from a transcript of the video content.
[0015] In some embodiments, a first entity is mentioned within a video at a first point in time, and a first entity card is provided for display in a second portion of a user interface at the first point in time.
[0016] In some embodiments, the method includes providing a shrunk second entity card that references a second entity mentioned within the video at a second point in time after the first point in time for display in a third portion of the user interface, and expanding the shrunk second entity card to fully display the second entity card in the third portion of the user interface while the video is being played, if the second entity is mentioned within the video at the second point in time, where the second entity card includes descriptive content regarding the second entity.
[0017] In some embodiments, the method further includes providing a second entity card for display in a second portion of the user interface while the video is being played, if the second entity is mentioned within the video while the video is being played, where the second entity card includes descriptive content regarding the second entity, and providing the second entity card for display in the second portion of the user interface by replacing the first entity card at the point in time when the second entity is mentioned within the video.
[0018] In some embodiments, the method includes providing one or more entity search user interface elements on the user interface that are configured to perform a search regarding the first entity when selected.
[0019] In some embodiments, the method includes providing, for display on a user interface, one or more search query user interface elements that, when selected, are configured to perform a search regarding a topic of a video other than a first entity.
[0020] In some embodiments, the method includes identifying a first entity and utilizing machine learning resources to generate a first entity card.
[0021] In some embodiments, the first entity is an entity among a plurality of entities mentioned in the video, and is determined by machine learning resources as the entity that is most likely to be searched by a user viewing the video among the plurality of entities mentioned in the video.
[0022] In some embodiments, the method is to provide a notification user interface element for display in a third portion of the user interface while the video is being played, if the first entity is mentioned in the video while the video is being played, the notification user interface element indicating that additional information related to the first entity is available, and in response to receiving a selection of the notification user interface element, providing a first entity card for display in a second portion of the user interface while the video is being played.
[0023] In some embodiments, the first entity card includes a text summary providing information related to the first entity and / or an image related to the first entity.
[0024] In one or more exemplary embodiments, a user computing device includes a display, one or more memories storing instructions, and one or more processors executing the instructions stored in the one or more memories, the one or more processors being configured to receive video for playback in a user interface, provide the video for display in a first portion of a user interface displayed on the display, and, while the video is being played, provide a first entity card for display in a second portion of the user interface while the video continues to be played if a first entity is mentioned within the video. The first entity card includes descriptive content related to the first entity, and the first entity is generated in response to an automatic recognition of the first entity from a transcript of the video content.
[0025] In one or more exemplary embodiments, the server system includes one or more memories that store instructions, obtains a transcript of content from video, applies machine learning resources to identify one or more entities that are most likely to be searched for by a user viewing the video based on the transcript of the content, and for each of the one or more entities, generates one or more entity cards, wherein each of the one or more entity cards includes descriptive content regarding the respective entity among the one or more entities, and provides a user interface that is displayed on a respective display of one or more user computing devices, wherein the user interface plays the video in a first portion of the user interface, and when the video is played and a first entity among the one or more entities is mentioned within the video, displays the first entity card in a second portion of the user interface, wherein the first entity card includes descriptive content associated with the first entity, and includes one or more processors that execute instructions stored in the one or more memories.
[0026] In one or more exemplary embodiments, there is provided a computer-readable medium (e.g., a non-transitory computer-readable medium) storing instructions executable by one or more processors of a user computing device and / or a server system. In some embodiments, the computer-readable medium may store instructions that cause one or more processors to perform one or more operations of any of the methods described herein (e.g., operations of the server system and / or operations of user computing). The computer-readable medium may also store additional instructions for performing other aspects of the server system and user computing device, as well as corresponding methods of operation, as described herein.
[0027] These and other features, aspects, and advantages of the various embodiments of the present disclosure will become better understood with reference to the following description, drawings, and appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.
[0028] A detailed description of exemplary embodiments directed to those skilled in the art is set forth in this specification with reference to the accompanying drawings.
Brief Description of the Drawings
[0029]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 5C
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8
Figure 9
Figure 10
Best Mode for Carrying Out the Invention
[0030] Here, embodiments of the present disclosure are referred to. One or more examples of the embodiments are shown in the drawings, and like reference numerals indicate like elements. Each example is provided as an illustration of the present disclosure and is not intended to limit the present disclosure. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the present disclosure without departing from the scope or spirit of the present disclosure. For example, features illustrated or described as part of one embodiment can be used with other embodiments to create still other embodiments. Accordingly, it is intended that the present disclosure cover modifications and variations that fall within the scope of the appended claims and their equivalents.
[0031] The terms used herein are used to describe exemplary embodiments and are not intended to limit and / or restrict the present disclosure. The singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly dictates otherwise. In the present disclosure, terms such as “including,” “having,” “comprising,” etc. are used to specify features, numbers, steps, operations, elements, components, or combinations thereof, but do not exclude the presence or addition of one or more of features, elements, steps, operations, elements, components, or combinations thereof.
[0032] The terms first, second, third, etc. may be used herein to describe various elements, but it will be understood that the elements should not be limited by these terms. Rather, these terms are used only to distinguish one element from another. For example, without departing from the scope of the present disclosure, the first element may be referred to as the second element, and the second element may be referred to as the first element.
[0033] When an element is referred to as being "connected" to another element, it will be understood that such a statement includes examples of direct connection or direct coupling, as well as connection or coupling with one or more other elements intervening therebetween.
[0034] The term "and / or" includes combinations of multiple related listed items, or any of the multiple related listed items. For example, the scope of the expression or phrase "A and / or B" includes the item "A", the item "B", and the combination of the item "A and B".
[0035] Furthermore, the scope of the expression or phrase "at least one of A or B" is intended to include all of (1) at least one of A, (2) at least one of B, and (3) at least one of A and at least one of B. Similarly, the scope of the expression or phrase "at least one of A, B, or C" is intended to include all of (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B, and at least one of C.
[0036] According to an exemplary embodiment, when a user is viewing a video in the user interface of a display of a user computing device and an entity (e.g., a concept, term, topic, etc.) is mentioned in the video, an entity card containing information about the entity that may be useful for the user to understand the content of the video is provided in the user interface. For example, when the user is viewing a video about the pyramids of Egypt, an entity card may be provided in the user interface to provide additional information about the term "sarcophagus", such as a definition of the term, a photograph of a sarcophagus, etc. The entity card is provided in the user interface while the video is being played, so that the user does not need to leave the application or web page that is playing the video to navigate away in order to learn more about the entity (e.g., the term "sarcophagus"). Thus, information about potentially difficult concepts, or concepts that the user may be trying to learn more about, is presented to the user, which may help the user quickly understand the concept without leaving the application or web page that is playing the video, and the user does not need to perform a separate search regarding that concept.
[0037] In some embodiments, the entire entity card may be visible to the user on the user interface, or a part of the entity card may be visible on the user interface, and the user can select a part of the entity card to expand the entity card and view the hidden part of the entity card for more information about the entity.
[0038] In some embodiments, the entity card can be displayed on the user interface simultaneously with the entity being mentioned in the video. That is, the display of the entity card is synchronized with the time when the entity is mentioned in the video. The entity card may be displayed on the user interface only when the entity is first mentioned in the video each time the entity is mentioned in the video, or may be selectively displayed when the entity is mentioned multiple times in the video.
[0039] In some embodiments, a user interface element separate from the entity card can be provided on the user interface to enable the user to perform a search regarding the entity. For example, if the user desires to obtain further information about the entity beyond the information provided in the entity card, the user can select a user interface element that causes a search to be performed regarding the entity, and a search results page can be displayed on the display of the user computing device.
[0040] In some embodiments, one or more user interface elements separate from the entity card can be provided on the user interface corresponding to each proposed search query. The one or more user interface elements enable the user to perform a search for information other than the entity itself, for example, regarding other entities or other topics covered in the video. For example, if the user desires to obtain further information about other entities or other topics covered in the video, the user can select the corresponding user interface element that causes a search to be performed regarding the corresponding entity or topic, and a search results page can be displayed on the display of the user computing device.
[0041] In some embodiments, there may be multiple concepts in a video that may potentially be difficult for a user to understand (or potentially of interest to the user). According to an exemplary embodiment, an entity card for each of the multiple concepts within the video may be displayed on the user interface as the user is viewing the video, and the multiple concepts are mentioned. For example, the user interface provides a separate entity card for each concept that includes information about the concept that may help the user to understand the content of the video as a whole. The entity cards are provided to the user interface while the video is being played, thereby eliminating the need for the user to navigate away from the application or web page that is playing the video. Thus, information about potentially difficult concepts, or concepts that the user may be interested in learning more about, is presented to the user and may help the user to quickly understand the concepts without having to navigate away from the application or web page that is playing the video, and the user does not need to perform separate searches for each of the concepts.
[0042] In some embodiments, all of the entity cards associated with the video may be visible to the user on the user interface simultaneously while the video is being played, or only some of the entity cards associated with the video may be visible to the user on the user interface simultaneously while the video is being played. For example, one or more of the entity cards may be fully expanded so that the user can view the entire content of the entity card, and some or all of the remaining entity cards may be displayed on the user interface in a collapsed or hidden form. For example, in a collapsed or hidden form, the user can view a part of the entity card, and the part of the entity card may include some identification information (e.g., identification of the corresponding entity) so that the user can understand the relevance of the entity card. For example, the user can select a part of the entity card to expand the entity card and view also the hidden part of the entity card for more information about the entity.
[0043] In some embodiments, an entity card associated with a video may be visible to a user on a user interface as the video progresses, and the user may not be able to view the entity card until the corresponding entity is mentioned in the video. For example, a first entity card for a first entity may be displayed on the user interface at a point in the video (e.g., at a first time point) when the first entity is mentioned in the video. The first entity card may be displayed for a predetermined time only (e.g., a time sufficient for an average user to read or view the content included in the first entity card) while the video continues to play, or until the next entity is mentioned in the video. At the point when the next entity is displayed, other entity cards are provided on the user interface. For example, a second entity card for a second entity may be displayed on the user interface at a point in the video (e.g., at a second time point) when the second entity is mentioned in the video.
[0044] In some embodiments, the second entity card may be displayed on the user interface by replacing the first entity card (i.e., by occupying some or all of the space on the user interface previously occupied by the first entity card).
[0045] In some embodiments, the second entity card may be in a shrunk or hidden form, and when the second entity is mentioned in the video at the second time point, the second entity card may be expanded so that the second entity card can be fully displayed on the user interface at the second time point. If the first entity card is not already in a shrunk or hidden form before the second entity card is displayed, the first entity card may be changed to a shrunk or hidden form when the second entity card is expanded. In some embodiments, each of the first and second entity cards may be fully displayed on the user interface.
[0046] In some embodiments, when an entity is mentioned within a video during video playback, the notification user interface element is displayed on the user interface while the video continues to play. For example, the notification user interface element indicates that additional information related to the entity is available. In response to receiving a selection of the notification user interface element, an entity card is displayed on the user interface while the video continues to play. The notification user interface element may include an image (such as a thumbnail image) of the entity to further inform the user that the notification user interface element is associated with the entity and that an entity card related to the entity is available.
[0047] In some embodiments, a timeline may be displayed on the user interface while a video is being played, and the timeline indicates one or more points along the timeline at which information about one or more entities is available via corresponding entity cards. As the video approaches or passes each point, the corresponding entity card is displayed on the user interface and information about the entity is provided. For example, a user interface element may be provided that, when selected, enables entity cards to cycle through the user interface as the video progresses along the timeline.
[0048] According to the exemplary embodiments disclosed herein, for one or more entities mentioned in a video, while the video is being played, one or more entity cards are provided on the user interface of the display of a user computing device. For example, the entity for which the entity card is provided can be identified from the video by using machine learning resources. For example, one of the plurality of candidate entities mentioned in the video is identified by the machine learning resources as the entity for which the entity card should be generated when the machine learning resources determine or predict that the entity is likely to be searched by the user watching the video (e.g., having a confidence value greater than a threshold, a probability of being searched greater than a threshold, and being determined to be the most likely to be searched by the user watching the video compared to other entities mentioned in the video). For example, the machine learning resources can select a candidate entity as the entity for which the entity card should be generated based on one or more of the relevance of each of the plurality of candidate entities to the topic of the video, the relevance of each of the plurality of candidate entities to one or more other candidate entities among the plurality of candidate entities, the number of mentions of the candidate entity in the video, and the number of videos in which the candidate entity is displayed across the corpus of videos stored in one or more databases).
[0049] According to the exemplary embodiments disclosed herein, a server system includes one or more servers that provide a video and entity cards for display on the user interface of the display of a user computing device.
[0050] According to the examples disclosed in this specification, entity cards are generated based on entities identified from videos. For example, a video can be analyzed using an automatic speech recognition program to perform automatic speech recognition and obtain a transcript of the speech (text transcript). The next operation may include associating some or all of the text from the transcript with knowledge graph entities to obtain a set of knowledge graph entities associated with the video. The training data for the machine learning resource may be obtained based on identifying those knowledge graph entities that also appear in search queries from the actual users watching the video. Additional operations for identifying entities from the video that may be important for understanding the video content and / or entities from the video that are likely to be searched by the user include determining how related an entity is to other entities within the video, determining how extensively an entity uses the tf-idf (term frequency-inverse document frequency) signal across the corpus of videos, and determining how related an entity is to the topic of the video (e.g., using signals of prominent terms based on the query). The machine learning resource can be trained by applying weights to candidate entities (e.g., the higher the frequency with which a term is mentioned within the video, the higher weight may be assigned to the entity, a lower weight may be assigned to entities that are overly broad and frequently appear in the video corpus, and a higher weight may be assigned to entities that are more related to the topic of the video, etc.). The machine learning resource can be applied to evaluate candidate entities and rank candidate entities among those identified within the video. For example, one or more of the top-ranked candidate entities can be selected as the entities for which the entity card is generated.
[0051] To generate an entity card, information about the entity can be obtained from various sources and text and / or image(s) providing more information about the entity can be input into the entity card. For example, the information about the entity can be obtained from one or more websites, electronic services that provide summaries of topics, etc.
[0052] In some embodiments, the entity card may be limited to less than a predetermined length and / or size (e.g., less than 100 words). The entity card may include information including a title (e.g., a title identifying the entity such as the title of "Barack Obama"), a subtitle (e.g., "former President of the United States" for the previous example of the entity and the title of Barack Obama), and one or more of attribution information (providing attribution to the source of the information). For example, the image may be limited to a thumbnail image size, a specified resolution, etc.
[0053] According to the examples disclosed herein, the following operations include rendering a user interface displayed on a display of a user computing device, the user interface including a video and one or more entity cards associated with the video, and the one or more entity cards may be displayed in the user interface at various points in the video or may be displayed (fully or partially) throughout the video.
[0054] For example, machine learning resources can be updated or adjusted based on an evaluation of user interactions with the user interface. For example, if a user typically does not interact with a particular entity card during a video, the entity can implicitly indicate that it is not a topic the user is interested in or a term or topic the user does not understand with respect to the content of the video. Thus, the machine learning resources may be adjusted to reflect the user interactions (or lack thereof) with the user interface. Similarly, if a user typically interacts with a particular entity card during a video, the entity can implicitly indicate that it is a topic the user is interested in or a term or topic the user does not understand with respect to the content of the video. Thus, the machine learning resources may be adjusted to reflect the user interactions with the user interface.
[0055] The systems and methods of the present disclosure provide several technical effects and advantages. In one embodiment, the present disclosure provides a way for a user to easily understand or further learn entities (e.g., terms, concepts, topics, etc.) associated with a video, and similarly, to easily identify content that the user may wish to consume for further details. By providing such a user interface, the user can more quickly understand the entities without having to perform separate searches and without having to pause the video, improving the user experience. The user can also determine whether they are interested in learning more about an entity by performing a search for the entity after a simple summary (e.g., snippet) and / or image related to the entity is provided via an entity card. In such a manner, the user can avoid performing searches for entities, loading search results, and reading various content items that may or may not be related to the entity and that are more computationally expensive than simply reading the information already presented, thus saving the time, processing, memory, and network resources of a computing system (server device, client device, or both). Similarly, the user's convenience and experience are improved because the user is more likely to watch the entire video without being discouraged by the complexity of the video content. Also, since machine learning resources predict entities that the user is likely to search for and information about those entities is automatically displayed during the video, the user is not discouraged by incorrect search results due to spelling mistakes, improving the user's convenience and experience. Accordingly, the user can avoid loading / viewing content from search result pages, which also saves the processing, memory, and network resources of the computing system.In addition, it is possible to avoid the inconvenience of the user switching between an application or web page for playing a video and a search result page, and instead continuously play the video without interruption. At the same time, since the entity card can be accessed or presented during the video, the operability and experience of the user are improved.
[0056] In some cases, a system of the type disclosed herein can learn content items, viewpoints, source type balance, and / or other preferred attributes through one or more various machine learning techniques (e.g., by training a neural network or other pre-trained machine learning model) based on different contexts such as different types of content, different user groups, timing, and location. For example, data (e.g., "click", "high rating", or the like) that describes actions performed by one or more users with respect to a user interface in scenarios viewed from various contexts can be stored and used as training data to train one or more pre-trained machine learning models (e.g., via supervised training techniques) to generate predictions that assist in providing content (e.g., entity cards) that meet the respective preferences of one or more users in the user interface. In this way, by reducing manual intervention, system performance is improved, user searches are reduced, and the processing, memory, and network resources of the computing system (server device, client device, or both) are further conserved.
[0057] Referring now to the drawings, FIG. 1 is an exemplary system according to one or more exemplary embodiments of the present disclosure. FIG. 1 shows an example of a system including a user computing device 100, a server computing system 300, a video data store 370, an entity data store 380, an entity card data store 390, and external content 400. Each of these can communicate with each other via a network 200.
[0058] For example, the user computing device 100 can include any one of a personal computer, a smartphone, a laptop, a tablet computer, and the like.
[0059] For example, the network 200 can include any type of communication network such as a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a personal area network (PAN), a virtual private network (VPN), and the like. For example, wireless communication between elements of the exemplary embodiments can be performed via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), ultra-wideband (UWB), Infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), a radio frequency (RF) signal, and the like. For example, wired communication between elements of the exemplary embodiments can be performed via a pair cable, a coaxial cable, an optical fiber cable, an Ethernet cable, and the like.
[0060] For example, the server computing system 300 can include servers that communicate with each other in a distributed manner, for example, or a combination of servers (such as a web server, an application server, etc.).
[0061] In an exemplary embodiment, the server computing system 300 may obtain data or information from one or more of a video data store 370, an entity data store 380, an entity card data store 390, and external content 400. The video data store 370, the entity data store 380, and the entity card data store 390 may be provided integrally with the server computing system 300 (e.g., as part of the memory 320 of the server computing system 300), or may be provided separately (e.g., remotely). Further, the video data store 370, the entity data store 380, and the entity card data store 390 may be combined as a single data store (database), or may be multiple respective data stores. Data stored in one data store (e.g., the entity data store 380) may overlap with some data stored in another data store (e.g., the entity card data store 390). In some embodiments, one data store (e.g., the entity card data store 390) may reference data stored in another data store (e.g., the entity data store 380).
[0062] The video data store 370 can store videos and / or information about videos. For example, the video data store 370 can store a collection of videos. The videos can be stored, grouped, or classified in any manner. For example, the videos can be stored according to genre or category, according to title, according to date (e.g., creation or last modification date, etc.). Information about the videos can include location information (e.g., Uniform Resource Locator (URL)) regarding where the videos can be stored or accessed. Information about a video may include the transcript information of the video. For example, a computing system may be configured to perform automatic speech recognition on a video stored in a video data store 370 or stored elsewhere (e.g., using one or more speech recognition programs) to obtain a transcript of the video (i.e., a text transcript), and the text transcript of the video may be stored in the video data store 370.
[0063] The entity data store 380 can store information about entities identified from the text transcript of a video. The identification of entities within a video will be described in more detail below. An entity card can be created or generated for a video with respect to an entity that is identified from the video and can be stored in the entity data store 380. For example, an entity for which an entity card is provided can be identified from a video by using machine learning resources. For example, one of a plurality of candidate entities mentioned in a video can be identified by machine learning resources as the entity for which an entity card should be generated when the machine learning resources determine or predict that the entity is likely to be searched by a user watching the video (e.g., having a confidence value greater than a threshold, a probability of being searched greater than a threshold, and being determined to be the most likely to be searched by a user watching the video compared to other entities mentioned in the video). For example, the machine learning resources can select a candidate entity as the entity for which an entity card should be generated based on one or more of the relevance of each of a plurality of candidate entities to the topic of the video, the relevance of each of the plurality of candidate entities to one or more other candidate entities among the plurality of candidate entities, the number of mentions of the candidate entity within the video, and one or more of the number of videos in which the candidate entity is shown across a corpus of videos stored in one or more databases (e.g., the video data store 370). For example, entities that can be identified from a video titled "The myth of Icarus and Daedalus" can include "Greek mythology", "Crete", "Icarus", and "Daedalus". These entities can be stored in the entity data store 380 and associated with the video titled "The myth of Icarus and Daedalus" that can be stored or referenced in the video data store 370.
[0064] The entity card data store 390 can store information regarding entity cards created or generated from entities identified from videos. The creation or generation of entity cards is described in more detail below. For example, an entity card can be created or generated by obtaining information regarding the entity from various sources (e.g., external content 400) and inputting into the entity card text and / or image(s) that provide more information regarding the entity. For example, the information regarding the entity can be obtained from one or more websites, an electronic service that provides a summary of a topic, etc. For example, the entity cards stored in the entity card data store 390 may be limited to a predetermined length and / or size less than (e.g., less than 100 words). The entity cards stored in the entity card data store 390 may include information including a title (e.g., a title that identifies the entity such as the title "Barack Obama"), a subtitle (e.g., "former President of the United States" regarding the previous entity and the title example of Barack Obama), and one or more of attribution information (that provides attribution to the source of the information). For example, the image forming part or all of the entity card may be limited to a thumbnail image size, a specified resolution, etc.
[0065] External content 400 can be any form of external content, including news articles, web pages, video files, audio files, image files, written descriptions, evaluations, game content, social media content, photos, commercial offers, transportation methods, weather conditions, or other suitable external content. User computing device 100 and server computing system 300 can access external content 400 via network 200. External content 400 can be searched by user computing device 100 and server computing system 300 according to known search methods, and the search results can be ranked according to relevance, popularity, or other suitable attributes (including location-specific filtering or promotion).
[0066] Referring to FIG. 1, in an exemplary embodiment, a user of user computing device 100 can send a request to view a video provided via server computing system 300. For example, the video may be available via a video sharing website, an online video hosting service, a streaming video service, or other video platforms. For example, the user can view the video on display 160 of the user computing device via video application 130 or web browser 140. In response to receiving the request, server computing system 300 is configured to perform, for display on the display of user computing device 100, playing the video in a first portion of the user interface and, if the video is played and a first entity among one or more entities is mentioned in the video, displaying a first entity card in a second portion of the user interface, the first entity card including descriptive content related to the first entity. FIGS. 4A through 7B, which will be discussed in more detail below, provide an exemplary user interface showing the display of the video in the first portion of the user interface and the display of the first entity card in the second portion of the user interface. In some embodiments, server computing system 300 can store or obtain a user interface including the video and the first entity card and provide the user interface in response to a request to view the video from user computing device 100. In some embodiments, server computing system 300 can store or obtain the video and the first entity card related to the video and, in response to a request to view the video from user computing device 100, dynamically generate a user interface including the video and the first entity card and send the user interface to user computing device 100.In some embodiments, the server computing system 300 may store or retrieve videos, may store or retrieve one or more entities identified from the videos, may dynamically generate one or more entity cards from the identified one or more entities, and may dynamically generate a user interface including the video and one or more entity cards (e.g., including the first entity card) in response to a request to view a video from the user computing device 100. In some embodiments, the server computing system 300 may store or retrieve videos, may dynamically identify one or more entities from the videos (e.g., after writing out the video or after obtaining the write-out of the video), may dynamically generate one or more entity cards from the identified one or more entities, and may dynamically generate a user interface including the video and one or more entity cards (e.g., including the first entity card) in response to a request to view a video from the user computing device 100.
[0067] Referring now to FIG. 2, an exemplary block diagram of a user computing device and a server computing system in accordance with one or more exemplary embodiments of the present disclosure is described herein.
[0068] The user computing device 100 can include one or more processors 110, one or more memory devices 120, a video application 130, a web browser 140, an input device 150, and a display 160. The server computing system 300 may include one or more processors 310, one or more memory devices 320, and a user interface generator 330.
[0069] For example, one or more processors 110, 310 may be any suitable processing device that can be included in the user computing device 100 or the server computing system 300. For example, such processors 110, 310 may include one or more of a processor, a processor core, a controller, and an arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field programmable array, a programmable logic unit, an application specific integrated circuit (ASIC), a microprocessor, a microcontroller, etc., and combinations thereof, and any other device capable of responding in a defined manner and executing instructions. The one or more processors 110, 310 may be a single processor, or multiple processors operably connected, for example, in parallel.
[0070] The memories 120, 320 can include one or more non-transitory computer-readable storage media such as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and flash memory, USB drives, volatile memory devices such as random access memory (RAM), hard disks, floppy disks, Blu-ray disks, or optical media such as CD ROMs and DVDs, and combinations thereof. However, the examples of the memories 120, 320 are not limited to the above description, and the memories 120, 320 can be implemented by various other devices and structures as understood by those skilled in the art.
[0071] For example, when executed, the memory 120 enables one or more processors 110 to receive video for playback in the user interface, provide video for display in a first portion of the user interface displayed on a display, and, while the video is being played back, if a first entity is mentioned within the video, provide a first entity card for display in a second portion of the user interface while the video continues to play. For example, the memory 320 can store instructions that, when executed, cause one or more processors 310 to play video in a first portion of the user interface and, if a first entity among one or more entities is mentioned within the video while the video is being played, display a first entity card in a second portion of the user interface, so as to provide a user interface displayed on each display of one or more user computing devices.
[0072] Memory 120 can also include data 122 and instructions 124 that are obtained, manipulated, created, or stored by one or more processors 110. In some exemplary embodiments, such data can be accessed and used as input for receiving video for playback in a user interface, providing video for display in a first portion of a user interface displayed on a display, and providing a first entity card for display in a second portion of the user interface while the video is being played, if a first entity is mentioned in the video while the video is being played. Memory 320 can also include data 322 and instructions 324 that are obtained, manipulated, created, or stored by one or more processors 310. In some exemplary embodiments, such data can be accessed and used as input for playing video in a first portion of a user interface and displaying a first entity card in a second portion of the user interface if a first entity among one or more entities is mentioned in the video while the video is being played, for providing a user interface that is displayed on a respective display of one or more user computing devices.
[0073] In FIG. 2, user computing device 100 includes a video application 130, also referred to as a video player or video streaming app. Video application 130 enables a user of user computing device 100 to view video provided via the application and displayed on the user interface of display 160. For example, the video can be provided to video application 130 via server computing system 300.
[0074] In FIG. 2, the user computing device 100 includes a web browser 140, which may also be referred to as an Internet browser or simply a browser. The web browser 140 may be any browser used to access a website or web page (e.g., via the World Wide Web). A user of the user computing device 100 may provide an input (e.g., a URL) to the web browser 140 to obtain content (e.g., a video) and display the content on the display 160 of the user computing device. For example, the rendering engine of the web browser 140 may display the content on a user interface (e.g., a graphical user interface). For example, the video may be provided or obtained using the web browser 140 via a video sharing website, an online video hosting service, a streaming video service, or other video platforms. For example, the video may be provided to the web browser 140 via the server computing system 300.
[0075] In FIG. 2, the user computing device 100 includes an input device 150 configured to receive input from a user, such as, for example, a keyboard (e.g., a physical keyboard, a virtual keyboard, etc.), a mouse, a joystick, buttons, switches, an electronic pen or stylus, a gesture recognition sensor (e.g., one that recognizes a user's gesture including movement of a body part), an input sound device or a voice recognition sensor (e.g., a microphone that receives voice commands), a trackball, a remote controller, a portable phone (e.g., a mobile phone or a smartphone), a tablet PC, a pedal or a foot switch, a virtual reality device, or one or more of the like. The input device 150 may further include a haptic device that provides haptic feedback to the user. The input device 150 may also be embodied, for example, by a touch-sensitive display having a touch screen function. The input device 150 can be used by a user of the user computing device 100 to provide input for viewing a video, to provide input for selecting user interface elements displayed on a user interface, to input a search query, and the like.
[0076] In FIG. 2, the user computing device 100 includes a display 160 that displays information visible to the user, for example, on a user interface (e.g., a graphical user interface). For example, the display 160 may be a non-touch-sensitive display or a touch-sensitive display. The display 160 may include, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, an active matrix organic light emitting diode (AMOLED), a flexible display, a 3D display, a plasma display panel (PDP), a cathode ray tube (CRT) display, and the like. However, the present disclosure is not limited to these exemplary displays and may include other types of displays.
[0077] According to the exemplary embodiments described herein, the server computing system 300 can include one or more of the aforementioned processors 310 and memory 320. The server computing system 300 may also include a user interface generator 330. For example, the user interface generator 330 may include a video provider 332, an entity card provider 334, and a search query provider 336. The user interface generator 330 can generate a user interface for display on the display 160 of the user computing device 100. The user interface can include various portions for displaying various content. For example, the user interface can display a video in a first portion of the user interface, display an entity card in a second portion of the user interface, and display various user interface elements in other portions of the user interface (e.g., a proposed search query is displayed in a third portion of the user interface).
[0078] The video provider 332 may include information or content (e.g., video content) that can be used to render the user interface so that the video can be displayed and played back on the user computing device 100. In some embodiments, the video provider 332 can be configured to obtain a video from the video data store 370 (e.g., in response to a request from the user computing device 100 for the video).
[0079] The entity card provider 334 can include information or content (e.g., an entity card including a summary of text and / or an image) that can be used to render a user interface so that the entity card can be displayed on the user interface on the user computer device 100 together with a video. In some embodiments, the entity card provider 334 can be configured to obtain an entity card associated with the video from the entity card data store 390.
[0080] The search query provider 336 can include information or content (e.g., a proposed search query) that can be used to render a user interface so that the search query can be displayed on the user interface on the user computer device 100 together with at least one of a video or an entity card. In some embodiments, the search query provider 336 can generate one or more proposed search queries based on entities identified with respect to the video and / or based on the content of the entity card(s) associated with the video. For example, the search query provider 336 can generate proposed search queries in response to viewing the video, based on previous user searches, based on previous user searches regarding the topic of the video, etc.
[0081] Additional aspects of the user computing device 100 and the server computing system 300 are discussed with reference to the following figures shown in FIGS. 3 through 7B, and the flow diagrams of FIGS. 8 through 10.
[0082] FIG. 3 shows an exemplary system 3000 for generating an entity card for a user interface according to one or more exemplary embodiments of the present disclosure. The exemplary system 3000 includes a video 3010, an entity generator 3020, external content 400, an entity card generator 3030, a user interface entity card renderer 3040, and a user interface generator 3050. Each of the video 3010, the entity generator 3020, the entity card generator 3030, the user interface entity card renderer 3040, and the user interface generator 3050 may be part of a server computing system 300. In some embodiments, the video 3010, the entity generator 3020, the entity card generator 3030, the user interface entity card renderer 3040, and the user interface generator 3050 may be part of a server computing system 300 (e.g., as part of a single server or distributed among multiple servers and / or data stores).
[0083] Referring to FIG. 3, an entity card may be generated based on entities identified from the video. For example, the video 3010 may be provided to the entity generator 3020 for one or more entities identified from the video 3010. The video 3010 may be obtained from, for example, a video data store 370 or stored in the video data store 370. The entity generator 3020 may include, for example, a video transcriptase 3022, a signal annotator 3024, and a machine learning resource 3026. In some embodiments, a video provider 332 may provide the video 3010 to the entity generator 3020, or a third-party application or service provider may provide the video 3010.
[0084] The video transcriber 3022 is configured to transcribe the video 3010. For example, the video transcriber 3022 may include a speech recognition program, which analyzes the video 3010, performs automatic speech recognition, and obtains a transcript of the speech (e.g., a text transcript) from the video 3010. In some embodiments, the transcript of the video 3010 may be obtained from the video data store 370 or from a third-party application or service provider that generates the transcript of the video, e.g., by automatic speech recognition, and the entity generator 3020 may not include the video transcriber 3022.
[0085] The signal annotator 3024 is configured to associate some or all of the text from the transcript obtained by the video transcriber 3022 or the video data store 370 with knowledge graph entities and obtain a set of knowledge graph entities associated with the video 3010. A knowledge graph generally can refer to an interconnected and / or interrelated description between objects, events, situations, or concepts, e.g., in the form of a graph-structured data model.
[0086] The machine learning resource 3026 is configured to predict which entity from the video 3010 (e.g., which knowledge graph entity from the video 3010) among a plurality of entities mentioned in or identified from the video is most likely to be searched for by a user viewing the video 3010 (e.g., having a higher probability of being searched than a threshold, having a confidence value greater than a threshold at which the user can perform a search regarding the entity, being determined to be the most likely to be searched for by a user viewing the video as compared to other entities mentioned in the video). The training data for the machine learning resource 3026 may be obtained by identifying or matching those knowledge graph entities that also appear in search queries from actual users viewing the video 3010.
[0087] The machine learning resource 3026 can be configured to identify entities from the video 3010, and for this entity, determine the degree to which the entity is related to other entities within the video 3010, determine the extent to which the entity uses the tf-idf (term frequency-inverse document frequency) signal over a corpus of videos (e.g., stored in the video data store 370), and determine the degree to which the entity is related to the topic of the video (e.g., using signals of prominent terms based on a query), whereby an entity card is generated. For example, the machine learning resource 3026 can select a candidate entity as an entity for which the entity card is to be generated based on one or more of the relevance of each of a plurality of candidate entities to the topic of the video, the relevance of each of the plurality of candidate entities to one or more other candidate entities among the plurality of candidate entities, the number of mentions of the candidate entity within the video, and the number of videos in which the candidate entity is shown over a corpus of videos stored in one or more databases (e.g., the video data store 370).
[0088] The machine learning resource 3026 can be trained by applying weights to candidate entities (for example, the higher the frequency with which a term is mentioned within the video 3010, the higher weight may be assigned to the entity, and a lower weight may be assigned to an entity that is overly broad and frequently appears in the video corpus, and a higher weight may be assigned to an entity more relevant to the topic of the video, etc.). The machine learning resource 3026 can be configured to evaluate candidate entities from among the candidate entities identified within the video 3010 and rank the candidate entities. For example, the entity generator 3020 can be configured to select one or more of the highest-ranked candidate entities identified by the machine learning resource 3026 as the entity for which the entity card is generated. For example, an entity among a plurality of candidate entities mentioned in or identified from the video 3010 can be identified by the machine learning resource 3026 as the entity for which the entity card should be generated when the machine learning resource 3026 determines or predicts that the entity is likely to be searched for by a user viewing the video 3010.
[0089] For example, the machine learning resource 3026 can be updated or adjusted based on an evaluation of user interaction with the user interface provided to the user computing device 100. For example, if a user does not typically interact with a particular entity card while watching video 3010, it can implicitly indicate that the entity is not a term or topic that the user is interested in or understands with respect to the content of video 3010. Thus, the machine learning resource 3026 may be adjusted to reflect the user interaction (or lack thereof) with the user interface. Similarly, if a user typically interacts with a particular entity card while watching video 3010, it can implicitly indicate that the entity is a term or topic that the user is interested in or understands with respect to the content of video 3010. Thus, the machine learning resource 3026 may be adjusted to reflect the user interaction with the user interface. For example, the machine learning resource 3026 can regenerate the entities associated with video 3010 according to a pre-set schedule (e.g., every two weeks), which may result in different entities being identified compared to the previous entities identified by the entity generator 3020. Thus, after the machine learning resource 3026 is updated or adjusted, different entity cards may be displayed in relation to video 3010. The machine learning resource 3026 can also regenerate the entities related to video 3010 according to a user input request to the entity generator 3020.
[0090] Entity generator 3020 may output one or more entities that machine learning resource 3026 determines or predicts are likely to be searched for by a user watching a video to entity card generator 3030. In response to receiving one or more entities from entity generator 30202, entity card generator 3030 may be configured to generate an entity card for each of the one or more entities. For example, entity card generator 3030 may be configured to obtain information about an entity from various sources based on the identity of the entity and the associated metadata of the entity, and input text and / or image(s) that provide more information about the entity into the entity card. For example, entity card generator 3030 may be configured to create or generate an entity card by obtaining information about the entity from various sources (e.g., external content 400) and inputting text and / or image(s) that provide more information about the entity into the entity card. For example, the information about the entity may be obtained from one or more websites, an electronic service that provides a summary of a topic, etc. For example, entity card generator 3030 may store the generated entity card in entity card data store 390. For example, the entity card may be limited to less than a predetermined length and / or size (e.g., less than 100 words). For example, the entity card may include information including a title (e.g., a title that identifies the entity, such as the title "Barack Obama"), a subtitle (e.g., "former President of the United States" for the previous entity and the title example of Barack Obama), and one or more of attribution information (that provides attribution to the source of the information). For example, an image that forms part or all of the entity card may be limited to a thumbnail image size, a specified resolution, etc.In some embodiments, the entity card generator 3030 may correspond to the entity card provider 334.
[0091] The user interface entity card renderer 3040 may be configured to render an entity card provided for at least a portion of the user interface for display on the display 160 of the user computing device 100.
[0092] The user interface generator 3050, which may correspond to the user interface generator 330, may be configured to combine the video 3010 and the rendered entity card to generate a user interface provided for display on the display 160 of the user computing device 100. In some embodiments, the rendering of the entity card, or at least the rendering of the user interface including the video and the entity card, may be performed on the user computing device 100. In some embodiments, the rendering of the entity card, or at least the rendering of the user interface including the video and the entity card, may be performed on the server computing system 300.
[0093] Figures 4A through 4C illustrate exemplary user interfaces in which one or more entity cards are presented during the display of a video, according to one or more exemplary embodiments of the present disclosure. Referring to Figure 4A, an exemplary user interface 4000 displayed on a display 160 of a user computing device 100 is shown. User interface 4000 includes a section in which a video 4010 is being played on a first portion 4012 of user interface 4000. User interface 4000 also includes a section 4020 captioned "In This Video" that is displayed on a second portion 4022 of user interface 4000 that summarizes the content of video 4010 at different points in time during video 4010. User interface 4000 also includes a section 4030 captioned "Related Topics" that includes an entity card 4040 displayed on a third portion 4032 of user interface 4000. In this example, the entity identified from video 4010 is "King Tutankhamuh," and entity card 4040 includes descriptive content related to King Tutankhamuh, including an image of a pharaoh and a text summary about the pharaoh. For example, the entity may be identified by applying machine learning resources to identify one or more entities that are most likely to be searched for by a user viewing video 4010, based on a transcript of the content from video 4010.
[0094] In some embodiments, a user of the user computing device 100 can scroll down the user interface 4000 to view the content shown in the entity card 4040 within the third portion 4032. Referring to FIG. 4B, an exemplary user interface 4000 displayed on the display 160 of the user computing device 100 is shown. For example, as shown in FIG. 4B, in response to the user scrolling down to view the entity card 4040, the video 4010 can be maintained (anchored) at the top of the user interface 4000' on the first portion 4012' of the user interface 4000'. The user interface 4000' also includes a section entitled "Related Topics" that includes the entity card 4040 displayed on the second portion 4022' of the user interface 4000'. In this example, the entity card 4040 includes a title 4042 (King Tutankhamuh) of the entity card 4040, a subtitle 4044 (Pharaoh), descriptive content 4046 (text summary and thumbnail image), and an attribute 4048 that cites the source of the descriptive content 4046.
[0095] For example, in other portions of the user interface 4000', additional user interface elements and entity cards may be provided. For example, a proposed search query 4060 may be provided on the user interface 4000'. Here, the proposed search query 4060 corresponds to an entity search user interface element that is configured to perform a search related to the entity when selected.
[0096] For example, in the third part 4032' of the user interface 4000', a second entity card 4070 (Valley of the Kings) is also displayed. Here, the second entity card 4070 is displayed in a contracted or folded form, in contrast to the entity card 4040 in the unfolded form. The second entity card 4070 displayed in the contracted or folded form includes sufficient identification information (e.g., the title of the entity card such as "Valley of the kings", and a thumbnail image related to the Valley of the Kings) so that the user can understand what theme, concept, or topic the second entity card 4070 is related to. For example, the user can expand the second entity card 4070 by selecting the user interface element 4080 to obtain a more complete description of the second entity (the Valley of the kings), and the user can fold the entity card 4040 by selecting the user interface element 4050.
[0097] For example, in the third part 4032' of the user interface 4000', additional proposed search queries 4090 (e.g., "Howard carter" and "Mummification process") are also displayed. Here, the proposed search query 4090, when selected, corresponds to a search query user interface element configured to perform a search regarding topics of videos 4010 other than the first or second entity (e.g., topics of videos different from any of the entities identified from the video 4010).
[0098] In an exemplary embodiment of the present disclosure, video 4010 continues to play while the user is viewing entity card 4040 and / or second entity card 4070. Thus, the viewing of video 4010 is not interrupted when the user seeks to learn more about an entity identified from video 4010. The user may obtain sufficient information about the entity from the entity card presented on user interface 4000'.
[0099] In some embodiments, the entity may be mentioned in video 4010 at a first point in time within video 4010, and entity card 4040 may be provided for display at a first point in time as part of the user interface (e.g., the third portion 4032 of user interface 4000 or the second portion 4022' of user interface 4000'). Thus, entity card 4040 may be provided for display at a time synchronized with the discussion of the entity within video 4010.
[0100] As shown in FIG. 4B, the second entity card 4070 is displayed on the third portion 4032 of the user interface 4000' in a contracted form. The second entity card 4070 may identify or reference a second entity (Valley of the kings) mentioned within the video 4010 at a second time point within the video 4010 after a first time point. In some embodiments, when the second entity is mentioned within the video 4010 at the second time point, while the video 4010 is being played, the contracted second entity card 4070 may be automatically expanded so as to be fully displayed on the third portion 4032' of the user interface 4000', and the second entity card 4070 includes descriptive content regarding the second entity. In some embodiments, when the second entity is mentioned within the video 4010 at the second time point, while the video 4010 is being played, the contracted second entity card 4070 may be automatically expanded so as to be fully displayed on the second portion 4022' of the user interface 4000', and the second entity card 4070 replaces the entity card 4040 on the user interface 4000' at the time when the second entity is mentioned within the video 4010, and also includes descriptive content related to the second entity.
[0101] Referring to FIG. 4C, an exemplary user interface 4000’’ displayed on the display 160 of the user computing device 100 is shown. The user interface 4000’’ displays a search result page 4092 obtained in response to the user selecting the proposed search query 4060 provided in the user interface 4000’. The proposed search query 4060 corresponds to an entity search user interface element, and the entity search user interface element is configured to perform a search related to an entity when selected. In the example of FIG. 4C, the entity search is for the entity King Tutankhamun, and the search box 4092 has the entity automatically entered, thereby eliminating the need for the user to spell the entity. Entities may be misspelled by the user, resulting in errors in search results, leading to frustration for some users and wasting computing resources in ineffective and error-prone searches.
[0102] As described above with reference to the examples of FIGS. 4A through 4C, in some embodiments, all of the entity cards associated with the video (e.g., entity card 4040 and the second entity card 4070) may be visible to the user on the user interface simultaneously while the video is being played, or only some of the entity cards associated with the video may be visible to the user on the user interface simultaneously while the video is being played. For example, one or more of the entity cards (e.g., entity card 4040 and / or the second entity card 4070) may be fully expanded so that the user can view the entire content of the entity card(s), and some or all of the remaining entity cards may be displayed on the user interface in a collapsed or folded form. For example, in the collapsed or folded form, the user can view a part of the entity card, and the part of the entity card may include some identification information (e.g., identification of the corresponding entity) so that the user can understand the relevance of the entity card. For example, the user can select a user interface or some parts of the visible part of the folded entity card to expand the entity card and view the hidden part of the entity card for more information about the entity. In some embodiments, the entity card can be automatically expanded to fully display the entity card on the user interface when the corresponding entity is mentioned in the video. To save space on the user interface, the previously shown entity card can be changed to a collapsed or folded form when the second entity card is expanded if it is not already in a collapsed or folded form before the second entity card is displayed. In some embodiments, each of all the entity cards can be fully displayed on the user interface throughout the video.
[0103] Figures 5A through 5C show exemplary user interfaces in which entity cards are presented during the display of video, according to one or more exemplary embodiments of the present disclosure. Referring to FIG. 5A, an exemplary user interface 5000 displayed on a display 160 of a user computing device 100 is shown. The user interface 5000 includes a section in which a video 5010 is being played on a first portion 5012 of the user interface 5000. For example, a transcript (i.e., closed caption) 5014 of the audio portion of the video may be provided for display within the video 5010 displayed on the user interface 5000. The user interface 5000 also includes a section 5020 titled "Topics to Explore" displayed in a second portion 5022 of the user interface 5000 that includes one or more entity cards displayed at different points in time during the video 5010. For example, the second portion 5022 may include a first sub-portion 5032 that includes a first entity card 5030 and a second sub-portion 5052 that includes at least a portion of a second entity card 5050.
[0104] In the example of FIG. 5A, the first entity identified from the video 5010 is "Greek mythology," and the first entity card 5030 includes descriptive content related to Greek mythology, including an image related to Greek mythology and a text summary related to Greek mythology. For example, the first entity card 5030 may also include information regarding the point in time (e.g., 8 seconds into the video 5010) within the video 5010 at which the first entity is being discussed. For example, the second entity identified from the video 5010 is "Crete," and the second entity card 5050 includes descriptive content related to Crete, including an image related to Crete and a text summary related to Crete. For example, the second entity card 5050 may also include information regarding the point in time (e.g., 10 seconds into the video 5010) within the video 5010 at which the second entity is being discussed.
[0105] For example, the first sub - portion 5032 may include one or more user interface elements. For example, a proposed search query 5040 may be provided for the first sub - portion 5032. Here, the proposed search query 5040 corresponds to an entity search user interface element, and the entity search user interface element is configured to execute a search related to an entity (e.g., a search for the first entity "Greek mythology") when selected.
[0106] As shown in FIG. 5A, the second entity card 5050 is only partially shown on the user interface 5000 because topics or concepts related to the second entity have not yet been discussed in the video 5010. For example, the entity cards within the second portion 5020 may rotate in a carousel manner, for example, as each entity is being discussed while the video is being played continuously. Thus, the video 5010 continues to play while the user is viewing the first entity card 5030 and the second entity card 5050. Therefore, when the user desires to know more about the entities identified from the video 5010 and obtain sufficient information about the entities from the entity cards provided on the user interface 5000 (or subsequent user interfaces 5000', 5000'', etc.) to understand the content of the video 5010, it is not necessary to interrupt the viewing of the video 5010, nor to perform other searches and / or stop the video 5010.
[0107] Referring to FIG. 5B, an exemplary user interface 5000' displayed on the display 160 of the user computing device 100 is shown. The user interface 5000' includes a section where a video 5010 is being played on a first portion 5012 of the user interface 5000'. The user interface 5000' also includes a section 5020 titled "Topics to Explore" displayed in a second portion 5022 of the user interface 5000', which includes one or more entity cards displayed at different points in time during the video 5010. For example, the second portion 5022 may include a first sub-portion 5032' that includes at least a part of a first entity card 5030, a second sub-portion 5052' that includes a second entity card 5050, and a third sub-portion 5062' that includes at least a part of a third entity card 5060.
[0108] In the example of FIG. 5B, the second entity card 5050 includes descriptive content related to Crete, including an image related to Crete and a text summary about Crete. For example, the second entity card 5050 may also include information about the point in time (e.g., 10 seconds into the video 5010) within the video 5010 when the second entity is being discussed. For example, a third entity identified from the video 5010 is "Icarus", and the third entity card 5060 includes descriptive content related to Icarus, including an image related to Icarus and a text summary about Icarus. For example, the third entity card 5060 may also include information about the point in time (e.g., 13 seconds into the video 5010 as shown in FIG. 5C) within the video 5010 when the third entity is being discussed.
[0109] For example, the second sub - portion 5052’ may include one or more user interface elements. For example, a proposed search query 5040’ may be provided for the second sub - portion 5052’. Here, the proposed search query 5040’ corresponds to an entity search user interface element, and the entity search user interface element is configured to execute a search related to an entity (e.g., a search for the second entity “Crete”) when selected.
[0110] As shown in FIG. 5B, the first entity card 5030 and the third entity card 5060 are only partially shown in the user interface 5000’ because topics or concepts related to the first entity have already been discussed in the video 5010, while topics or concepts related to the third entity have not yet been discussed in the video 5010. For example, the entity cards within the second portion 5020 may rotate in a carousel - like manner as each entity is being discussed while the video is playing continuously. Thus, the video 5010 continues to play while the user is viewing the first entity card 5030, the second entity card 5050, and the third entity card 5060. Therefore, when the user desires to know more about the entities identified from the video 5010 and obtains sufficient information about the entities from the entity cards provided on top of the user interfaces 5000, 5000’, etc. to understand the content of the video 5010, it is not necessary to interrupt the viewing of the video 5010, perform other searches, and / or stop the video 5010.
[0111] Referring to FIG. 5C, an exemplary user interface 5000’ displayed on the display 160 of the user computing device 100 is shown. The user interface 5000’’ includes a section where video 5010 is being played on the first portion 5012 of the user interface 5000’’. The user interface 5000’’ also includes a section 5020 titled “Topics to Explore” displayed in the second portion 5022 of the user interface 5000’’, which includes one or more entity cards displayed at different points in time during the video 5010. For example, the second portion 5022 may include a first sub-portion 5052’’ that includes at least a portion of the second entity card 5050, a second sub-portion 5062’’ that includes the third entity card 5060, and a third sub-portion 5072’’ that includes at least a portion of the fourth entity card 5070.
[0112] In the example of FIG. 5C, the third entity card 5060 includes descriptive content related to Icarus, including an image related to Icarus and a text summary about Icarus. For example, the third entity card 5060 may also include information about the point in time (e.g., 13 seconds into the video 5010) within the video 5010 where the third entity is being discussed. For example, the fourth entity identified from the video 5010 may include “Daedalus”, and the fourth entity card 5070 includes descriptive content related to Daedalus, including an image related to Daedalus and a text summary about Daedalus. For example, the fourth entity card 5070 may also include information about the point in time (e.g., 20 seconds into the video 5010) within the video 5010 where the fourth entity is being discussed.
[0113] For example, the second sub - portion 5062’’ may include one or more user interface elements. For example, a proposed search query 5040’’ may be provided for the second sub - portion 5062’’. Here, the proposed search query 5040’ corresponds to an entity search user interface element, and the entity search user interface element is configured to perform a search related to an entity (e.g., a search for the second entity “Icarus”) when selected.
[0114] As shown in FIG. 5C, the second entity card 5050 and the fourth entity card 5070 are only partially shown in the user interface 5000’’ because topics or concepts related to the second entity have already been discussed in the video 5010, while topics or concepts related to the fourth entity have not yet been discussed in the video 5010. For example, the entity cards within the second portion 5020 may rotate in a carousel manner, for example, while the video is being played and each entity is being discussed. Thus, the video 5010 continues to play while the user views the first entity card 5030, the second entity card 5050, the third entity card 5060, the fourth entity card 5070, etc. Therefore, when the user wishes to know more about the entities identified from the video 5010 and obtains sufficient information about the entities from the entity cards provided on the user interfaces 5000, 5000’, 5000’’, etc. to understand the content of the video 5010, it is not necessary to interrupt the viewing of the video 5010, perform other searches, and / or stop the video 5010.
[0115] As described above with respect to the examples of FIGS. 5A through 5C, in some embodiments, entity cards associated with a video (e.g., entity cards 5030, 5050, 5060, 5070) may be visible to a user on a user interface as the video progresses, and the user may not be able to fully view an entity card until the corresponding entity is mentioned within the video. For example, a first entity card regarding a first entity may be displayed on the user interface at a point in the video (e.g., at a first point in time) when the first entity is mentioned within the video. The first entity card may be displayed for a predetermined amount of time (e.g., an amount of time sufficient for an average user to read or view the content included in the first entity card) while the video continues to play, or until the next entity is mentioned within the video, and at the point at which the next entity is displayed, other entity cards are provided on the user interface. For example, a second entity card regarding a second entity may be displayed on the user interface at a point in the video (e.g., at a second point in time) when the second entity is mentioned within the video. In some embodiments, the second entity card may be displayed on the user interface by replacing the first entity card (i.e., by occupying some or all of the space on the user interface previously occupied by the first entity card).
[0116] Figures 6A through 6C are diagrams showing an exemplary user interface in which notification user interface elements are presented to display one or more entity cards during the display of a video, according to one or more exemplary embodiments of the present disclosure. Referring to FIG. 6A, an exemplary user interface 6000 displayed on a display 160 of a user computing device 100 is shown. User interface 6000 includes a section in which a video 6010 is being played on a first portion 6012 of user interface 6000. User interface 6000 also includes a section 6020 captioned "Related Searches" displayed on a second portion 6022 of user interface 6000, the section including one or more proposed search queries 6030 related to video 6010.
[0117] Referring to FIG. 6B, an exemplary user interface 6000' displayed on the display 160 of the user computing device 100 is shown. The user interface 6000' has a configuration similar to that of the user interface 6000 of FIG. 6A, except that the notification user interface element 6040 is displayed in response to an entity (for which an entity card exists) mentioned in the video 6010. For example, when an entity is mentioned in the video 6010 during playback of the video 6010, the notification user interface element 6040 is displayed on the user interface 6000' while the video 6010 continues to play. For example, the notification user interface element 6040 indicates that additional information related to the entity is available. In response to receiving a selection of the notification user interface element 6040, the entity card is displayed on the user interface while the video 6010 continues to play. The notification user interface element 6040 may include some identification information of the entity (e.g., identification of the corresponding entity, and / or an image such as a thumbnail image) so as to further let the user recognize that the notification user interface element 6040 is associated with the entity, thereby enabling the user to understand the relevance of the available entity card.
[0118] Referring to FIG. 6C, an exemplary user interface 6000'' displayed on the display 160 of the user computing device 100 is shown. For example, the user interface 6000' can be displayed in response to a user selecting a notification user interface element 6040 displayed on the user interface 6000'. The user interface 6000'' includes a section where a video 6010 is being played on a first portion 6012 of the user interface 6000. The user interface 6000 also includes a section 6020 captioned "Related Search" displayed in a second portion 6022 of the user interface 6000'', but the second portion 6022 is obscured by a new section 6050 captioned "Mentioned Topics" overlaid (e.g., as a pop-up window) on the second portion 6022.
[0119] The section 6050 captioned "Mentioned Topics" includes an entity card 6060 overlaid on the second portion 6022. In this example, the entity card 6060 includes a title 6062 (King Tutankhamuh), a subtitle 6064 (Pharaoh), descriptive content 6066 (text summary and thumbnail image), and an attribute 6068 citing the source of the descriptive content 6066.
[0120] For example, in other parts of the section 6050 captioned "Mentioned Topics", additional user interface elements and entity cards can be provided. For example, a proposed search query 6070 can be provided. Here, the proposed search query 6070 corresponds to an entity search user interface element, and the entity search user interface element is configured to execute a search related to the entity when selected. For example, at least a portion of a second entity card 6080 is also provided. The second entity card 6080 can be related to the next entity discussed in the video.
[0121] In some embodiments, the notification user interface element 6040 may be displayed on the user interface 6000' at the same time an entity is mentioned within the video 6010. In some embodiments, the notification user interface element 6040 may be displayed on the user interface 6000' throughout the video 6010, and selection of the notification user interface element 6040 may cause a section 6050 entitled "topics mentioned" to be displayed. For example, the section 6050 entitled "topics mentioned" may be overlaid (e.g., as a pop-up window) on the second portion 6022 and may remain open until (e.g., via the user interface element 6090) it is closed. For example, the second entity card 6080 may be fully displayed (e.g., by replacing the entity card 6060) when the entity associated with the second entity card 6080 is discussed in the video 6010. That is, the display of the entity card may be synchronized with the time when the relevant entity is mentioned within the video. For example, the entity card may be displayed on the user interface only when the relevant entity is first mentioned within the video, or selectively displayed when the entity is mentioned multiple times within the video, each time the relevant entity is mentioned within the video.
[0122] As discussed above with respect to the examples of FIGS. 6A through 6C, in some embodiments, the availability of entity cards associated with a video may be indicated using a notification user interface element while the video is being played, e.g., when the relevant entity is discussed during the video. Thus, the user may have the option to view the entity card while watching the video by deciding to select the notification user interface element.
[0123] Figures 7A and 7B illustrate an exemplary user interface in which a timeline is presented to display one or more entity cards during the display of a video, according to one or more exemplary embodiments of the present disclosure.
[0124] Referring to FIG. 7A, an exemplary user interface 7000 displayed on a display 160 of a user computing device 100 is shown. For example, the user interface 7000 includes a section in which a video 7010 is being played on a first portion 7012 of the user interface 7000. The user interface 7000 also includes a section 7020 titled "Related Searches" displayed on a second portion 7022 of the user interface 7000, which section includes one or more proposed search queries 7030 related to the video 7010. The user interface 7000 further includes a persistent timeline section 7042 overlaid (e.g., as a pop-up window) on the second portion 7022 to hide at least a portion of the second portion 7022.
[0125] The persistent timeline section 7042 includes a persistent timeline 7040, at least a portion of a first entity card 7050, and at least a portion of a second entity card 7060. This persistent timeline 7040 displays the timeline of the video 7010 and may include one or more points 7044 indicating when an entity is discussed during the video 7010. For example, the next entity to be discussed (e.g., Howard Carter) can be indicated by at least a portion of the second entity card 7060 shown within the persistent timeline section 7042 and includes an image of the entity.
[0126] The first entity card 7050 and / or the second entity card 7060 may be selectable such that the entity card is expanded, as shown in FIG. 7B.
[0127] Referring to FIG. 7B, the user interface 7000' may be displayed in response to the user selecting the first entity card 7050 displayed on the user interface 7000, or in some embodiments, the user interface 7000' may be displayed when the entity (King Tutankhamuh) is being discussed in the video 7010. In FIG. 7B, the user interface 7000' includes the video 7010 displayed in the first portion 7012' of the user interface 7000'. The user interface 7000' also includes a section titled "Topics to Explore", which includes at least a portion of the first entity card 7050 and the second entity card 7060 displayed in the second portion 7022' of the user interface 7000'. The user interface 7000' also includes a section 7020 titled "Related Searches" displayed on the third portion 7032' of the user interface 7000', which includes one or more proposed search queries 7030 related to the video 7010.
[0128] As shown in FIG. 7B, the first entity card 7050 is expanded to include descriptive content about the entity and, similar to the previous example, includes a title (King Tutankhamuh), a subtitle (Ancient Egyptian King), descriptive content (text summary and thumbnail), and an attribute citing the source of the descriptive content. The second entity card 7060 may be at least partially displayed next to the first entity card 7050 within the second portion 7022' of the user interface 7000'. The user interface element 7080 may be included in the user interface 7000' as a selectable element, and when selected, the entity cards associated with the video 7010 can be cycled, for example, in a carousel manner as the video is played.
[0129] In some embodiments, the section entitled "Topics to Explore" may also include additional user interface elements. For example, a proposed search query 7070 may be provided. Here, the proposed search query 7070 corresponds to an entity search user interface element, and the entity search user interface element is configured to perform a search related to the entity when selected.
[0130] As discussed above with respect to the examples of FIGS. 7A and 7B, in some embodiments, while a video is being played, a persistent timeline section may be displayed to show the user entity cards available in the video. The entity card related to the entity currently being discussed may be displayed centrally (i.e., prominently) within the persistent timeline section, while a next entity card for the next entity discussed in the video may also be provided. For example, the user may select an entity card displayed in the persistent timeline section (e.g., in a collapsed or folded form) to expand the entity card for further information regarding the entity. In some embodiments, the entity card may be automatically expanded to fully display the entity card on the user interface when the corresponding entity is mentioned within the video.
[0131] FIGS. 8 - 10 show flow diagrams of exemplary non - limiting computer - implemented methods according to one or more exemplary embodiments of the present disclosure.
[0132] Referring to the figure, in FIG. 8, method 800 includes an operation 810 in which a user computing device (e.g., user computing device 100) displays a video on a first portion of a user interface displayed on a display 160 of the user computing device 100. In operation 820, if a first entity is mentioned in the video while the video is being played, the user interface displays a first entity card on a second portion of the user interface while continuing to play the video, and the first entity card includes descriptive content regarding the first entity. For example, the first entity is generated in response to an automatic recognition of the first entity from a transcript of the video content.
[0133] Referring to FIG. 9, method 900 includes operation 910 of a server computing system (e.g., server computing system 300) that obtains a transcript of content from a video (e.g., via automatic speech recognition). In operation 920, the method includes applying machine learning resources to identify one or more entities that are most likely to be searched for by a user viewing the video based on the transcript of the content. For example, in a video about ancient Egypt, the identified entities may include "valley of the kings", "King Tutankhamuh", and "sarcophagus". In operation 930, the method is to generate one or more entity cards for each of the one or more entities, each of the one or more entity cards including descriptive content about the respective entity among the one or more entities. In operation 940, the method includes providing (or generating) a user interface that is displayed on a respective display of one or more user computing devices for playing the video in a first portion of the user interface and, when the video is played and a first entity among the one or more entities is mentioned in the video, displaying the first entity card in a second portion of the user interface, the first entity card including descriptive content related to the first entity.
[0134] Referring to FIG. 10, method 1000 includes operation 1010 of a server computing system (e.g., server computing system 300) that obtains a transcript of content from a video (e.g., via automatic speech recognition). In operation 1020, the method includes associating text from the transcript with a knowledge graph to obtain a set of knowledge graph entities. In operation 1030, the method includes obtaining training data for building a machine learning model for machine learning resources by identifying (or matching) the knowledge graph entities with entities that appear in an actual search query of a user watching the video. In operation 1040, the method includes weighting the entities based on, for example, the relevance of the entity to other entities in the video, the breadth of the entity, the relevance of the entity to the topic of the video, etc. For example, the machine learning resources may be trained by applying weights to candidate entities (e.g., the higher the frequency with which a term is mentioned in the video, the higher the weight that may be assigned to the entity, and a lower weight may be assigned to entities that are overly broad and frequently appear in the video corpus, and a higher weight may be assigned to entities that are more relevant to the topic of the video, etc.). In operation 1050, the method includes applying the machine learning resources to evaluate candidate entities from among the candidate entities identified in the video and rank the candidate entities. In operation 1060, the method includes the machine learning resources predicting one or more entities for which entity cards are generated for the video. For example, the machine learning resources may select a predetermined number of entities (e.g., three or four) that are the highest-ranked candidate entities as the entities for which entity cards are generated. For example, the machine learning resources may select these entities that are predicted to be searched for by the user (e.g., with a specified confidence, with a probability of being searched for exceeding a threshold level, etc.).The identified entity may then be provided to an entity card generator to generate an entity card, and / or the identified entity may be stored in a database (e.g., entity data store 380).
[0135] Terms such as "module", "unit", "provider", and "generator" may be used herein in connection with various features of the present disclosure. Such terms may refer to software or hardware components or devices such as, but not limited to, a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC) that performs a particular task. A module or unit may be configured to exist on an addressable storage medium or may be configured to execute on one or more processors. Thus, a module or unit may include, by way of example, components such as software components, object oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functions provided by the components and modules / units may be combined into fewer components and modules / units or further separated into additional components and modules.
[0136] Aspects of the exemplary embodiments described above may be recorded on a computer-readable medium (e.g., a non-transitory computer-readable medium) including program instructions for performing various operations implemented by a computer. The medium may also include program instructions, data files, data structures, etc., alone or in combination. Examples of non-transitory computer-readable media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD ROM disks, Blu-Ray disks, and DVDs, magneto-optical media such as optical disks, and other hardware devices specially configured to store and execute program instructions, such as semiconductor memories, read-only memories (ROMs), random access memories (RAMs), flash memories, USB memories, etc. Examples of program instructions include both machine code such as that generated by a compiler and files containing higher-level code that can be executed by a computer using an interpreter. The program instructions may be executed by one or more processors. The hardware devices described may be configured to function as one or more software modules for performing the operations of the above embodiments, or vice versa. Further, the non-transitory computer-readable storage medium may be distributed among computer systems connected via a network, and the computer-readable code or program instructions may be stored and executed in a decentralized manner. Further, the non-transitory computer-readable storage medium may also be embodied in at least one application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA).
[0137] Each block of the flowchart illustrations may represent a unit, module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions represented by the blocks may occur in a different order. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the functionality contained.
[0138] Although the present disclosure has been described with respect to various exemplary embodiments, each example has been provided for illustrative purposes only and is not intended to limit the present disclosure. Those of ordinary skill in the art will readily be able to make changes, modifications, and equivalents to such embodiments upon reaching the foregoing understanding. Accordingly, the present disclosure does not exclude including such modifications, variations, and / or additions to the subject matter of the disclosed invention as would be readily apparent to those of ordinary skill in the art. For example, features illustrated or described as part of one embodiment may be used in other embodiments to create still other embodiments. Accordingly, the present disclosure is intended to cover such changes, modifications, and equivalents.
Claims
1. A computer-implemented method for a server system, comprising: obtaining a transcript of content from a video; applying machine learning resources to identify one or more entities that are most likely to be searched for by a user viewing the video, based on the transcript of the content; generating, for each of the one or more entities, one or more entity cards, each of the one or more entity cards including descriptive content regarding the respective entity among the one or more entities; providing a user interface to be displayed on a respective display of one or more user computing devices, playing the video in a first portion of the user interface; when the video is played and a first entity among the one or more entities is mentioned within the video, displaying a first entity card in a second portion of the user interface, the first entity card including descriptive content associated with the first entity; providing the user interface for; A computer-implemented method comprising.
2. Applying the machine learning resources to identify the one or more entities includes obtaining training data to train the machine learning resources based on observational data of users who perform searches in response to viewing only the video. The computer-implemented method according to claim 1, comprising.
3. Applying the machine learning resources to identify the one or more entities includes identifying a plurality of candidate entities from the video by associating text from the transcript with a knowledge graph; the relevance of each of the candidate entities to the topic of the video; the relevance of each of the candidate entities to one or more other candidate entities among the plurality of candidate entities; the number of mentions of the candidate entities within the video; and the number of videos in which the candidate entities appear across a corpus of videos stored in one or more databases. Ranking the candidate entities to obtain the one or more entities based on one or more of the same; The computer-implemented method according to claim 2, further comprising.
4. Applying the machine learning resource to identify the one or more entities, Evaluating a user interaction with the user interface; Determining at least one adjustment to the machine learning resource based on the evaluation of the user interaction with the user interface; The computer-implemented method according to claim 3, further comprising.
5. The first entity is mentioned in the video at a first time point within the video, The first entity card is displayed in the second portion of the user interface at the first time point. The computer-implemented method according to claim 1.
6. The one or more entities include a second entity, and the one or more entity cards include a second entity card, The method comprising providing the user interface displayed on the respective displays of the one or more user computing devices, While continuing to play the video, displaying the second entity card in a collapsed form in a third portion of the user interface, the collapsed second entity card referring to the second entity mentioned in the video at a second time point after the first time point; displaying; When the second entity is mentioned in the video at the second time point, while continuing to play the video, displaying the second entity card in a fully expanded form in the third portion of the user interface, the fully expanded second entity card including descriptive content regarding the second entity; displaying; The computer-implemented method according to claim 5, further comprising providing the user interface for.
7. The one or more entities include a second entity, and the one or more entity cards include a second entity card, providing the user interface that is displayed on respective displays of the one or more user computing devices, when the second entity is mentioned in the video while the video is being played, displaying the second entity card in the second portion of the user interface while the video is being played, the second entity card including descriptive content regarding the second entity, and the second entity card being displayed in the second portion of the display by replacing the first entity card at the time when the second entity is mentioned in the video The computer-implemented method according to claim 1, further comprising providing the user interface for this purpose.
8. providing the user interface that is displayed on respective displays of the one or more user computing devices, when the first entity is mentioned in the video while the video is being played, displaying a notification user interface element in a third portion of the user interface while the video is being played, the notification user interface element indicating that additional information related to the first entity is available in response to the first entity being mentioned in the video while the video is being played and in response to receiving a selection of the notification user interface element, displaying the first entity card in the second portion of the user interface while the video is being played The computer-implemented method according to claim 1, further comprising providing the user interface for this purpose.
9. The computer-implemented method according to claim 1, wherein the first entity card includes at least one of a text summary providing information related to the first entity or an image related to the first entity.
10. A computer-implemented method for a user computing device, Receiving a video for playback in a user interface; Providing the video for display in a first portion of the user interface that is displayed on a display of the user computing device; When a first entity is mentioned in the video while the video is being played; Providing a first entity card for display in a second portion of the user interface while the video is continuing to be played; comprising; wherein the first entity card includes descriptive content related to the first entity; A computer-implemented method, wherein the first entity is generated in response to an automatic recognition of the first entity from a transcript of the content of the video.
11. wherein the first entity is mentioned in the video at a first point in time within the video; The computer-implemented method according to claim 10, wherein the first entity card is provided for display in the second portion of the user interface at the first point in time.
12. Providing a shrunk second entity card that references a second entity mentioned in the video at a second point in time within the video after the first point in time for display in a third portion of the user interface; When the second entity is mentioned in the video at the second point in time, expanding the shrunk second entity card to fully display the second entity card in the third portion of the user interface while the video is continuing to be played, wherein the second entity card includes descriptive content related to the second entity; The computer-implemented method according to claim 11, further comprising.
13. When a second entity is mentioned in the video while the video is being played, further comprising providing a second entity card for display in the second portion of the user interface while the video is continuing to be played, wherein the second entity card includes descriptive content related to the second entity; The computer-implemented method according to claim 10, wherein the second entity card is provided for display in the second part of the user interface by replacing the first entity card at the time when the second entity is mentioned in the video.
14. The computer-implemented method according to claim 10, further comprising providing one or more entity search user interface elements for display on the user interface, which are configured to perform a search regarding the first entity when selected.
15. The computer-implemented method according to claim 14, further comprising providing one or more search query user interface elements for display on the user interface, which are configured to perform a search regarding a topic of the video other than the first entity when selected.
16. The computer-implemented method according to claim 10, further comprising using machine learning resources to identify the first entity and generate the first entity card.
17. The computer-implemented method according to claim 16, wherein the first entity is an entity among a plurality of entities mentioned in the video, and is determined by the machine learning resources as the entity that is most likely to be searched by a user viewing the video among the plurality of entities mentioned in the video.
18. When the first entity is mentioned in the video while the video is being played, providing a notification user interface element for display in a third part of the user interface while the video continues to be played, wherein the notification user interface element indicates that additional information related to the first entity is available. In response to receiving a selection of the notification user interface element, while continuing to play the video, providing the first entity card for display in the second portion of the user interface; The computer-implemented method according to claim 10, further comprising.
19. The computer-implemented method according to claim 10, wherein the first entity card includes a text summary providing information related to the first entity and / or an image related to the first entity.
20. A user computing device, A display, One or more memories storing instructions, One or more processors for executing the instructions stored in the one or more memories, Receiving a video for playback in a user interface, Providing the video for display in a first portion of the user interface displayed on the display, When a first entity is mentioned within the video while the video is being played, While continuing to play the video, providing a first entity card for display in a second portion of the user interface; One or more processors for performing; and comprising, The first entity card includes descriptive content related to the first entity, The first entity is generated in response to an automatic recognition of the first entity from a transcript of the content of the video. A user computing device.
Citation Information
Patent Citations
Contextual digital media processing system and method
JP2021523480A
Using automated content analysis for audio / video content consumption
US20080140385A1
Queue to Display Additional Information for Entities in Captions
US20150095938A1