Pairing media items with products

A model-based approach for automatic product detection in media content items addresses the challenges of identifying and purchasing associated products, enhancing user experience and system efficiency.

JP2025525887APending Publication Date: 2025-08-07GOOGLE LLC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025505868
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-01
Filing Date
2023-07-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Conventional systems face difficulties in efficiently identifying, accessing, and purchasing products associated with media content items, leading to time-consuming and resource-intensive user experiences.

Method used

Implementing a model-based approach to automatically identify products within media content items, using machine learning to enhance product detection from text and images, and providing intuitive user interfaces for seamless product information and purchasing.

Benefits of technology

Enhances user experience by reducing time and computing resources required to discover and purchase associated products, improving accuracy and efficiency in product identification and information access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025525887000001_ABST
    Figure 2025525887000001_ABST
Patent Text Reader

Abstract

The method includes presenting a user interface (UI) including one or more graphical representations of one or more videos. Each graphical representation of the respective videos is selectable to initiate playback of the respective video. The graphical representations are displayed with a UI element in a collapsed state. The UI element includes information identifying a plurality of products covered by the respective videos. The method further includes changing the presentation of the UI element from the collapsed state to an expanded state. The expanded UI element includes a plurality of visual components each associated with one of the plurality of products. The method further includes initiating presentation of the video upon a user selection of one of the plurality of visual components.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Aspects and embodiments of the present disclosure relate to methods and systems for facilitating pairing of media items with related objects, and more particularly to systems for facilitating pairing of media items with products. [Background technology]

[0002] A platform (e.g., a content sharing platform) can transmit (e.g., stream) media items over a network to client devices connected to the platform. Different types of client devices may be optimized for different tasks, preferred by users for different tasks, etc. A media item may include references to one or more products. Summary of the Invention

[0003] The following summary is a simplified summary of the disclosure in order to provide a basic understanding of some aspects of the disclosure. This summary is not an exhaustive overview of the disclosure. It is not intended to identify key or critical elements of the disclosure or to delineate the scope or claims of particular embodiments of the disclosure. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.

[0004] Systems and methods are disclosed for facilitating user participation in one or more engagement activities related to one or more products of media items. In some implementations, the method includes presenting a user interface (UI) including one or more graphical representations of one or more videos. Each graphical representation of the respective videos is selectable to initiate playback of the respective video and is displayed with a collapsed UI element. The collapsed UI element includes information identifying multiple products covered by the respective video. The method further includes continuing to present the graphical representations of the respective videos while changing the display of the UI element from a collapsed state to an expanded state. The change in display of the UI element is performed in response to user interaction with the collapsed UI element. The expanded UI element includes multiple visual components, each associated with one of the multiple products. The method further includes initiating playback of a video covering the product associated with the selected visual component in response to a user selection of a respective one of the multiple visual components of the expanded UI element.

[0005] In some embodiments, the presentation of the UI element may be changed to a product-focused state. The product-focused state may include detailed information about one or more products associated with the content item. The presentation of the UI element may be changed to the product-focused state in response to a user selection of a UI element or a UI element component. The presentation of the UI element may be changed to the product-focused state in response to a user selection of one of multiple visual components of a UI element in an expanded state.

[0006] In some embodiments, the presentation of the UI elements and graphical representations may be performed in response to a user device requesting a home feed of content items for display, the presentation of the UI elements and graphical representations may be performed in response to a search query of the user device, or the presentation of the UI elements and graphical representations may be performed in response to a user request to view a content item.

[0007] The expanded UI element may include one or more tabs. A first tab may be associated with multiple products. A second tab may be associated with multiple portions (e.g., chapters) of the respective content item. The first tab may be presented by default.

[0008] In some embodiments, the presentation of the UI element may change to a transaction state. The change to the transaction state may be performed in response to a user selection of the UI element or a component of the UI element. The UI element in the transaction state may facilitate a transaction related to one of a plurality of products.

[0009] One or more UI elements may be overlaid. The UI elements may be overlaid on a content item being presented. The UI elements may be overlaid on a graphical representation of the content item. User interaction with the UI elements may include hovering over the UI elements.

[0010] In another aspect, a method includes providing one or more graphical representations of one or more content items to a device for display by a UI of the device. Each graphical representation of the respective content items is selectable to initiate presentation of the respective content item. The method further includes providing instructions to the device to display a UI element in a collapsed state along with a first graphical representation of a first content item. The UI element may include information identifying multiple products covered by the first content item. The method further includes providing instructions to the device to modify the presentation of the UI element. The instructions may be provided in response to receiving notification of a user interaction with the UI element. The instructions may facilitate transitioning the UI element from the collapsed state to an expanded state. The expanded UI element includes multiple visual components. Each of the visual components is associated with one of the multiple products. The method further includes providing instructions to the device to facilitate presentation of the content items. The instructions to facilitate presentation of the content items may be provided in response to receiving notification of a user selection of one of the multiple visual components of the expanded UI element.

[0011] In some embodiments, the method includes providing instructions to a device that facilitate changing a presentation of a UI element from an expanded state to a product-focused state. The change of the UI element can be performed in response to receiving notification of a user selection of one of a plurality of visual components of the UI element in the expanded state. The UI element in the product-focused state can include detailed information about a product associated with the selected visual component.

[0012] In some embodiments, the method includes obtaining identifiers of a plurality of products covered by the first content item as output from a trained machine learning model, the trained machine learning model receiving data related to the first content item as input.

[0013] In some embodiments, the method includes providing instructions to a device that facilitate a change in presentation of UI elements to a transaction state. The transaction state includes one or more components that facilitate a transaction related to one of a plurality of products associated with the content item. The instructions that facilitate the change in presentation of UI elements to the transaction state can be provided in response to a user selection of a component of the UI elements associated with the product.

[0014] In some embodiments, the instructions facilitating the presentation of the content item may include notification of a portion of the content item to present. The portion of the content item may cover one of a plurality of products associated with one of a plurality of visual components. In some embodiments, one or more instructions may be provided to the device in response to obtaining a user history. The user history may include one or more interactions with the content item related to the product. In some embodiments, the content item is a live streaming video. In some embodiments, the content item is a short-form video.

[0015] In another aspect, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform operations including presenting a UI. The UI includes one or more graphical representations of one or more videos. Each graphical representation of the respective videos is selectable to start playing the respective video and is displayed with a collapsed UI element. The UI element includes information identifying multiple products covered by the respective videos. The operations further include continuing to present the graphical representations of the respective videos while changing the display of the UI element from a collapsed state to an expanded state. The change is performed in response to user interaction with the collapsed UI element. The expanded UI element may include multiple visual components. Each of the visual components may be associated with one of the multiple products. The operations further include starting playback of the respective videos covering the products associated with the selected visual component. Initiating playback may be performed in response to a user selection of one of the multiple visual components.

[0016] In some embodiments, the action includes changing the presentation of the UI element from an expanded state to a product-focused state. The change may be performed in response to a user selection of one of a plurality of visual components of the expanded UI element. The product-focused UI element may include detailed information about a product associated with the selected visual component.

[0017] In some embodiments, the operations further include changing the presentation of the UI element to a transaction state. Changing the presentation of the UI element may be performed in response to a user selection of a component of the UI element associated with the product. The UI element in the transaction state facilitates a transaction associated with the product.

[0018] In some embodiments, the UI element may include a tab. The UI element may include a first tab associated with a plurality of products. The UI element may include a second tab associated with a plurality of portions of the respective videos. The content of the first tab may be presented by default. In some embodiments, initiating playback of the respective videos includes presenting a portion of the video depicting the product. The product may be associated with a selected visual component of the UI element.

[0019] Optional features of one aspect may be combined with other aspects where appropriate.

[0020] Aspects and embodiments of the present disclosure will be more fully understood from the following detailed description of various aspects and embodiments of the disclosure and from the accompanying drawings, which should not be construed to limit the disclosure to any particular aspect or embodiment, but are for purposes of illustration and understanding only. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 illustrates an exemplary system architecture for providing related and associated product information, according to some embodiments. [Figure 2] FIG. 2 is a block diagram of a system 200 including a dataset generator for creating datasets for one or more models, according to some embodiments. [Figure 3A] 1 is a block diagram illustrating a system for generating output data, such as associations between products and content items, according to some embodiments. [Figure 3B] 1 is a block diagram of an example system for generating data describing associations between content items and one or more products, according to some embodiments. [Figure 4A] FIG. 1 illustrates a device presenting an exemplary user interface (UI) including UI elements showing products associated with a content item, according to some embodiments. [Figure 4B] FIG. 1 illustrates a device presenting an exemplary UI including UI elements showing products associated with a content item, according to some embodiments. [Figure 4C] FIG. 1 illustrates a device presenting an exemplary UI including a UI element presenting a content item and a UI element presenting information about a product related to the content item, according to some embodiments. [Figure 4D] FIG. 1 illustrates a device presenting an exemplary UI including UI elements that present content items and UI elements that facilitate transactions related to products, according to some embodiments. [Figure 4E] FIG. 1 illustrates an exemplary device having UI elements overlaid on content presentation elements, according to some embodiments. [Figure 5A] FIG. 1 is a flow diagram of a method for generating a dataset for a machine learning model, according to some embodiments. [Figure 5B] 1 is a flow diagram of a method for updating metadata of a content item according to some embodiments. [Figure 5C] FIG. 1 is a flow diagram of a method for training a machine learning model related to content item and product pairings, according to some embodiments. [Figure 5D] 1 is a flow diagram of a method for adjusting metadata associated with a content item, according to some embodiments. [Figure 5E] FIG. 1 is a flow diagram of a method for presenting UI elements related to one or more products, according to some embodiments. [Figure 5F] FIG. 1 is a flow diagram of a method for instructing a device to present one or more UI elements related to a product, according to some embodiments. [Figure 6] FIG. 1 is a block diagram illustrating an exemplary computer system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0022] Aspects of the present disclosure relate to methods and systems for facilitating pairing of media items (e.g., content items, content, etc.) with related products. A platform (e.g., a content sharing platform, etc.) may be used to enable users to access media items (e.g., video items, audio items, etc.) hosted by the platform (e.g., via client devices connected to the platform). The platform may provide the user's client device with access to the media items over a network (e.g., the Internet) (e.g., by transmitting the media items to the user's client device). A media / content item may have one or more additional associated activities that can deepen a user's engagement with the content item. For example, engagement activities may include a comment section for the content item, a live chat associated with the content item, etc. Some content items may be associated with one or more products. For example, a product may be displayed as part of the content item, a product may be reviewed within the content item, a content item may be associated with one or more products (e.g., sponsored by a company related to the products), etc.

[0023] In conventional systems, identifying products associated with a content item can be difficult, inconvenient, time-consuming, etc. Content platforms (e.g., platforms that provide content for presentation to users) can have difficulty identifying products associated with a content item. In some systems, a content creator may identify one or more products associated with a content item, a content channel, a list of content items, etc. In some systems, the content creator may include information about one or more products in a content item, e.g., a photo of the item, an item displayed in a video, etc. In some systems, the content creator may include information about one or more products in one or more fields associated with the content item, e.g., a title of the content item, a description of the content item, comments about the content item (e.g., pinned comments), etc.

[0024] In conventional systems, a user may not be able to easily see whether a content item has one or more related products. The content item may not have an indicator of the related products. The user may need to be presented with the content item (e.g., watch a video) to see whether the content item has one or more related products.

[0025] In conventional systems, users may have difficulty identifying products associated with a content item. There may be no immediate indication (e.g., while selecting the content item to consume, view, etc.) that a content item has associated products. Product information may be difficult to find, e.g., scattered among different fields such as the content item title, the content item description, the comments section, etc. Product information may be included in a content item; for example, a video may include audio describing one or more products, an image item may include images of one or more products, etc. Extracting this information from a content item may be difficult, time consuming, error-prone, etc.

[0026] In conventional systems, receiving additional information about a product featured in a content item can be difficult, time-consuming, etc. In some systems, a product may be displayed in a content item (e.g., it may be displayed in a video). A user may not be provided with additional information (e.g., product name, distributor name, etc.) and may perform a search independent of the content platform to learn more about the product. A name and / or distributor associated with the product may be provided (e.g., in the description or title of the content item). A user can perform a search (e.g., independent of the content platform) to learn more information about the product, such as product variations, related products, availability, and prices. Guidance (e.g., user guidance, instructions to a processing device in the form of a link to a distributor's website, etc.) may be provided to facilitate learning more about the product. A user may also receive additional information from another source (e.g., a website) independent of the content platform.

[0027] In conventional systems, information about products included in content items can become outdated. For example, within a content item or associated fields (e.g., content item title, content item description, content item comments, etc.), a content creator may include additional information about one or more products, such as pricing information, distributor information, availability, information about alternative versions or variations, etc. Some information included in a content item or associated fields may not be updated as this information changes or may be dependent on updates from the content creator.

[0028] Conventional systems may have obstacles that make it difficult, time-consuming, tedious, or otherwise cumbersome for a user to purchase one or more items related to a content item. A user may search for a product or search for a merchant that stocks or sells the product. In some embodiments, the content item or an associated field may include instructions on how to purchase the item (e.g., a description of a content item may include one or more links to products related to the content item). A user may be directed to another platform independent of the content platform to complete the purchase of one or more products.

[0029] It may take a significant amount of time and computing resources for a user to find information about products covered by a content item. For example, a video may be long, and the product of interest may be featured near the end. It may take a significant amount of time for a user to consume the video to obtain accurate information about the product of interest, thereby increasing the usage of computing resources on the client device. Furthermore, the computing resources of the client device used to enable a user to consume a media item may not be available for other processes, potentially reducing the overall efficiency of the client device and increasing overall latency.

[0030] Aspects of the present disclosure may address one or more of these shortcomings of conventional methods. In some embodiments, aspects of the present disclosure may enable automatic identification of products featured and / or included in a content item. In some embodiments, aspects of the present disclosure may enable use of a model to identify products from text associated with a content item. The text may include a title of the content item, a description of the content item, a caption associated with the content item, etc. The caption may be a machine-generated caption, for example, generated by a speech-to-text model for a video or audio content item. The model may be a machine-learning model. The model may output a confidence value indicating the likelihood that the content item contains the product.

[0031] In some embodiments, aspects of the present disclosure may enable the use of a model to identify a product from an image of a content item. The content item may be or include a photograph. The content item may be or include a video. One or more images of the content item may be provided to a model configured to identify a product from an image. The model may reduce the number of dimensions of the image. The model may search for similar images related to the product. The model may search for similar reduced-dimensional images within the reduced-dimensional space. The model may be a machine learning model. The model may output a confidence value indicating the likelihood that the content item contains the product.

[0032] In some embodiments, aspects of the present disclosure enable the use of a model (e.g., a fusion model) to verify whether a product is included in a content item. The fusion model may receive notification of one or more products detected by a model that receives text related to the content item as input. The fusion model may receive notification of a confidence value that one or more products are included in the text. The fusion model may receive notification of one or more products detected by a model that receives an image of the content item as input. The fusion model may receive notification of a confidence value that one or more products are included in the image. The fusion model may determine a confidence level that one or more products appear in the content item and associated data (e.g., title, description, etc.). The fusion model may be a machine learning model.

[0033] In some embodiments, product detection may be utilized to enhance content or related information (e.g., descriptions, captions). The model may be provided with data related to the content item. For example, the model may be provided with machine-generated captions for the content item. The model may identify one or more captions that may be misrepresentations of the product name (e.g., the machine-generated captions may include the closest English equivalent to the spoken product name). The model may be provided with one or more images of the content item. The model may determine the likelihood that one or more products are included in the image. Information related to the content item (e.g., metadata, descriptions, captions, etc.) may be updated to account for the one or more detected products.

[0034] In some embodiments, aspects of the present disclosure enable indicators of content items that have one or more associated products. A list of content items may include one or more indicators that one or more of the content items in the list include associated products. The indicators may include a visual indicator (e.g., a "shopping" symbol or text indicating one or more products displayed in association with the content item), additional fields (e.g., a panel containing product information), etc. In some embodiments, the list of content items may be presented to a user via a user interface (UI). The UI element may be associated with the content platform (e.g., presented via an application associated with the content providing platform). The UI may include an element that indicates that the content item is associated with one or more products. In some embodiments, user interaction with the UI element may present an additional UI element with additional information, for example, a list of products related to the content item.

[0035] In some embodiments, aspects of the present disclosure enable a user to identify products related to a content item. In some embodiments, a list of content items for presentation to a user may include a list of products related to one or more content items. For example, the list of content items may be displayed to a user via a UI. The UI may include one or more UI elements that present one or more products related to the content item to the user. The UI elements may be provided via an application associated with the content item's content provision platform. In some embodiments, a UI element listing products associated with a content item may be presented in response to detecting a user interaction with a UI element that indicates the content item has related products. In some embodiments, user interaction with a product on the list of products may cause the UI to present additional information to the user.

[0036] In some embodiments, aspects of the present disclosure enable providing a user with additional information about one or more products associated with a content item. UI elements that provide additional information about the products may be provided to the user. The UI elements may provide information such as product variations (e.g., color variations, size variations, etc.), related products, availability and / or price (e.g., associated with one or more retailers). The UI elements may be provided via an application associated with the content item, a content providing platform, etc. UI elements that display additional product information may be presented in response to user interaction with other UI elements, e.g., selecting a content item for viewing, selecting a UI element that includes one or more related products, etc. In some embodiments, upon user interaction with the UI elements, further UI elements may be provided, e.g., to facilitate purchasing the product.

[0037] In some embodiments, aspects of the present disclosure enable automatic updates of information linked to one or more products associated with a content item. In some embodiments, a content providing platform associated with a content item may include, communicate with, be connected to, etc., one or more memory devices containing product data. For example, the content platform may maintain and update a database of information related to the products, and changes made to the database may be reflected in UI elements presented to a user.

[0038] In some embodiments, aspects of the present disclosure enable a simplified purchasing procedure for a user. Upon receiving notification of a user's intent to purchase a product (e.g., upon user interaction with a UI element associated with the product), UI elements may be presented to the user to facilitate the purchase of the product. In some embodiments, the user may purchase the product through an application, content providing platform, etc., associated with the content item. In some embodiments, the user may be directed to one or more external merchants (e.g., a merchant application, a merchant webpage, etc.) to purchase the product.

[0039] In some embodiments, product-related UI elements may be provided in response to various operations of an application (e.g., an application associated with a content platform). Product-related UI elements may be provided as part of a list of content items tailored to a user, e.g., tailored to a user account associated with the user. Product-related UI elements may be provided as part of a list of content items related to previously presented content items, e.g., a watch next list, a recommended video list, etc. Product-related UI elements may be provided as part of a list of content items generated in response to a user search. Product-related UI elements may be provided as part of a product-specific list, e.g., a shopping section of an application associated with a content platform. Inclusion of product-related UI elements, product-related content items, the number of UI elements or content items presented in association with a product, etc. may be based on several metrics. Metrics may include user viewing history, user search history, etc.

[0040] In some embodiments, aspects of the present disclosure may enable quick access to one or more portions of a content item related to a product. For example, a UI element may include a list of products associated with the content item. One or more of the lists of products may be associated with portions of the content item, e.g., a timestamp of a video content item. Upon interaction with a portion of the UI element related to a product, the content item may present the relevant portion of the content item (e.g., a video may begin playing a portion of the video associated with the content item). In some embodiments, the list of products associated with a content item may be updated as the content item is presented. For example, the list of products associated with a video may be rearranged as the video plays. For example, a product currently highlighted by the video may be at the top of the list of products, products currently on screen may be grouped together in the product presentation UI element, etc.

[0041] Aspects of the present disclosure may provide technical advantages over previous solutions. Aspects of the present disclosure may enable automatic product detection within a content item, which can improve the content creator's experience by automatically associating one or more products with the content item (e.g., removing the burden of associating products with the content item from the creator). The automatically generated product associations may have improved accuracy due to the use of multiple sources (e.g., product text search and image search), fusion models, etc. Model-based product detection may be used to improve the content item, associated data, etc. For example, the use of object detection may improve the machine-generated caption or description of a content item. More accurate captions can improve a user's experience when consuming a content item. These improvements may reduce the time required for content creators to generate accurate content, reduce the time users spend discovering content items and / or products of interest, increase the accuracy of machine-generated information related to the content item, etc. Thus, computing resources on content creators, content viewers, and / or client devices associated with the platform are reduced and available for other processes, thereby improving overall system efficiency and reducing overall latency.

[0042] Aspects of the present disclosure may improve a user's experience when searching for content items, viewing content items, scrolling through lists of content items, etc. A user may be able to identify a content item that has related products from a list of content items, for example, without being presented with the content item. A user may be provided with a seamless way to increase engagement with products associated with a content item. For example, interacting with a UI element that indicates that a content item has products associated with it may present a content item containing more information about the one or more products, further interaction may prompt a purchase of the one or more products, etc. The presentation of one or more UI elements may streamline the user's experience. A user may be able to easily obtain additional information, for example, within an application associated with the content platform / content item. A user may be able to more easily purchase a product associated with a content item. A user may be directed to relevant portions of a content item based on an interest they have expressed in a product associated with the content item. A user may be able to easily view pricing information, availability, product variations, related products, etc., within the context of a single application. Such implementations may allow for saving users time and frustration, simplifying the shopping and / or purchasing process, simplifying the product search, review, and / or selection process, and the like.

[0043] 1 illustrates an exemplary system architecture 100 for providing content and related product information, according to some embodiments. The system architecture 100 includes a client device 110, one or more networks 105, a content platform system 102, and a product identification system 175. The content platform system 102 includes one or more server machines 106, one or more data stores 140, and may include various platforms for performing tasks (e.g., tasks related to content delivery). The content platform system 102 platform may be hosted by one or more server machines 106. The content platform system 102 platform may include and / or be hosted on one or more computing devices (e.g., rack-mounted servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.) and one or more data stores (e.g., hard disks, memory, and databases), and may be coupled to one or more networks 105. In some embodiments, components of the content platform system 102 (e.g., server machines 106, data store 140, hardware associated with one or more platforms, etc.) may be directly connected to one or more networks 105. In some embodiments, one or more components of the content platform system 102 may access the network 105 through other devices, e.g., hubs, switches, etc. In some embodiments, one or more components of the content platform system 102 may communicate directly with other components shown in FIG. 1 , e.g., components of product identification system 175, such as server machines 170 and / or 180. Data store 140 may be included in one or more server machines 106, may include external data storage, etc.The platform of the content platform system 102 may include an advertising platform 165, a social network platform 160, a recommendation platform 157, a search platform 145, and a content providing platform 120. The product identification system 175 includes a server machine 170, a server machine 180, and a set of models 190. The product identification system 175 may include additional devices, such as a data store, additional servers, etc. Multiple operations of the product identification system 175 may be performed by a single physical or virtual device.

[0044] The one or more networks 105 may include one or more public networks (e.g., the Internet), one or more private networks (e.g., a local area network (LAN), a wide area network (WAN), one or more wired networks (e.g., an Ethernet network), one or more wireless networks (e.g., an 802.11 network), one or more cellular networks (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or combinations thereof. In one embodiment, some components of architecture 100 are not directly connected to each other. In one embodiment, system architecture 100 includes separate networks 105.

[0045] The one or more data stores 140 may reside in memory (e.g., random access memory), cache, drive (e.g., hard drive), flash drive, etc., or may be part of one or more database systems, one or more file systems, or other types of components or devices capable of storing data. The one or more data stores 140 may include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers). A data store may be persistent storage on which data can be stored. The persistent storage may be a local or remote storage unit, an electronic storage unit (e.g., main memory), or similar storage unit. The persistent storage may be a monolithic device or a distributed set of devices.

[0046] Content items 121A-121C (e.g., media content items) may be stored in one or more data stores. The data stores may be part of one or more platforms. Examples of content items 121 may include, but are not limited to, digital videos, digital movies, animated images, digital photos, digital music, digital audio, digital video games, collaborative media content presentations, website content, social media updates, e-books, e-journals, digital audiobooks, web blogs, software applications, etc. Content items 121A-121C may also be referred to as media items. Content items 121A-121C may be pre-recorded or live-streamed. For brevity and simplicity, video may be used throughout this document as an example of content item 121 (e.g., content item 121A). Video may include pre-recorded video, live-streaming video, short-form video, etc.

[0047] Content items 121A-121C may be provided by a content provider. The content provider may be a user, a company, an organization, etc. The content provider may provide content item 121 (e.g., content item 121A) that is a video. The content provider may provide content item 121 that includes live streamed content; for example, content item 121 may include a live streaming video, a live chat related to the video, etc.

[0048] Client devices 110 may include devices such as televisions, smartphones, personal digital assistants, portable media players, laptop computers, e-book readers, tablet computers, desktop computers, gaming consoles, set-top boxes, and the like.

[0049] The client device 110 may include a communications application 115. Content items 121 (e.g., content item 121A) may be consumed by a user via the communications application 115. For example, the communications application 115 may access one or more networks 105 (e.g., the Internet) via the hardware of the client device 110 to provide the content items 121 to the user. As used herein, “media,” “media item,” “online media item,” “digital media,” “digital media item,” “content,” “media content item,” and “content item” may include electronic files that can be executed or loaded using software, firmware, and / or hardware configured to present the content item. In one embodiment, the communications application 115 may be an application that enables a user to create, send, and receive content items 121 (e.g., videos) over a platform (e.g., the content providing platform 120, the recommendation platform 157, the social networking platform 160, and / or the search platform 145), and / or a combination of platforms and / or networks.

[0050] In some embodiments, the communication application 115 may be (or may include aspects of) a social networking application, a video sharing application, a video streaming application, a video game streaming application, a photo sharing application, a chat application, or a combination of such applications. The communication application 115 associated with the client device 110 may render, display, present, and / or play one or more content items 121 to one or more users. For example, the communication application 115 may provide a user interface 116 (e.g., a graphical user interface) displayed on the endpoint device 110 for receiving and / or playing video content. In some embodiments, the communication application 115 is associated with (and managed by) the content platform 102 or the content providing platform 120.

[0051] In some embodiments, the communication application 115 may include a content viewer 113 and a related products component 114. A user interface 116 (UI) may display the content viewer 113 and the related products component 114. The related products component 114 may be used to display a UI element that displays information about one or more products (e.g., one or more products related to a content item). The related products component 114 may display a UI element to inform the user that a content item has one or more related products (e.g., the related products component 114 may display a “shopping” symbol displayed near the content viewer 113; the related products component 114 may display an element near an element for selecting a content item to present from a list of content items indicating that the content item has one or more related products, etc.). The related products component 114 may cause the UI to display information about one or more products associated with the content item (e.g., the related products component 114 may display a UI that may include a list of products related to the content item, images of the products related to the content item, pricing information, or information connecting the products to the content item, such as a timestamp of a video segment related to the products). The related products component 114 may cause the UI element to display additional information about one or more products, such as product variations (e.g., color variations, size variations, etc.), related products, recommended products, etc. The related products component 114 may cause the UI element to display one or more options for purchasing one or more products. In some embodiments, a user may be able to navigate between different views provided by the related products component 114. Further description of exemplary UI elements related to products associated with a content item may be found in relation to FIGS. 4A through 4E . In some embodiments, the related products component 114 may cause more than one UI element to be displayed, for example, if two or more content items with related products are displayed.In some embodiments, the related products component 114 may suppress the display of UI elements that indicate related products based on, for example, user settings, user preferences, user history, etc. In some embodiments, multiple content viewers 113 and / or related products components 114 may be associated with a single user interface, communication application, client device, etc. For example, multiple content items may be displayed to a user at one time. In some implementations, the communication application 115 is a web browser that can access, retrieve, present, and / or navigate content (e.g., web pages, such as HyperText Markup Language (HTML) pages, digital media items, etc.), and may include the related products component 114 and the content viewer 113. The content viewer 113 may be an embedded media player embedded in a user interface 116 (e.g., a web page associated with viewing content) provided by the content providing platform 120. Alternatively, the application 115 is not a web browser, but rather a standalone application (e.g., a mobile application, a desktop application, a gaming console application, a television application, etc.) that is downloaded from a platform (e.g., the content providing platform 120, the recommendation platform 157, the social networking platform 160, or the search platform 145) or is pre-installed on the client device 110. The standalone application 115 can provide a user interface 116 that includes a content viewer 113 (e.g., an embedded media player) and associated product components 114.

[0052] In some embodiments, the content platform system 102 may include a product information platform 161 (e.g., hosted by the server machine 106). The product information platform 161 may store, retrieve, provide, receive, etc., data related to one or more products associated with one or more content items. The content platform system 102 may provide the data to the related products component 114 of the client device 110. The product information platform 161 may include information provided by a content creator, e.g., a content creator may provide a list of products related to content items that the content creator provided to the content platform system 102. The product information platform 161 may also include information provided by one or more users, e.g., one or more users may identify products associated with a content item, e.g., in response to being presented with the content item. The product information platform 161 may include information provided by a product identification system 175, e.g., one or more machine learning models may be utilized to identify products featured in a content item and provide notifications of related products and content items to the product information platform 161.

[0053] In some embodiments, a communication application 115 installed on a client device 110 may be associated with a user account, e.g., a user may be signed in to an account on the client device 110. In some embodiments, multiple client devices 110 may be associated with the same client account. In some embodiments, the provision of information regarding product association(s) with one or more content items may be conditioned on the user account, e.g., account settings, account history (e.g., history of interaction with UI elements that include associated product information), etc.

[0054] In some embodiments, client device 110 may include one or more data stores. The data stores may include commands (e.g., instructions that, when executed by a processing device, cause an action) for rendering a UI (e.g., user interface 116). The instructions may include interactive components, e.g., instructions for rendering UI elements with which a user can interact to present additional information about one or more products associated with a content item. In some embodiments, these instructions may cause the processing device to render UI elements that present information about one or more products associated with one or more content items (e.g., several videos reviewing products in which a user is interested may be presented to the user along with UI elements that present more information about the products).

[0055] In some embodiments, one or more server machines 106 may include computing devices such as rack-mounted servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc., and may be coupled to one or more networks 105. The server machines 106 may be standalone devices or part of any platform (e.g., content serving platform 120, social networking platform 160, etc.).

[0056] Social network platform 160 may provide online social networking services. Social networking platform 160 may offer communication applications 115 for users to create profiles and perform activities with their profiles. Activities may include updating profiles, exchanging messages with other users, rating (e.g., liking, commenting, sharing, recommending) status updates, photos, videos, etc., and receiving notifications related to other users' activities. In some embodiments, additional product information (e.g., provided by product information platform 161) may be shared by users to one or more other users via social network platform 160.

[0057] The recommendation platform 157 may be used to generate and provide content recommendations (e.g., articles, videos, posts, news, games, etc.). Recommendations may be based on search history, content consumption history, followed / subscribed channel content, linked profiles (e.g., friends list), popular content, etc. The recommendation platform 157 may be utilized to generate, for example, a user's home feed, a user's watchlist, a user's playlist, etc. One or more UI elements that notify users of related products, display a list of related products, present product information, present one or more options for purchasing a product, etc. may be presented in combination with, as part of, associated with, accessible from, etc. the home feed, watchlist, playlist, watchnext list, etc. Presentation of the one or more UI elements may be based on user history, user settings, data indicative of user preferences (e.g., demographic data), etc.

[0058] Using the search platform 145, a user may be able to query one or more data stores 140 and / or one or more platforms and receive query results. The search platform 145 may be utilized by a user to search for content items, search for a topic, etc. For example, the search platform 145 may be utilized by a user to search for content items having one or more related products. The search platform 145 may be utilized to search for content items related to a type of product (e.g., headphone review videos). The search platform 145 may be utilized to search for content items related to a specific product (e.g., a specific make and / or model of headphones). In response to receiving a search query, one or more UI elements may be displayed to the user. The type, style, etc. of the displayed UI element may be based on the content of the user's search. For example, in response to searching for content related to a type of product (e.g., headphones), a UI element may be displayed indicating that the content items suggested as search results have one or more related products. As a further example, in response to searching for content items related to a more specific product (e.g., "best headphones for podcasts"), a different UI element may be displayed to provide information about products related to the content item. As a further example, in response to a search for content items related to a particular product (e.g., a particular make and / or model), different UI elements may be displayed to provide specific information about the searched product and indicate that the product is related to the content item recommended as a search result.

[0059] The content serving platform 120 may be used to provide one or more users with access to the content items 121 and / or to provide the content items 121 to one or more users. For example, the content serving platform 120 may allow users to consume, upload, download, and / or search for the content items 121. In other examples, the content serving platform 120 may allow users to rate the content items 121, such as approving ("liking"), disliking, recommending, sharing, rating, and / or commenting on the content items 121. In other examples, the content serving platform 120 may allow users to edit the content items 121. The content serving platform 120 may also include a website (e.g., one or more web pages) and / or one or more applications (e.g., a communication application 115) that may be used to provide one or more users with access to the content items 121. For example, the communication application 115 may be used by the client device 110 to access the content items 121. The content providing platform 120 may include any type of content distribution network that provides access to the content items 121 .

[0060] The content providing platform 120 may include multiple channels (e.g., channel A 125, channel B 126, etc.). A channel may be a collection of content available from a common source, a collection of content having a common topic or subject, etc. The data content may be digital content selected by a user, digital content made available by a user, digital content uploaded by a user, digital content selected by a content provider, content selected by a broadcaster, etc. For example, channel A 125 may include two videos (e.g., content items 121A and 121B). A channel may be associated with an owner, which may be a user who can perform actions on the channel. The content may be one or more content items 121. The data content of a channel may be pre-recorded content, live content, etc., although a channel is described as one embodiment of a content providing platform, and embodiments of the present disclosure are not limited to content sharing platforms that provide content items 121 via a channel model.

[0061] Each of product identification system 175, server machine 170, and server machine 180 may include one or more computing devices such as a rack-mounted server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application-specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc. Operations of prediction server 112, server machine 170, server machine 180, data store 140, etc. may be performed by a cloud computing service, a cloud data storage service, etc.

[0062] Product identification system 175 may include one or more models 190. Models 190 included in product identification system 175 may perform tasks related to identifying one or more products from content items. One or more of models 190 may be trained machine learning models. Operations for generating trained machine learning models, including model training, validation, and testing, are described in connection with FIGS. 3A and 5C.

[0063] Models 190 may include one or more text analysis models 191. Text analysis models 191 may be configured to receive text as input and generate one or more notifications of products associated with the text as output. For example, a first one of text analysis models 191 may be configured to predict related products from a title of a content item, a second one of text analysis models 191 may be configured to predict related products from a (e.g., written) description of a content item, a third one of text analysis models 191 may be configured to predict related products from a caption of a content item (e.g., an auto-generated caption, a machine-generated caption, a user-provided caption, etc.), etc. In some embodiments, all of the operations of text analysis models 191 may be performed by a single model. In some embodiments, a model of text analysis models 191 may be configured to generate, as output, product context information, e.g., information indicating that a content item is associated with one or more products. For example, the product context information may indicate that a content item includes a category of products, e.g., a group of different products (e.g., a type of product, a brand of product, a category of products such as “electronics,” etc.).

[0064] Model 190 may include one or more image analysis models 192. Image analysis model 192 may be configured to identify a product from one or more images. Image analysis model 192 may include one or more models aimed at identifying whether an image contains a product, a model configured to separate product images from content item images (e.g., by removing background elements), a model configured to determine the identity of a product within a content item image, etc. The operations of image analysis model 192 may be performed by a single model. Image analysis model 192 may include one or more models configured to provide images to a product identification model. For example, image analysis model 192 may include a model configured to extract a portion of a still image of a content item, a model configured to extract one or more frames of a video content item, etc. Models of image analysis model 192 may be provided with one or more frames of a video, one or more portions of one or more frames of a video, etc., and may generate as output one or more products and one or more confidence values associated with the one or more products. For example, a model in image analysis models 192 may receive as input one or more frames of a video content item and generate as output a list of products with confidence values indicating the likelihood that the product is included in the image of the content item. Image analysis models 192 may include one or more models that determine which images from a content item to utilize. For example, image analysis models 192 may include one or more models that select frames from a video content item for image identification.

[0065] Image analysis models 192 may include one or more models configured to reduce the dimensionality of an image. For example, an image may be reduced to a vector of values. In some embodiments, one or more models of image analysis models 192 may be configured to reduce the dimensionality of an image in a way that similar images (e.g., images of the same product) may be represented similarly (e.g., by similar vectors) after dimensionality reduction. One or more models of image analysis models 192 may be configured to compare a dimensionally reduced image from a content item (e.g., a vector of values generated from one or more frames of a content item video) with dimensionally reduced images of known products (e.g., via product information platform 161).

[0066] Models 190 may include text correction model 193. Text correction model 193 may be configured to provide corrections to text associated with a content item. Text correction model 193 may be configured to adjust text associated with a content item to include one or more products referenced in the content item. One or more models in text correction model 193 may be configured to adjust computer-generated, machine-generated, automatically-generated, etc. text associated with a content item. One or more models in text correction model 193 may be configured to update captions of a video (e.g., inaccurate captions) to include one or more products. In some embodiments, machine-generated text (e.g., captions) associated with a content item may be incorrect. For example, the name of a product may be replaced with an approximation when generating a caption (e.g., the name of the product may not be a word in the language of the caption, the name of the product may be a word in a language different from the language of the caption, etc.). Models in text correction model 193 may be configured to identify portions of text that may be incorrect and recommend corrections, make the corrections, alert a user or other system, etc. For example, a model in text correction model 193 may receive a machine-generated video caption, identify portions of the caption that may incorrectly replace words in the caption's language for a product name, provide data indicating the potentially inaccurate text to other models, systems, users, etc.

[0067] Model 190 may include fusion model 194. Fusion model 194 may receive as input one or more indications of products associated with a content item. In some embodiments, fusion model 194 receives as input the output from one or more other models (e.g., text analysis model 191, image analysis model 192, etc.). Fusion model 194 may receive one or more indications of products and one or more indications of confidence values associated with a content item. For example, fusion model 194 may receive indications of one or more products detected in the title of a content item and confidence values associated with the one or more products. Fusion model 194 may further receive indications of one or more products detected in the description of a content item and confidence values associated with the one or more products. Fusion model 194 may further receive indications of one or more products detected in the caption of a content item and confidence values associated with the one or more products. Fusion model 194 may further receive indications of one or more products detected in an image of a content item and confidence values associated with the one or more products. The fusion model 195 may generate as output one or more products detected in association with the content item. The fusion model 195 may further generate a confidence value associated with the confidence that the product appears in the content item, is associated with the content item, etc. Further actions may be performed based on the output of the fusion model 195 (e.g., product identification information and confidence value) (e.g., a UI element describing the product associated with the content item may be presented).

[0068] One type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. Artificial neural networks generally include a feature representation component with a classifier or recurrent layer that maps features to a desired output space. For example, a convolutional neural network (CNN) hosts multiple layers of convolutional filters. Pooling can be performed to address nonlinearities in lower layers, and a multi-layer perceptron is typically added on top to map the features extracted by the convolutional layers to a decision (e.g., a classification output).

[0069] A recurrent neural network (RNN) is another type of machine learning model. A recurrent neural network model is designed to interpret a series of inputs where the inputs are intrinsically related to each other, e.g., time trace data, continuous data, etc. The output of the perceptron in the RNN is fed back as input to the perceptron to generate the next output.

[0070] Deep learning is a type of machine learning algorithm that uses a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Deep neural networks contain a hierarchy of layers, with different layers learning different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a slightly more abstract and complex representation. For example, in an image recognition application, the raw input may be a matrix of pixels; the first representation layer may abstract the pixels and encode edges; the second layer may construct and encode the edge configuration; the third layer may encode higher-level shapes (e.g., teeth, lips, gums, etc.); and the fourth layer may recognize the role of the scan. Notably, the deep learning process can learn on its own which features are best placed at which levels. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, deep learning systems have a significant depth of credit assignment pathways (CAPs). A CAP is a chain of transformations from input to output. A CAP describes the potential causal relationships between input and output. For feedforward neural networks, the depth of the CAP may be the depth of the network or the number of hidden layers plus one. For recurrent neural networks, where a signal can propagate through layers multiple times, the depth of the CAP is potentially infinite.

[0071] In some embodiments, product identification system 175 further includes server machine 170 and server machine 180. Server machine 170 includes a dataset generator 172 that can generate datasets (e.g., a set of data inputs and a set of target outputs) to train, validate, and / or test model(s) 190, including one or more machine learning models. Some operations of dataset generator 172 are described in more detail below with respect to FIGS. 2 and 5A . In some embodiments, dataset generator 172 may divide historical data (e.g., existing content item data, content items with one or more specified associated products, content items with product allocations provided by one or more users, etc.) into a training set (e.g., 60 percent of the historical data), a validation set (e.g., 20 percent of the historical data), and a test set (e.g., 20 percent of the historical data).

[0072] In some embodiments, components of product identification system 175 may generate multiple sets of features. For example, the features may be rearrangements of the input data, combinations of the input data, reductions in the dimensionality of the input data, subsets of the input data, etc. One or more data sets may be generated based on one or more features of the input data.

[0073] Server machine 180 includes a training engine 182, a validation engine 184, a selection engine 185, and / or a test engine 186. The engines (e.g., training engine 182, validation engine 184, selection engine 185, and test engine 186) may relate to hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing device, etc.), software (e.g., instructions executed on a processing device, general-purpose computer system, or dedicated machine), firmware, microcode, or a combination thereof. Training engine 182 may be capable of training one or more models 190 using one or more sets of features associated with a training set from dataset generator 172. Training engine 182 may generate multiple trained models 190, each trained model 190 corresponding to a distinct set of features in the training set. The dataset generator 172 receives the output of the trained models (e.g., the fusion model 194 may be trained based on the output of the text analysis model 191 and / or the image analysis model 192) and sorts the data into training, validation, and testing datasets that may be used to train a second model (e.g., the fusion model 194).

[0074] The validation engine 184 may be capable of validating the trained models 190 using the corresponding set of features of the validation set from the dataset generator 172. For example, a first trained machine learning model 190 trained using the first set of features of the training set may be validated using the first set of features of the validation set. The validation engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of features of the validation set. The validation engine 184 may discard trained models 190 with an accuracy that does not meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting one or more trained models 190 with an accuracy that meets the threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting the trained model 190 with the highest accuracy of the trained models 190.

[0075] The test engine 186 may be capable of testing the trained model 190 using a corresponding set of features of a test set from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of features of the training set may be tested using a first set of features of the test set. The test engine 186 may determine the trained model 190 with the highest accuracy among all the trained models based on the test set.

[0076] In the case of a machine learning model, model 190 may refer to a model artifact created by training engine 182 using a training set that includes data inputs and corresponding target outputs (correct answers for each training input). Patterns in a dataset that map data inputs to target outputs (correct answers) can be found, and machine learning model 190 is provided with a mapping that captures these patterns. Machine learning model 190 may use one or more of the following: support vector machine (SVM), radial basis function (RBF), clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, k-nearest neighbor algorithm (k-NN), linear regression, random forest, decision forest, neural network (e.g., artificial neural network, recurrent neural network), linear model, function-based model (e.g., NG3 model), etc. Synthetic data generator 174 may include one or more machine learning models that may include one or more of the same type of model (e.g., artificial neural network).

[0077] Automatic (e.g., model-based) detection of products from content items and related data offers significant technical advantages over other methods. In some embodiments, a content item featuring a product (e.g., a product review video reviewing a product) may be linked to or associated with the product (e.g., data linking the product and the content item may be generated) without the content creator's attention, action, time, etc. In some embodiments, a content item promoting a product (e.g., a content item may be sponsored or may promote one or more products) may be linked to or associated with the product. Model-based detection of products within a content item can generate product associations for products present in the content item that are not specifically featured in the content item (e.g., products that a user may be interested in purchasing may be displayed on-screen within a video content item). Model-based detection of products within a content item may generate associations for products advertised within the content item. Based on model-based detection, a user may be directed to products present in the content item by providing a UI element to the user, e.g., indicating that the product is associated with the content item.

[0078] One or more models 190 may be run on the inputs to generate one or more outputs. The models may determine (e.g., extract) confidence data from the outputs that indicates a level of confidence that the model's output is an accurate depiction of the content item. For example, the model may determine that a first product is associated with the content item and determine a confidence that the first product was correctly discovered by the model in the content item. One or more components of the product identification system 175 may use the confidence data to determine whether to update data associated with the content item, e.g., whether to associate one or more products with the content item, whether to update one or more captions for the content item, etc.

[0079] The confidence data may include or indicate a level of confidence that the model's output (e.g., one or more products) is an accurate indication of a product associated with the content item. For example, the level of confidence output by the model (e.g., with respect to a product identified in the content item) may be a real number between 0 and 1. 0 may indicate no confidence that the predicted product is associated with the content item, while 1 may indicate absolute confidence that the predicted product is associated with the content item. In response to the confidence data indicating a level of confidence below a threshold level for a predetermined number of instances (e.g., a percentage of instances, a frequency of instances, a total number of instances, etc.), the product identification system 175 may re-train one or more trained models 190 (e.g., based on updated and / or new data for training, validation, testing, etc.). Re-training may include generating one or more datasets (e.g., via the dataset generator 172).

[0080] By way of example and not limitation, aspects of the present disclosure describe methods for training one or more machine learning models 190 using historical data and inputting current data (e.g., newly updated content items, content items not previously associated with products, etc.) into the one or more trained machine learning models to determine an output indicative of a content item-product association. In other embodiments, a heuristic model, a physics-based model, or a rule-based model is used to determine that one or more products are associated with a content item (e.g., without using a trained machine learning model). In some embodiments, such models may be trained using historical data. In some embodiments, these models may be retrained utilizing historical data. Any of the information described with respect to data input 210 in FIG. 2 may be monitored or otherwise used in a heuristic model, a physics-based model, or a rule-based model.

[0081] In some embodiments, the functionality of client device 110, product identification system 175, content platform system 102, server machine 170, server machine 180, and server machine 106 may be provided by fewer machines. For example, in some embodiments, server machines 170 and 180 may be combined into a single machine, while in some other embodiments, server machine 170, server machine 180, and server machine 106 may be combined into a single machine. In some embodiments, client device 110 and server machine 106 may be combined into a single machine. In some embodiments, the functionality of client device 110, server machine 106, server machine 170, server machine 180, and data store 140 may be performed by a cloud-based service.

[0082] In general, functions described in one embodiment as being performed by client device 110, server machine 106, server machine 170, and server machine 180 may also be performed by server machine 106 in other embodiments, where appropriate. In addition, functionality assigned to a particular component may be performed collaboratively by different or multiple components. For example, in some embodiments, product identification system 175 may determine associations between products and content items. In other examples, content platform system 102 may determine associations between content items and one or more products.

[0083] Additionally, the functionality of a particular component may be performed by different or multiple components in cooperation. One or more of server machine 106, server machine 170, or server machine 180 may be accessed as a service offered to other systems or devices via an appropriate application programming interface (API).

[0084] In embodiments of the present disclosure, a "user" may be represented as a single individual. However, other embodiments of the present disclosure encompass a "user" that is a set of users and / or an entity controlled by an automated source. For example, a set of individual users integrated as a community within a social network may be considered a "user." In another example, an automated consumer may be an automated ingestion pipeline, such as a topic channel, for one or more platforms, one or more content items, etc. In addition to the above, users may be provided with controls that allow them to choose both whether and when the systems, programs, or features described herein may enable collection of user information (e.g., information regarding the user's social network, social actions, or activities, occupation, user preferences, or the user's current location) and whether content or communications are sent to the user from a server. Furthermore, certain data may be processed in one or more ways to remove personally identifiable information before being stored or used. For example, a user's identifying information may be processed so that personally identifiable information cannot be determined, or if location information is obtained, the user's geographic location may be generalized (such as to the city, zip code, or state level) so that the user's specific location cannot be determined. Thus, users have control over what information is collected about them, how that information is used, and what information is provided to them.

[0085] 2 is a block diagram of a system 200 including a dataset generator 272 for creating datasets for one or more models, according to some embodiments. The dataset generator 272 may create a dataset (e.g., data inputs 210, target outputs 220) using historical data. A dataset generator similar to the dataset generator 272 may be utilized to train an unsupervised machine learning model, e.g., the target outputs 220 may not be generated by the dataset generator 272. A dataset generator similar to the dataset generator 272 may be utilized to train a semi-supervised machine learning model, e.g., the target outputs 220 corresponding to a subset of the data inputs 210 may be generated by the dataset generator 272.

[0086] The dataset generator 272 may generate datasets for training, testing, and validating models. The dataset generator 272 may generate datasets for machine learning models. The system 200 may generate datasets for training, testing, and / or validating a fusion model, for example, to determine the likelihood that one or more products will appear in a content item. Systems similar to the system 200 may generate datasets for training, testing, and / or validating models having different functions, with corresponding changes to the input data and / or output data included in the datasets. Models for text analysis (e.g., extracting one or more references to products from text associated with a content item), image analysis (e.g., extracting one or more references to products from images associated with a content item), text correction (e.g., including one or more references to products in machine-generated text associated with a content item), etc. may have datasets for training, testing, and / or validating models generated by a dataset generator similar to the dataset generator 272.

[0087] In some embodiments, a dataset generator, such as dataset generator 272, may be associated with two or more separate models (e.g., the dataset may be used to train an ensemble model). For example, an input dataset may be provided to a first model, an output of the first model may be provided to a second model, and a target output may be provided to the second model to train, test, and / or validate the first and second models (e.g., an ensemble model).

[0088] The dataset generator 272 may generate one or more datasets to provide to the model, for example, during training, validation, and / or testing operations. A machine learning model may be provided with a set of historical data. A machine learning model (e.g., a fusion model) may be provided with sets of historical text analysis data 264A-264Z as data input. The text analysis data may be provided by the machine learning model and may include, for example, one or more products and confidence values identified by the machine learning model on text associated with a content item. A machine learning model may be provided with sets of historical image analysis data 265A-265Z as data input. The image analysis data may be provided by the machine learning model and may include, for example, one or more products and confidence values identified by a trained machine learning model on one or more images associated with a content item.

[0089] In some embodiments, dataset generator 272 may be configured to generate datasets for training, text propagation, validation, etc. of a fusion model. A dataset generator similar to dataset generator 272 may generate a set of text data (e.g., title text of a content item, description text of a content item, caption text of a content item, etc.) as data input for training a machine learning model to determine one or more products associated with the content item. A dataset generator similar to dataset generator 272 may generate a set of image data (e.g., one or more frames from a video content item, portions of one or more frames from a video content item, etc.) as data input for training a machine learning model to determine one or more products associated with the content item.

[0090] In some embodiments, the dataset generator 272 generates a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210 (e.g., training inputs, validation inputs, test inputs). The data inputs 210 may be provided to the training engine 182, the validation engine 184, or the test engine 186 of FIG. 1. The dataset may be used to train, validate, or test a model (e.g., a fusion model, a text analysis model, an image analysis model, etc.). The dataset generator 272 may generate a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210. The data inputs 210 may be referred to as "features," "attributes," "vectors," or "information."

[0091] In some embodiments, the dataset generator 272 can generate a first data input corresponding to the first set of historical text analysis data 264A and / or the first set of historical image analysis data 265A to train, validate, or test a first machine learning model. The dataset generator 272 can generate a second data input corresponding to the second set of historical metrology data 264B and / or the second set of design rule data 265B to train, validate, or test a second machine learning model. Some embodiments of generating training sets, test sets, validation sets, etc. are further described with respect to FIG. 5A .

[0092] In some embodiments, the dataset generator 272 can generate target outputs 220 that provide one or more machine learning models for training, testing, validation, etc. The dataset generator 272 may generate product association data 268 as the target output 220. The product association data 268 may include identifiers of one or more products associated with a content item (e.g., human-labeled product associations). The product association data 268 may include an input-output mapping, e.g., the set of historical text analysis data 264A may be associated with the first set of product association data 268, etc. The machine learning model may be updated (e.g., trained) by providing input data, generating outputs, and comparing the outputs to provided target outputs (e.g., “ground truths”). Various weights, biases, etc. of the model are then updated so that the model more closely matches the training data. This process may be repeated multiple times to generate a model that provides accurate outputs for a threshold portion of the provided inputs. The target outputs 220 may share one or more characteristics of the data inputs 210, for example, the target outputs 220 may be organized into attributes or vectors, the target outputs 220 may be organized into sets AZ, etc.

[0093] In some embodiments, a dataset generator similar to dataset generator 272 may be utilized in conjunction with a text analysis model configured to determine one or more product associations with a content item. Product associations may include contextual associations such as product brands, product types, product classifications, etc. The dataset generator may generate, as a target output, a list of products, product types, product brands, etc. associated with a content item, text of the content item, etc. A dataset generator similar to dataset generator 272 may be utilized in conjunction with an image analysis model configured to determine one or more product associations with a content item. The dataset generator may generate, as a target output, a list of products, product classifications, product categories, etc. associated with an image input. A dataset generator similar to dataset generator 272 may be utilized in conjunction with a text correction model. The text correction model may be configured to recognize one or more words in a target language that are erroneously inserted in place of a product name in machine-generated text. The text correction model may be provided with one or more sets of machine-generated text as input data and with products associated with the text (e.g., products that were not correctly captured by the machine-generation of the text) as a target output.

[0094] In some embodiments, after generating a dataset and using the dataset to train, validate, or test a machine learning model, the model may be further trained, validated, or tested, or tuned (e.g., by adjusting weights or parameters associated with the model's inputs, such as connection weights in a neural network). The model may be tuned and / or retrained based on data that differs from the original training operation, for example, data generated after training, validating, and / or testing the model.

[0095] 3A is a block diagram illustrating a system 300A for generating output data (e.g., product / content item association data) according to some embodiments. System 300A may be used in conjunction with a fusion model to generate product / content item association data and confidence data based on potential product / content item associations generated by other models (e.g., a text-based model that detects products in text associated with a content item, an image-based model that detects products from images associated with a content item, etc.). Systems similar to system 300A may also be used to generate output data from other types of models, such as text analysis models, image analysis models, text correction models, etc.

[0096] At block 310, system 300A (e.g., a component of product identification system 175 of FIG. 1 ) performs data splitting (e.g., via dataset generator 272 of FIG. 2 ) of data used in training, validating, and / or testing machine learning models. In some embodiments, training data 364 includes historical data, such as historical associations between products and content items based on text, historical associations between products and content items based on images, etc. In some embodiments, for example, when system 300A is used to generate output from a fusion model, training data 364 may include data generated by one or more trained machine learning models, e.g., models configured to detect products in text or images associated with content items. Training data 364 may undergo data splitting at block 310 to generate training set 302, validation set 304, and test set 306. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data.

[0097] The generation of the training set 302, validation set 304, and test set 306 may be tailored for a particular application. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data. System 300A may generate multiple sets of features for each of the training set, validation set, and test set. For example, if training data 364 includes product associations extracted from text data from two or more text sources (e.g., titles associated with content items and descriptions associated with content items), the input training data may be split into a first set of features containing products identified from text from the first source and a second set of features containing products identified in text from the second source. Either the target input, the target output, or both, or neither, may be split into sets. Multiple models may be trained with different sets of data.

[0098] In block 312, the system 300A performs model training (e.g., via the training engine 182 of FIG. 1 ) using the training set 302. Training of the machine learning model may be achieved in a supervised learning manner. This involves providing a training dataset containing labeled inputs through the model, observing its outputs, defining an error (by measuring the difference between the output and the label value), and adjusting the model weights to minimize the error using techniques such as deep gradient descent and backpropagation. In many applications, repeating this process across many labeled inputs in the training dataset results in a model that can generate correct outputs when presented with inputs different from those present in the training dataset. In some embodiments, training of the machine learning model may be achieved in an unsupervised manner; for example, labels or classifications may not be supplied during training. Unsupervised models may be configured to perform anomaly detection, clustering of results, etc.

[0099] For each training data item in the training dataset, the training data item may be input to a model (e.g., a machine learning model). The model can then process the input training data items (e.g., one or more product notifications detected in association with a content item and associated confidence values, etc.) to generate output. The output may include, for example, a list of products that may be associated with the content item and corresponding confidence values. The output may be compared to the labels of the training data item (e.g., a set of human-labeled products associated with the content item).

[0100] Processing logic may then compare the generated output (e.g., predicted product / content item associations) with the labels included in the training data items (e.g., a human-generated list of product / content item associations). Processing logic determines an error (i.e., classification error) based on the difference between the output and the label(s). Processing logic adjusts one or more weights and / or values of the model based on the error.

[0101] When training a neural network, an error term, or delta, may be determined for each node of the artificial neural network. Based on this error, the artificial neural network adjusts one or more of its parameters (weights of one or more inputs of the node) for one or more of its nodes. Parameters may be updated in a back-propagation manner, such that nodes in the top layer are updated first, followed by nodes in the next layer, and so on. An artificial neural network includes multiple layers of "neurons," each layer receiving values from neurons in the previous layer as inputs. The parameters of each neuron include weights associated with values received from each of the neurons in the previous layer. Adjusting a parameter may therefore include adjusting weights assigned to each of the inputs of one or more neurons in one or more layers within the artificial neural network.

[0102] The system 300A may train multiple models using multiple sets of features in the training set 302 (e.g., a first set of features in the training set 302, a second set of features in the training set 302, etc.). For example, the system 300A may train models to generate a first trained model using a first set of features in the training set (e.g., a subset of the training data 364, such as data associated with only a subset of models configured to generate product / content item associations, etc.) and generate a second trained model using a second set of features in the training set. In some embodiments, the first trained model and the second trained model may be combined to generate a third trained model (e.g., which may be superior to the first or second trained models alone). In some embodiments, the sets of features used in comparing the models may overlap (e.g., a first set of features that are products based on the title, description, and some images of the content item, and a second set of features that are products detected based on a description of the content item, a different set of images of the content item, and the detected context of the content item (e.g., the type of product associated with the content item)). In some embodiments, hundreds of models may be generated, including models with various permutations of features and combinations of models.

[0103] At block 314, the system 300A performs model validation (e.g., via the validation engine 184 of FIG. 1 ) using the validation set 304. The system 300A may validate each of the trained models using the corresponding set of features in the validation set 304. For example, the system 300A may validate a first trained model using a first set of features in the validation set and a second trained model using a second set of features in the validation set. In some embodiments, the system 300A may validate hundreds of models (e.g., models having various combinations of features, model combinations, etc.) generated at block 312. At block 314, the system 300A may determine the accuracy of each of the one or more trained models (e.g., through model validation) and may determine whether one or more of the trained models have an accuracy that meets a threshold accuracy. In response to determining that none of the trained models have an accuracy that meets the threshold accuracy, flow returns to block 312, where system 300A performs model training using a different set of features from the training set, an updated or expanded training set provided by the dataset generator. In response to determining that one or more of the trained models have an accuracy that meets the threshold accuracy, flow continues to block 316. System 300A can discard trained models that have an accuracy below the threshold accuracy (e.g., based on a validation set).

[0104] At block 316, the system 300A performs model selection (e.g., via selection engine 185 of FIG. 1 ) to determine which of the one or more trained models that meet the accuracy threshold has the highest accuracy (e.g., selected model 308 based on validation of block 314). In response to determining that two or more of the trained models that meet the accuracy threshold have the same accuracy, flow may return to block 312, where the system 300A performs model training using a more refined training set corresponding to a more refined set of features to determine the trained model with the highest accuracy.

[0105] At block 318, the system 300A performs model testing (e.g., via the test engine 186 of FIG. 1 ) using the test set 306 to test the selected model 308. The system 300A tests the first trained model using a first set of features in the test set and determines that the first trained model meets a threshold accuracy (e.g., based on the first set of features in the test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 overfits the training set 302 and / or the validation set 304 and is not applicable to other datasets, such as the test set 306), flow continues to block 312, where the system 300A performs model training (e.g., retraining) using a different training set corresponding to a different set of features, different content items, etc. In response to a determination that the selected model 308 has an accuracy based on the test set 306 that meets the threshold accuracy, flow continues to block 320. At least in block 312, the model can learn patterns in the training data to make predictions, and in block 318, the system 300A can apply the model to the remaining data (e.g., test set 306) to test the predictions.

[0106] In block 320, the system 300A receives current data 322 (e.g., newly uploaded content items, newly created content items, content items not included in the training, test, or validation sets of the selected model 308, etc.) using a trained model (e.g., the selected model 308) and determines (e.g., extracts) output data 324 (e.g., product / content item associations and corresponding confidence values) from the output of the trained model. Corrective actions associated with the content items and / or related data may be performed in light of the output data 324. For example, instructions may be updated to include presenting a UI element with a content item that specifies that the content item includes one or more related products, instructions may be updated to include presenting a UI element with a content item that includes additional information about the related products, etc. The instructions may depend on additional factors of the content item, such as user preferences, presentation environment (e.g., a search page, a home page, etc.), etc. In some embodiments, the current data 322 may correspond to the same types of features in the historical data used to train the machine learning model. In some embodiments, the current data 322 corresponds to a subset of the types of features in the historical data used to train the selected model 308 (e.g., the machine learning model is trained using product associations and / or contextual information, as well as confidence values, from several sources, such as text-based and image-based sources, and this subset of data is provided as the current data 322).

[0107] In some embodiments, the performance of the machine learning model (e.g., selected model 308) may be adjusted, improved, and / or updated over time. For example, additional training data may be provided to the model to improve the model's ability to correctly classify product associations with content items. In some embodiments, a portion of the current data 322 may be provided to retrain the model (e.g., via training engine 182 of FIG. 1 ). A portion of the current data 322 may be labeled (e.g., human-labeled), and the labels may be provided to retrain the model (e.g., as current target output data 346). The current data 322 and the current target output data 346 may be utilized to update and / or improve the selected model 308 periodically, continuously, etc.

[0108] In some embodiments, one or more of acts 310 through 320 may occur in various orders and / or with other acts not presented and described herein. In some embodiments, one or more of acts 310 through 320 may not be performed. For example, in some embodiments, one or more of the data partitioning of block 310, the model validation of block 314, the model selection of block 316, or the model testing of block 318 may not be performed.

[0109] System 300A is described in terms of a fusion model that accepts one or more indications of a product detected in association with a content item (e.g., from text associated with the content item, from an image associated with the content item, etc.) and a confidence value (e.g., confidence that the product was actually referenced in the content item), and generates as an output an overall likelihood that the product was referenced in the content item based on the multiple inputs. Systems similar to system 300A may be utilized to perform other machine learning-based tasks, such as text analysis or image analysis models, the outputs of which may be provided as inputs to the fusion model and, with appropriate substitutions of input data, output data, etc., may be operated in a manner similar to that described in connection with system 300A.

[0110] 3B is a block diagram of an exemplary system 300B for generating associations between a content item and one or more products, according to some embodiments. System 300B may include multiple modules, such as image identification 330, image verification 340, text identification 350, and fusion 360. In some embodiments, multiple modules of system 300B may work in conjunction to identify products from content items. For example, products detected in images of a content item and content item metadata (e.g., title, description, caption, etc.) may be provided to a fusion model, which may determine, based on multiple input channels, the likelihood that one or more products are associated with the content item. Actions may be taken in response to the likelihood that a product will appear in the content item as determined by the fusion model (e.g., updating the content item metadata to include a product notification).

[0111] The image identification module 330 may be utilized to identify one or more products from images associated with a model; for example, images from a content item may be compared to images in a database of products (e.g., tens of thousands of products) to identify products present in the video. The presented products may include the specific subject matter of the content item (e.g., products reviewed in the content item), products included in the content item (e.g., products that appear incidentally, products that appear without being specifically highlighted, etc.), etc. The text identification module 350 may identify one or more products associated with a content item from metadata / text data associated with the content item, for example, from text including the title of the content item, a description of the content item, captions associated with the content item, etc. The image verification module 340 may verify products identified in a content item using one or more images. For example, the image verification module 340 may function similarly to the image identification module 330, but may also be used to confirm the presence of one or more products identified by a separate module (e.g., by comparing an image of a potential product to a more limited range of product images provided by another module). The fusion module 360 can receive candidate products and associated confidence values contained in a content item and determine the likelihood that one or more products will appear in the content item based on various inputs.

[0112] Image identification 330 may be utilized to determine products associated with a content item having a visual component, e.g., a video. Image identification 330 may include frame selection 332. Frame selection 332 may be utilized to select one or more frames of the video from which to search for images of the product. Frame selection 332 may be performed via random sampling, periodic sampling, intelligent sampling methods, etc. For example, a content item (e.g., a video) may be provided to a machine learning model, and the machine learning model may be trained to predict frames of the video that are likely to include one or more products.

[0113] One or more frames may be provided to an object detection model 334. The object detection 334 may extract predicted objects from the one or more frames. For example, the object detection 334 may separate potential products of the image data of a content item from people, animals, background, etc. The object detection 334 may be or include a machine learning model.

[0114] Images of detected objects may be provided to embedding 336. Embedding 336 may include converting one or more images to a lower dimensionality. Embedding 336 may include providing one or more images to a dimensionality reduction model. The dimensionality reduction model may be a machine learning model. The dimensionality reduction model may be configured to reduce the dimensionality of similar images in a similar manner. For example, embedding 336 may receive an image as input and generate a vector of values as output. Embedding 336 may be configured, trained, etc., such that similar images (e.g., images of the same or similar products) are represented similarly (e.g., by Euclidean distance, cosine distance, other distance metrics, etc.) in a reduced-dimensional vector space. Embedding 336 may generate reduced-dimensional data.

[0115] The reduced-dimensionality image data may be provided to product identification 338. Product identification 338 may identify one or more products associated with the reduced-dimensionality representation provided by embedding 336. Product identification 338 may compare the reduced-dimensionality image data (e.g., provided by embedding 336) with reduced-dimensionality image data of products included in product image index 339 (e.g., generated from images of the products by the same machine learning model used by embedding 336). Product image index 339 may be stored as part of a data store. Product image index 339 may include, for example, many products (e.g., hundreds of products, tens of thousands of products, or more). Product image index 339 may include associations between stored image data (e.g., reduced-dimensionality image data) and product identifiers, product indicators, etc. Product image index 339 may be segmented, e.g., the stored reduced-dimensionality data may be classified into one or more categories, classes, etc. For example, product identification 338 may compare the data received from embedding 336 to products of a particular category, type, classification, etc. In some embodiments, the category, type, classification, etc. may be provided by one or more users, one or more content creators, automatically detected (e.g., by one or more machine learning models), etc. A content item or one or more products related to a content item may be associated with a category (e.g., a general category such as electronics, a more specific category such as screen devices, a classification of products such as tablets, a manufacturer or brand, a model, or the like). Product identification 338 may generate one or more notifications of products detected in the content item's image (e.g., a list of products that may match the product represented by product image index 339) and one or more notifications of confidence values (e.g., the confidence that each of the listed products was accurately detected).The output of image identification 330 may be utilized to update the metadata of the content item (e.g., to include an association with one or more products, to include one or more product identifiers or indicators, etc.). The output of image identification 330 may be provided to image verification 340, e.g., to verify the presence within the content item of images identified by image identification 330. The output of image identification 330 may be provided to fusion 360, e.g., to generate an overall and / or multi-input determination of products included in the content item via fusion model 366. The output of image identification 330 may be provided to text identification 350 (not shown), e.g., to narrow the scope of products queried, searched, compared, etc. by text identification module 350. In some embodiments, image identification 330 is utilized to identify products detected in one or more frames of a video content item. For example, image identification module 330 may be configured to generate a list of all products detected in any selected frames and provide a confidence value for each product in each selected frame. The image identification module 330 can generate image-based product data, eg, one or more identifiers for products, which are identified based on the image of the content item.

[0116] Image verification module 340 may be configured to verify the presence of an identified product of a content item using one or more images of the content item. For example, image verification module 340 may include a model configured to confirm the presence of a product identified by another model. Image verification module 340 may include secondary identification 345. Secondary identification 345 may include similar components as image identification 330. In some embodiments, image identification module 330 may communicate directly with product candidate image index 344 instead of or in addition to secondary identification 345 communicating with product candidate image index 344. In some embodiments, secondary identification 345 may perform a similar role to image identification 330 but may include a different model, a model trained using different training data, a model configured to select different frames or detect different objects, etc.

[0117] Image verification 340 may include a synthetic model 341. The synthetic model 341 may receive notifications of products identified by image identification module 330, text identification module 350, secondary identification 345 (data flow not shown), etc. The synthetic model 341 may select an object detection model 342 to provide data to (e.g., an object detection model specifically configured for a product category or classification). The synthetic model 341 may include image selection; for example, the synthetic model 341 may provide one or more images to object detection 342, select one or more frames to provide to object detection 334, etc. For example, the synthetic model 341 may provide one or more frames that may contain a product to object detection 342 based on data received from image identification 330 and text identification 350. The object detection model 342 may perform similar functions to the object detection model 334, e.g., modified by the functions of the synthetic model 341. The embedding 343 may perform similar functions to embedding 336, e.g., to reduce the dimensionality of the detected product image. In some embodiments, the product candidate image index 344 may include reduced-dimensionality image data (e.g., vectors of values) detected by other modules (e.g., image identification 330, text identification 350, etc.). Secondary identification 345 may compare the reduced-dimensionality image data (e.g., embedded image data) with candidate data in the product candidate image index 344 (e.g., products identified by modules other than image verification module 340) to verify the presence of products in the content item.

[0118] Text identification 350 may be configured to identify one or more products from text data (e.g., metadata) associated with a content item. Text identification 350 may generate metadata-based product data, e.g., one or more product identifiers, based on the metadata of the content item. Text identification 350 may generate text-based product data, e.g., one or more product identifiers, based on text data associated with the content item. Text identification 350 may identify products from one or more of the title of the content item, the description of the content item, a caption associated with the content item (e.g., a machine-generated caption), a comment associated with the content item, and / or other text data or metadata of the content item. The text data (e.g., metadata) associated with the content item may be provided to text analysis model 352. Text analysis model 352 may be a machine learning model. Text analysis model 352 may be configured to detect or predict products from text data associated with the content item. Text analysis model 352 may be configured to detect one or more products having product identifiers stored in product identifier 354. The text analysis model 352 may provide output (e.g., a list of detected candidate products, associated confidence values, etc.) to the image verification module 340. The text analysis model 352 may provide the output of the composition model 341. The text analysis model 352 may provide output that influences the product of image verification, for example, the output of the text analysis model 352 may cause the products detected by the text identification module 350 to be added to the product candidate image index 344. The image verification module 340 may query an index (e.g., the product candidate image index 344) that includes products detected by other modules, e.g., the image identification module 330, the text identification module 350, etc.

[0119] The fusion module 360 may receive output data (e.g., detected products, associated confidence values) from one or more sources (e.g., image identification module 330, image verification module 340, text identification module 350, etc.). The fusion module 360 may further receive data from other sources, such as context term extraction 362 or additional feature extraction 363. The context term extraction 362 may, for example, provide context to potential products of a content item, such as detecting categories or themes related to some products. The context term extraction 362 may be performed by one or more machine learning models. The context term extraction 362 may detect contextual information from text associated with the content item, metadata associated with the content item, etc. The additional feature extraction 363 may provide additional details that can be used to determine whether one or more products appear in the content item. The additional features may include video embeddings. Additional features may include other metadata of the content item, such as the date the content item was uploaded to the content providing platform (e.g., compared to the release date of a product), the classification of the content item (e.g., shopping or product review videos are more likely to include products than other types of videos), etc.

[0120] Data from multiple sources may be provided to the fusion model 366. The fusion model 366 may be configured, for example, to receive data including one or more products and confidence values and determine the one or more products along with confidence values that indicate the likelihood that the products will appear in a content item. In some embodiments, the content item may be a video. In some embodiments, the content item may be a live-streamed feed, for example, a live-streamed video feed (e.g., product review feeds, unboxing feeds, etc.). The content item may be a short-form video.

[0121] 4A through 4E illustrate exemplary UIs presented on a user device, including UI elements that indicate related products, according to some embodiments. FIGS. 4A through 4E may include UIs provided as part of applications on devices 400A through 400E, such as web browser applications, mobile applications associated with / provided by a content platform, etc. User interaction with various elements in FIGS. 4A through 4E may change the presented UI elements. For example, interacting with a UI element that indicates a content item has related products may cause a second UI element to be displayed (e.g., replacing a first UI element, expanding the first UI element, etc.) to present additional information (e.g., about the related products). UI elements may include elements that, when interacted with, cause the UI to display less information about the related products (e.g., collapsing a panel describing one or more related products). Interacting with a UI element related to one or more products may result in a different effect, such as a transition to a UI environment presenting the content item. The new UI environment may include one or more UI elements related to one or more products of the content item. Interacting with a UI element associated with a product may display a UI element that facilitates a transaction (e.g., purchase) of the product. Various interactions between the UIs and UI elements presented in Figures 4A through 4E are possible (e.g., interacting with an element of a first UI layout may transition to a second UI layout), and any transitions between sample UIs, similar UIs, inclusion of similar UI elements, etc. are within the scope of this disclosure. Although Figures 4A through 4E are described in connection with video content items, other types of content items (e.g., image content, text content, audio content, etc.) may be presented in a similar UI. Any optional features, elements, etc. presented with respect to one or more of Figures 4A through 4E may be included in similar systems as other figures of these figures, if desired.

[0122] 4A illustrates a device 400A presenting an example UI 402 including a UI element 404 showing one or more related products, according to some embodiments. A UI including elements of UI 402 and / or similar UI elements may be presented by device 400A as part of presenting one or more content items for user selection. In FIG. 4A, UI element 404 is shown in a collapsed state, e.g., a default collapsed state.

[0123] UI 402 includes a first content item selector 406 (e.g., a video thumbnail) and a second content item selector 408. In some embodiments, more or fewer content items may be selectable, UI 402 may be scrolled to display additional content items, etc. UI element 404 is associated with the content item indicated by content item selector 406. UI element 404 (and other UI elements in FIGS. 4A-4D related to products) may be presented above, next to, on top of, within, below, etc. the content item associated with UI element 404 or a content item selector for the content item. In some embodiments, product information (e.g., product / content item associations) may be provided by a content creator. Product information may be provided by one or more users. Product information may be provided by an administrator. Product information may be obtained from the content item, for example, via one or more machine learning models, via a system such as system 300B, etc.

[0124] In some embodiments, a user may be presented with a replaced UI element, an updated UI element, etc. by interacting with UI element 404. For example, a user may interact with expand element 410 to display more information about a product associated with the content item. In some embodiments, expanding UI element 404 may open a panel containing additional information about one or more products associated with the content item. By expanding UI element 404, interacting with UI element 404, etc., the presentation of UI 402 may be adjusted to include, for example, the elements shown in FIGS. 4B-4D.

[0125] UI elements 404 may include, for example, an expansion element 410, related product notifications (e.g., how many products are associated with a content item), visual notifications of products 412 (e.g., visual notifications that a transaction or purchase is available, that a link to a product retailer is available, etc.), etc. User interaction with one or more components of UI element 404 may cause device 400A to change the presentation of UI element 404, e.g., user selection of expansion element 410 may cause the presentation of UI element 404 to change to an expanded state.

[0126] The UI including the UI element 404 may be presented in response to device 400A sending a request for a content item to a content providing platform (e.g., content providing platform 120 of FIG. 1 ). The UI element 404 may be presented within a home feed (e.g., a list of content items suggested to a user or user account). The UI element 404 may be presented within a suggested feed (e.g., a list of suggested content items based on one or more recently presented content items). The UI element 404 may be presented as part of a playlist (e.g., a list of presented content items populated by a user, creator, etc.). The UI element 404 may be presented within a search feed (e.g., in response to a user-generated search query sent by device 400A to the content providing platform). For example, UI elements related to one or more products may be presented based on the inclusion of a product, product category, product brand, etc. in the search query. The UI element 404 may be presented within a product-focused feed (e.g., a shopping content feed). The UI element 404 may be responsive to a user selection of a content item, e.g., may be displayed when a user selects to watch a related video, may present related content items, etc. The UI element 404 may be presented based on detection of a user's expressed interest in a content item. For example, a user may dwell on a thumbnail of a video (e.g., when the user places a cursor over the thumbnail, when the user pauses scrolling while the thumbnail is presented, etc.). The UI element 404 may be presented when the dwell meets one or more conditions (e.g., a duration condition, a thumbnail position condition, etc.). The UI element 404 can be presented in response to additional data, e.g., user account history, user settings, user preferences, etc. The UI element 404 may be presented with a list of presented content items or while presenting a content item (e.g., while playing a video related to the product of the UI element 404).

[0127] FIG. 4B shows a device 400B presenting an example UI 420 including a UI element 422 showing related products, according to some embodiments. The UI element 422 includes information about one or more products related to the content item. In FIG. 4B, the UI element 422 is presented in an expanded state. The UI element 422 may include several components. For example, a first component may include information about a first product (e.g., a photo, product name, price, timestamp, etc.), a second component may include information about a second product, etc. In some embodiments, the UI element 422 may be scrollable, for example, to access information about additional products. The UI element 422 may include multiple tabs (e.g., the UI element 422 may be associated with product information and one or more other types of information). For example, the UI element 422 may include a products tab 424 and a chapters tab 426. In some embodiments, the products tab 424 may be open by default (e.g., the content of the products tab 424 may be presented by default). In some embodiments, the chapters tab 426 may be open by default (e.g., the content of the chapters tab 426 may be presented by default). In some embodiments, other tabs may be open by default. For example, the chapters tab 426 may be open by default, except for content items with related products, and the tab that is open by default may be selected based on user history (e.g., user history, such as interactions with elements such as chapters tabs, product tabs, etc.), based on a search query (e.g., a search including a product name, related terms or phrases such as "product reviews"), etc. UI elements 422 may include additional elements, e.g., elements that a user may utilize to control the presentation of UI 420, UI elements 422, etc.For example, UI element 422 may include a "close" element that presents UI 420 without UI element 422, a "collapse" element that displays less information (e.g., to collapse a panel to mimic UI element 404 of FIG. 4A , to change the presentation of UI element 422 to a collapsed state, etc.), one or more elements for displaying more information (e.g., one or more listed products, listed product icons, etc. may be selected to present additional information about the products, facilitate the purchase of the products, etc.), etc.

[0128] The products tab 424 of UI element 422 may include one or more photos of the product, information about the product (e.g., product name, product description, etc.), one or more product prices, and timestamps of content items associated with the product. In some embodiments, the product information may be provided by a content creator, one or more users, a system administrator, etc. In some embodiments, the product information may be obtained by one or more models, e.g., machine learning models. For example, the presence of a product in a content item, the association of a product in a content item, the timestamp or location at which the product appears in the content item, etc. may be determined by one or more machine learning models. A system such as system 300B of FIG. 3B may be utilized to determine one or more products associated with a content item (e.g., a video). The portion of the content item (e.g., a timestamp) associated with the content item may be determined via, for example, the timing of a caption associated with the product, the timing of the display of an image or video frame including the product, etc. In some embodiments, selecting a product may present the portion of the content item associated with the product, present the content item from the time indicated by the timestamp associated with the product, etc.

[0129] In some embodiments, portions (e.g., visual components) of UI element 422 related to a particular product may be displayed by default, may be displayed in a different manner (e.g., highlighted), etc. For example, when a search query including the name of a product is received from a user, UI element 422 may be displayed that includes a display related to the searched product.

[0130] In some embodiments, one or more associations between content items and products may be stored, for example, as metadata associated with the content item (metadata associated with a content item may further include the content item's title, description, viewing history, captions associated with the content item, etc.). In response to a device (e.g., device 400B) executing instructions to display a list of content items for presentation, present content items to a user, present a UI element (e.g., UI element 422) including information about one or more products, etc., the device may retrieve information about the products based on the metadata associating the products with the content items. Information about the products (e.g., images, related products, e.g., color variations, availability, price, etc.) may be retrieved from a data store. The data store may contain information about the products and may be updated as information changes, e.g., as product price, and the UI may retrieve and display the updated information based on the content item / product associations.

[0131] UI element 422 may be presented as part of a home feed (e.g., a list of content items suggested to a user or user account), a suggested feed (e.g., a list of suggested content items based on one or more recently presented content items), a playlist, a list of search results, a shopping content page, etc. UI element 422 may be presented upon selection of a content item for presentation, upon presentation of a content item, etc. UI element 422 may be presented upon dwell (e.g., pausing scrolling on a content thumbnail, pausing scrolling on a less detailed element such as UI element 404 in FIG. 4A , etc.). In some embodiments, UI element 422 may be removed or replaced upon a user action, such as scrolling, and UI element 422 may collapse to a shape similar to UI element 404 (e.g., to facilitate selection of a content item from a list of content items, to simplify scrolling through a list of content items, etc.). UI element 422 may be displayed while presenting a list of content items, while presenting a single content item (e.g., while playing a video related to the product of UI element 422), etc.

[0132] FIG. 4C illustrates a device 400C presenting an exemplary UI 430 including a UI element 432 presenting a content item and a UI element 434 presenting information about related products, according to some embodiments. The UI element 434 may include more detailed information about the product related to the content item than the UI element 422 of FIG. 4B. The UI element 434 may be product-focused, e.g., function to display product information to a user. The UI element 434 may include one or more components, e.g., a component related to a first product, a component related to a second product, etc. In some embodiments, the UI element 434 may include a list of products related to the content item. The UI element 434 may be navigable, scrollable, etc. The UI element 434 may include one or more control elements (e.g., non-product-related), e.g., a back button for returning to a previous view, a close button for closing the UI element 434 to see a different set of UI elements, etc. In some embodiments, the UI element 434 may be removed from the UI 430 in response to other user actions, e.g., the user scrolling past the related content item. In some embodiments, a user selection of a product presented via UI 434 may present a UI element that facilitates the purchase of the product. In some embodiments, UI element 434 may be displayed in response to a determination that the user is interested in one or more products related to the content item (e.g., based on user history, one or more terms in a user search query, the user selecting and / or being presented with a shopping content item, a user selection of a product or a UI element related to the product, etc.).

[0133] The UI element 434 may include one or more photos and / or additional information about one or more products associated with a content item (e.g., the content item presented via the UI element 432). The photos and / or information may be provided by a content creator, one or more users, retrieved from a database (e.g., based on product / content item association metadata), etc. In some embodiments, the UI element 432 may scroll automatically. For example, the UI element 432 may scroll as content items are presented, e.g., to bring into view products associated with a portion of the currently presented content item. The UI element 432 may be presented in response to a user selection of a presented content item. The UI element 432 may be presented in response to other factors, such as user history. The UI element 432 may be presented when a user hovers over a related content item, a related UI element, etc.

[0134] In some embodiments, UI element 434 may present information about a single product, such as a product selected by a user (e.g., via UI element 422 of FIG. 4B ). UI element 434 may be product-focused, may be single-product-focused, may display product variations (e.g., as described in connection with UI element 444 of FIG. 4D ), or may include other elements, components, and / or information described in connection with other UI elements described herein.

[0135] 4D is a diagram illustrating a device 400D presenting an example UI 440 including a UI element 442 presenting a content item and a UI element 444 facilitating a transaction related to the product, according to some embodiments. The UI element 444 can provide one or more fields relevant to a user conducting a transaction, such as, for example, purchasing a product. For example, the UI element 444 can include an alternative product panel 446 that can include information, photos, prices, etc. about alternative products (e.g., products related to one or more products associated with the content item, such as color variations, size variations, variations in a product bundle, related products such as a similar product from another brand, etc.).

[0136] UI elements 444 may include transaction elements 448. In some embodiments, transaction elements 448 may facilitate transactions (e.g., purchases) within the application providing UI 440. In some embodiments, transaction elements 448 may facilitate transactions via other applications, other websites, etc. For example, interacting with transaction element 448 may direct a user to a merchant website or cause device 400D to open an application related to the purchase of an item.

[0137] UI element 444 may be navigable, scrollable, etc. UI element 444 may include one or more control elements, such as a back button for returning to a previous view, a close button for closing UI element 444 and displaying a different set of UI elements via UI 440, etc. In some embodiments, UI element 444 may be displayed as part of a UI 440 presenting content items, etc., as part of a list of content items for user selection. UI element 444 may be displayed in response to a user determining an interest in a transaction (e.g., purchase) related to one or more products included in a content item, e.g., selection of a product from a UI element such as UI element 434 of FIG. 4C , selection of a content item, dwelling on a content item, inclusion of a product or product-related terms in a search query, user navigation to a shopping-specific listing of content items, etc. UI element 444 may receive information about one or more products from a database, for example, based on metadata associating a content item with one or more products.

[0138] 4A-4D may be organized in various configurations. For example, a UI element such as UI element 404 may be presented, user interaction with UI element 404 may cause an element such as UI element 422 to be presented, user interaction with UI element 422 may cause UI element 434 to be presented, user interaction with UI element 434 may cause UI element 444 to be presented, etc. One or more UI elements may include navigation elements for instructing the device to display different UI elements; for example, interacting with expansion element 410 may cause a UI element such as UI element 422 to be presented, interacting with a different one of UI elements 404 may cause UI elements such as UI element 434, UI element 444, etc.

[0139] Other connections between UI elements are possible. For example, interacting with a UI element such as UI element 404 may cause a UI element such as UI element 422, UI element 434, UI element 444, etc. Interaction with a UI element, or portion thereof, such as UI element 422, may cause a UI element, such as UI element 404, UI element 434, UI element 444, etc. User interaction with a UI element, or portion thereof, such as UI element 434, may cause a UI element, such as UI element 404, UI element 422, UI element 444, etc. User interaction with a UI element, or portion thereof, such as UI element 444, may cause a UI element, such as UI element 404, UI element 434, UI element 444, etc. The default UI element that is presented may depend on the environment in which the UI element is presented (e.g., a list of content items presented in response to a search, a home feed, a shopping feed, a watch feed, etc., the environment containing the content item being presented, or the like). For example, the selection of the format of the UI element may be based on several factors. In some embodiments, when a search query includes a product name, category, etc., default UI elements may change; for example, the UI may default to displaying UI elements that include information about the product, UI elements that include options to purchase the product, etc. The determination of the format of UI elements to display may be based on user history, user account history, user actions (e.g., opening the home feed, presenting a watch feed, submitting a search query, selecting a shopping feed, etc.). The transition between formats of product-related UI elements may be determined by additional data similar to the data used to determine the format of the UI elements to be presented.

[0140] 4E is a diagram illustrating an example device 400E with UI elements overlaid on a content presentation element 452, according to some embodiments. Device 400E includes a UI 450. UI 450 may be provided by an application, for example, an application associated with a content provision platform. Presentation element 452 may present a content item (e.g., a video). UI 450 may present a list of additional content items, additional information related to the presented content item (e.g., title, description, comments, live chat, etc.), additional UI elements related to the product (e.g., UI elements such as UI elements 404, 422, 434, 444, or variations).

[0141] One or more UI elements may be overlaid on the presentation element 452. The UI element 454 may show a product included in a content item (e.g., a content item displayed in a video). The UI element 454 may perform functions similar to the other UI elements in FIGS. 4A through 4D , such as presenting information about the product, enabling the display of more information about the product, or facilitating the purchase of the product. The placement of the overlay element may be determined by one or more users, a content creator, a model (e.g., utilizing a machine learning model configured to detect products or objects in an image to avoid content item display areas that include objects), etc. The UI element 454 may include a visual indicator of where in the content item (e.g., where in the video) one or more related products are located. The UI element 454 may be presented during the presentation of the content item, e.g., while the video is playing. The UI element 454 may be displayed and / or removed in response to the presence of related products in the content item, e.g., displayed while the content item is in the video. UI element 454 may show multiple products, identify the products (e.g., display the name of one or more products), display information about the products, etc. Multiple UI elements such as UI element 454 may be displayed, for example, on a thumbnail of the video, throughout the presentation of the video, simultaneously during the video, etc.

[0142] Overlaid UI elements such as UI element 454 may be presented in combination with other UI elements related to the product; for example, UI element 456 may open a panel containing information about multiple products related to the content item, UI element 454 may display a UI element containing information about the product depicted in a photograph, etc.

[0143] The overlaid UI element 454 may be displayed over (e.g., in front of, visually preceding, etc.) the visual representation of the content item. For example, the UI element 454 may be overlaid over a video thumbnail. The UI element 454 may be displayed over a content item. For example, the UI element 454 can be overlaid over a playing video. The presentation of the UI element 454 may be performed in response to a user action. For example, once it is determined that a user is interested in one or more products (e.g., via a search query, via interaction with product-related UI elements, via user history, etc.), the UI element 454 may be displayed overlaid over other UI elements. In some embodiments, the content item may be a live streaming video. In some embodiments, the content item may be a short-form video.

[0144] In some embodiments, UI element 454 can perform similarly to the description of the performance of UI element 404, e.g., can notify the user that one or more products are associated with the content item. UI element 454 can respond to interactions from the user similarly to UI element 404, e.g., can open or expand a panel containing product information, change the presentation of UI element 404 to show more or different information, expand the UI element to include more information, begin presenting the content item, etc. UI element 454 can respond to interactions from the user similarly to UI element 434, e.g., can open or expand a panel that facilitates a transaction.

[0145] 5A through 5F are flow diagrams of methods 500A through 500F relating to content items having related products, according to some embodiments. Methods 500A through 500F may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, a processing device, etc.), software (e.g., a processing device, instructions executing on a general-purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. In some embodiments, methods 500A through 500F may be performed in part by content platform system 102, product identification system 175, and / or client device 110 of FIG. 1. Method 500A may be performed in part by product identification system 175 (e.g., server machine 170 and dataset generator 172 of FIG. 1, dataset generator 272 of FIG. 2). Product identification system 175 may use method 500A to generate a dataset for at least one of training, validating, or testing a machine learning model, according to embodiments of the present disclosure. Methods 500B through 500D may be performed by product identification system 175 (e.g., system 300B of FIG. 3B ) and / or server machine 180 (e.g., training, validation, and testing operations may be performed by server machine 180). Method 500E may be performed by client device 110. Method 500E may be utilized by client device 110 to display one or more UI elements related to a product, for example, to facilitate user identification of a product included in a content item. Method 500F may be performed by content platform system 102, for example, by processing logic of content providing platform 120 to facilitate presentation of one or more UI elements related to a product by client device 110. In some embodiments, a non-transitory machine-readable storage medium stores instructions that, when executed by a processing device (e.g., a processing device such as product identification system 175, server machine 180, etc.), cause the processing device to perform one or more of methods 500A through 500F.

[0146] For ease of explanation, methods 500A-500F are shown and described herein as a series of acts. However, acts in accordance with the present disclosure may occur in various orders and / or concurrently, and with other acts not shown or described herein. Moreover, not all of the acts shown may be performed to implement methods 500A-500F in accordance with the subject matter of the present disclosure. Furthermore, those skilled in the art will understand and appreciate that methods 500A-500F may alternatively be represented as a series of interrelated states via a state diagram or events.

[0147] 5A is a flow diagram of a method 500A for generating a dataset for a machine learning model, according to some embodiments. Referring to FIG. 5A, in some embodiments, at block 401, processing logic implementing method 500A initializes a training set T to an empty set.

[0148] At block 502, processing logic generates a first data input (e.g., a first training input, a first validation input) that may include one or more of product data, image data, metadata, text data, confidence data, etc. In some embodiments, the first data input may include a first set of features for a type of data, and the second data input may include a second set of features for a type of data (e.g., as described with respect to FIG. 3A ). In some embodiments, the input data may include historical data.

[0149] In some embodiments, at block 503, the processing logic optionally generates a first target output for one or more of the data inputs (e.g., the first data input). In some embodiments, the input includes one or more predicted products detected in the content item and associated confidence intervals, and the target output may include labels of the products included in the content item. In some embodiments, the input includes one or more sets of data related to the content item (e.g., image data such as frames or portions of frames of video, metadata such as title text or caption text, etc.), and the target output is a list of products included in the content item. In some embodiments, the first target output is predicted data. In some embodiments, the input data may be in the form of caption text data, and the target output may be a list of possible corrections to the caption to include product names / references for a machine learning model configured to correct captions by including product information. In some embodiments, no target output is generated (e.g., an unsupervised machine learning model capable of grouping input data or finding correlations in input data and not requiring the provision of a target output).

[0150] At block 504, processing logic optionally generates mapping data indicating an input / output mapping. The input / output mapping (or mapping data) may refer to data inputs (e.g., one or more of the data inputs described herein), target outputs for the data inputs, and associations between the data input(s) and the target outputs. In some embodiments, block 504 may not be performed, such as in connection with a machine learning model that does not provide a target output.

[0151] At block 505, processing logic, in some embodiments, adds the mapping data generated at block 504 to dataset T.

[0152] At block 506, processing logic branches based on whether dataset T is sufficient for at least one of training, validating, and / or testing a machine learning model, such as one of models 190 in FIG. 1. If so, execution proceeds to block 507; if not, execution continues back to block 502. Note that in some embodiments, whether dataset T is sufficient may be determined based solely on the number of inputs that map to outputs in the dataset, while in some other embodiments, whether dataset T is sufficient may be determined based on one or more other criteria in addition to or instead of the number of inputs (e.g., a measure of diversity of the data examples, accuracy, etc.).

[0153] At block 507, processing logic provides dataset T (e.g., to server machine 180 of FIG. 1 ) for training, validating, and / or testing machine learning model 190. In some embodiments, dataset T is a training set and is provided to training engine 182 of server machine 180 to perform training. In some embodiments, dataset T is a validation set and is provided to validation engine 184 of server machine 180 to perform validation. In some embodiments, dataset T is a test set and is provided to test engine 186 of server machine 180 to perform testing. In the case of a neural network, for example, input values (e.g., numerical values associated with data inputs 210) of a given input / output mapping are input to the neural network, and output values (e.g., numerical values associated with target outputs 220) of the input / output mapping are stored in output nodes of the neural network. Connection weights in the neural network are then adjusted according to a learning algorithm (e.g., backpropagation, etc.), and the procedure is repeated for other input / output mappings in dataset T. After block 507, the model (e.g., model 190) may be at least one of trained using training engine 182 of server machine 180, validated using validation engine 184 of server machine 180, or tested using test engine 186 of server machine 180. The trained model may be implemented by product identification system 175 to generate output data, for example, used by product information platform 161 to provide product data to users, provided to a fusion model, utilized to update metadata of content items to include one or more product associations, etc.

[0154] FIG. 5B is a flow diagram of a method 500B for updating metadata of a content item, according to some embodiments. At block 510, processing logic (e.g., a processing device, a computer processor, etc.) receives first data. The first data includes a first identifier (e.g., an indicator, a pointer to further data, a code identifying the product, etc.) of a first product determined to be associated with the content item based on the content item's metadata. The content item may be or include visual content, audio content, text content, video content, etc. The first product may have been determined by providing the content item's metadata to one or more trained machine learning models (e.g., text identification module 350 of FIG. 3B ). The metadata may include text data, such as the content item's title, description, caption, comments, live chat, etc. The first data may further include a first confidence value associated with the first product and the content item. The first confidence value may indicate a likelihood that the first product is associated with the first content item, e.g., a likelihood that the first product will appear or be referenced in the content item's metadata. The first data may further include an identifier of the second product and a second trust value associated with the second product. The first data may include a list of products and associated trust values.

[0155] At block 512, processing logic receives second data including a second identifier for the first product. The second identifier has been determined to be associated with the content item based on image data of the content item (e.g., one or more frames of a video, one or more portions of an image, etc.). The second data also includes a second confidence value associated with the first product and the content item. The confidence value and identifier may be generated by one or more machine learning models (e.g., image identification module 330 of FIG. 3B). The machine learning model may include a system configured to reduce the dimensionality of the image data. One or more candidate product images may be dimensionally reduced (e.g., images converted into vectors of values via a trained machine learning model). The machine learning model may perform operations including comparing the dimensionally reduced image data from the content item with dimensionally reduced product images in the data store to determine, for example, the likelihood that the image contains the product. The second data may include a list of products (e.g., product identifiers) and a list of confidence values that include at least the first product. The confidence value may indicate the likelihood that the related product will appear in the content item, be referenced by the content item, etc.

[0156] In some embodiments, one or more images of a content item are analyzed for potential products contained in the images. The presence of the potential product may be verified, for example, by providing the image of the content item for further product image detection analysis (e.g., a different frame of a video content item, additional frames, etc.), by text or metadata verification, etc. For example, after a candidate product is found, further analysis aimed at verifying the candidate product may be performed, for example, by searching for other evidence of the identified product. In some embodiments, text data and / or metadata associated with the content item may be analyzed for potential / candidate products. The presence of the potential product may be verified, for example, by image-based verification, text verification, etc.

[0157] In some embodiments, the processing logic may be further provided with one or more timestamps, e.g., a timestamp of a frame of a video that includes the detected candidate product, a timestamp of a caption associated with the video or audio content in which the detected candidate product appears, etc. The processing logic may utilize the timestamps for further analysis to adjust metadata for the content item, generate UI elements associated with the presentation of the content item, etc.

[0158] At block 514, processing logic provides the first data and the second data to a trained machine learning model. The trained machine learning model may be a fusion model. The trained machine learning model may be provided with one or more lists of products with associated confidence values.

[0159] At block 516, processing logic receives a third confidence value associated with the first product from the trained machine learning model. In some embodiments, processing logic may receive a list of confidence values associated with a list of products that includes the first product.

[0160] At block 518, processing logic adjusts metadata associated with the content item taking into account the third confidence value. In some embodiments, adjusting the metadata may include adding one or more connections between the content item and the product to the metadata. For example, adjusting the metadata may include adding a notice that a particular product is associated with, featured in, included in, promoted by, or the like. Adjusting the metadata may include, for example, adjusting captions for the content item to include one or more references to products that were incorrectly transcribed during caption generation.

[0161] 5C is a flow diagram of a method for training a machine learning model 500C related to product pairings of content items, according to some embodiments. In some embodiments, the machine learning model trained using method 500C may be a fusion model. Similar methods can be utilized to train different models connected to product pairings of media items, such as an image identification model, an image verification model, a text identification model, a caption update model, etc.

[0162] At block 520, processing logic receives product image data related to multiple content items. The product image data may include data associating products with content items, where the association may be derived from one or more images, e.g., frames of a video. The product image data includes an indication of one or more products (e.g., potential products, candidate products) detected (e.g., determined) in the image and a confidence value for one or more product images.

[0163] At block 522, processing logic receives product text data associated with a plurality of content items. The product text data may include data associating products with the content items. The association may be derived from text associated with the content items, e.g., metadata associated with the content items. The product text data includes one or more product notices detected in the text (associated with the content items) and one or more product text confidence values.

[0164] The data received (or, in some embodiments, obtained) by the processing logic at blocks 520 and 522 may be used as training inputs for training the fusion model. Training the machine learning models to perform different functions may include processing logic receiving different data as training inputs.

[0165] At block 524, processing logic receives data indicating products included in the plurality of content items. For example, each of the plurality of content items used to train the model (e.g., data associated with the content items may be used to train the model) may include a list of associated products, e.g., labeled by one or more users, labeled by a content creator, etc. The data received by the processing logic at block 524 may be used as target outputs for training the fusion model. Training the machine learning model to perform different functions may include processing logic receiving different data as target outputs.

[0166] At block 526, processing logic provides the product image data and product text data to the machine learning model as training inputs. Processing logic may provide different types of data to train different machine learning models. In some embodiments, a machine learning model for frame selection may be trained by providing frames of video as training inputs to the model. In some embodiments, a machine learning model for object detection may be trained by providing images (possibly including products) as training inputs to the machine learning model. A machine learning model for embedding may be trained by providing one or more images of an object (e.g., a product) as training inputs to the model. In some embodiments, a text analysis model may be trained by providing text (e.g., metadata) associated with a content item as training inputs. In some embodiments, a model configured to correct captions may be provided with machine-generated captions as training inputs.

[0167] At block 528, processing logic provides data indicating products included in the plurality of content items (e.g., a list of products included in each content item of the plurality of content items) to the machine learning model as a target output. The processing logic may provide different types of data to train different machine learning models. In some embodiments, a machine learning model for frame selection may be trained by providing data indicating which frames of one or more videos include products as a target output. In some embodiments, a machine learning model for object detection may be trained by providing labels of objects in images that are provided to the model as a target output. In some embodiments, a text analysis model may be trained by providing content items referenced by text of the content item as a target output. In some embodiments, a model configured to correct captions may be provided with a corrected caption (e.g., including one or more products) as a target output. In some embodiments, no target output is provided for training a machine learning model (e.g., an unsupervised machine learning model).

[0168] 5D is a flow diagram of a method 500D for adjusting metadata associated with a content item, according to some embodiments. At block 530, processing logic obtains first metadata associated with the content item. The metadata may include text data. The metadata may include a title, description, caption, comments, live chat, etc. of the content item. At block 531, processing logic provides the first metadata to a first model. In some embodiments, the model is a trained machine learning model. In some embodiments, the model is a product detection model, e.g., the model is configured to receive the metadata and generate notifications for products associated with the content item (e.g., taking the metadata into account).

[0169] At block 532, processing logic obtains a first product identifier based on the first metadata and a first confidence value associated with the first product identifier as output of the first model. The product identifier may be an ID number, an indicator, a product name, or any data that (uniquely) distinguishes a product. The first product identifier may identify the first product. In some embodiments, processing logic may obtain a list of products (e.g., candidate products, potential products) and associated confidence values.

[0170] At block 533, processing logic obtains image data for the content item. In some embodiments, the image data may include or be extracted from one or more frames of video. In some embodiments, the image data may be obtained from an object detection model. In some embodiments, the image data may include one or more products associated with the content item.

[0171] At block 534, processing logic provides the image data to a second model. In some embodiments, the second model is a machine learning model. In some embodiments, the second model is a model configured to identify a product from an image. In some embodiments, the second model is a model configured to verify the presence of a product identified from the image. In some embodiments, the second model may reduce the dimensionality of the provided image data. In some embodiments, the second model may compare the reduced-dimensionality image data to second reduced-dimensionality image data (e.g., output by a machine learning model obtained from a data store).

[0172] At block 535, processing logic obtains a second product identifier based on the image data and a second confidence value associated with the second product identifier as output of the second model. In some embodiments, the second product identifier indicates a second product. In some embodiments, the second product is the same as the first product. In some embodiments, processing logic may obtain a list of products and associated confidence values.

[0173] At block 536, processing logic provides data including the first product identifier, the first confidence value, the second product identifier, and the second confidence value as input to a third model, which may be a fusion model.

[0174] At block 537, processing logic obtains a third product identifier and a third confidence value as output of the third model. In some embodiments, the third confidence value may indicate the likelihood that the product indicated by the third product identifier is associated with (e.g., present in) the content item. In some embodiments, the third model may output a list of products and associated confidence values. In some embodiments, the third product identifier identifies a third product. In some embodiments, the third product is the same as the second product. In some embodiments, the third product is the same as the first product. In some embodiments, the first product, the second product, and the third product are all the same product.

[0175] At block 538, processing logic adjusts second metadata associated with the content item taking into account the third product identifier and the third confidence value. Adjusting the metadata may include supplementing the metadata with one or more product associations, e.g., notifications of associated products. Adjusting the metadata may include, for example, updating captions to include products that may have been mistranscribed (e.g., mistranscribed by a machine-generated caption model). In some embodiments, processing logic may further receive one or more timestamps (e.g., times in the video at which the products are detected within images of the video) associated with the content item and the one or more products. Updating the metadata may include adding to the metadata a notification of the time at which the products are found within the content item.

[0176] FIG. 5E is a flow diagram of a method 500E for presenting UI elements related to one or more products, according to some embodiments. At block 540, processing logic (e.g., of a user device, a client device, etc.) presents a UI. The UI includes one or more graphical representations of one or more content items (e.g., videos). A graphical representation (e.g., a video thumbnail) of a content item can be selected to initiate presentation of related content items. The one or more graphical representations of the content items may be displayed along with UI elements related to one or more products. Each graphical representation of each content item may be displayed along with UI elements associated with one or more products. The UI element(s) may be presented / displayed in a collapsed state, e.g., a collapsed default state. The UI elements include information identifying multiple products covered by each content item. The UI elements may identify that one or more products are associated with the content item. The UI elements may identify one or more products (e.g., via name, photo, etc.) associated with the content item (e.g., content covered by a video).

[0177] The UI may present selectable graphical representations of content items. The represented content items may be part of a home feed provided in response to a search, part of a watchlist, part of a playlist, or part of a shopping feed. In some embodiments, a UI element may be overlaid on top of and / or in front of one or more other elements of the UI. For example, a UI element (e.g., in a collapsed state) may be overlaid on a graphical representation of a content item, may be overlaid on the content item (e.g., while the content item is being presented), etc.

[0178] At block 542, in response to a user interaction with the collapsed UI element, processing logic continues to facilitate the presentation of the graphical representation of the respective video while changing the presentation of the UI element from the collapsed state to the expanded state. Interaction with the UI element may include selecting the UI element. Interaction with the UI element may include dwelling on the UI element (e.g., hovering a cursor over the UI element, scrolling to the UI element, pausing scrolling, etc.). The expanded UI element may include multiple visual components. Each of the visual components may be associated with one of multiple products. The visual components may include photos, descriptions, prices, timestamps, etc. associated with various products.

[0179] In some embodiments, a UI element (e.g., in an expanded state) may include multiple tabs. For example, the UI element may include a tab for products, tabs for chapters or portions of a content item, etc. The UI element may display / open the product tab by default for content items that have associated products. The UI element may display the product tab by default in response to user actions and / or history.

[0180] At block 544, in response to a user selection of one of the plurality of visual components of the expanded UI element, processing logic initiates presentation of a respective content item covering a product related to the selected visual component. Processing logic may initiate playback of a video covering the product associated with the selected visual component. Processing logic may initiate presentation of a portion of the content item related to the product of the selected visual component (e.g., initiate playback of a portion of the video).

[0181] In some embodiments, interacting with a UI element can change the UI element to a product focus state. The product focus state may present additional information, detailed information, etc. about one or more products. Interacting with a UI element in a collapsed state can change the UI element to a product focus state. Interacting with a UI element in an expanded state (e.g., interacting with a visual component of the UI element associated with a product) can change the UI element to a product focus state.

[0182] In some embodiments, interacting with a UI element can change the presentation of the UI element to a transaction state. Interacting with a UI element in a collapsed state can change the presentation of the UI element to a transaction state. Interacting with a UI element in an expanded state (e.g., selecting a product-related component) can change the presentation of the UI element to a transaction state. Interacting with a UI element in a product-focus state can change the presentation of the UI element to a transaction state. The determination of whether a selection or interaction with a UI element, UI element component, etc. results in a transition to a transaction state can be performed based on user history, user preferences, content items, content item feeds (e.g., search results, watch feeds, etc.), etc.

[0183] In some embodiments, UI elements may be overlaid on the presented content item. For example, UI elements identifying one or more products may be displayed overlaid on a video while the video is playing, while the video displays one or more products, etc. Selecting the overlaid UI element can cause additional UI elements to be displayed, change the overlaid UI element to a different state, change a separate UI element to a different state, etc. The overlaid UI element may be a collapsed state, an expanded state, a product-focus state, a transaction state, etc. The presence and / or location of the overlaid UI element may be determined by one or more trained machine learning models, e.g., one or more models configured to detect products.

[0184] FIG. 5F is a flow diagram of a method 500F for instructing a device to present one or more UI elements related to a product, according to some embodiments. At block 550, processing logic provides a UI to the device, including one or more graphical representations of one or more content items. The graphical representations are provided for display / presentation by the UI of the device. Each graphical representation of each content item is selectable to initiate presentation of the respective content item. The content items may include videos. The content items may include live streaming videos. The graphical representations may be provided in response to a request by the device. The graphical representations may include a home feed, a watch feed, a playlist, a search result list, a shopping feed, etc. Instructions sent to the device (e.g., including instructions related to any step of method 500F), UI sent to the device, UI elements sent to the device, etc. may be determined / selected based on obtaining a user history. The user history may include a history of interactions and / or selections with content items, including content items including related products. The user history may include a history of interactions and / or selections with UI elements or components of UI elements related to the product. The user history may include one or more searches of the user, for example, searches including product names. Instructions may be provided to the device in response to processing logic receiving the user history.

[0185] One or more of the graphical representations of the content items are displayed with a collapsed UI element. In some embodiments, each graphical representation is displayed with a collapsed UI element. In some embodiments, a subset of the graphical representations are displayed with a collapsed UI element. The collapsed UI element is presented / displayed with a first graphical representation of a first content item. The UI element includes information identifying multiple products covered by the first content item. The UI element may identify how many products are associated with the content item, identify one or more products by name, and identify a category or classification of products covered by the content item. In some embodiments, the multiple products are obtained as output from one or more trained machine learning models. The trained machine learning models may be similar to those described in connection with FIG. 3B.

[0186] At block 554, in response to receiving notification of a user interaction with the collapsed UI element, processing logic causes the device to change the presentation of the UI element. The presentation of the UI element can be changed from a collapsed state to an expanded state. The expanded UI element can include multiple visual components, each associated with one of multiple products. The visual components can include a photo, a name, a description, a price, a timestamp, etc.

[0187] At block 556, in response to receiving notification of a user selection of one of the plurality of visual components of the expanded UI element, processing logic facilitates presentation of the first content item. The processing logic may provide instructions to facilitate presentation of a portion of the first content item related to the first product, such as a product related to one of the plurality of visual components. The processing logic may provide instructions to display a portion of a video related to the product (e.g., instructions to start playing the video from a selected point in the video based on a timestamp associated with the product).

[0188] In some embodiments, the processing logic may further provide instructions to the device to change the presentation of the UI element to a product-focused state. For example, upon selection of a visual component of a UI element in an expanded state, the UI element may change to a product-focused state. The product-focused state may include additional details about one or more products covered by, included in, related to, etc. the content item.

[0189] In some embodiments, the processing logic may further provide instructions to the device to change the presentation of the UI element to a transaction state. The transaction state may facilitate a user initiating a transaction related to a product (e.g., purchasing a product). The transaction state may be presented in response to a user action, a user history, a user selection of one or more UI elements, etc. The UI element of the transaction state may include one or more components that facilitate a transaction related to one or more products.

[0190] 6 is a block diagram illustrating a computer system 600, according to some embodiments. In some embodiments, computer system 600 may be connected to other computer systems (e.g., via a network such as a local area network (LAN), an intranet, an extranet, or the Internet). Computer system 600 may operate in the capacity of a server or in a client computer in a client-server environment, or as a peer computer system in a peer-to-peer or distributed network environment. Computer system 600 may be provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a server, a network router, switch, or bridge, or any device capable of executing a set (series or otherwise) of instructions that specify actions to be taken by that device. Furthermore, the term "computer" is intended to include any collection of computers that, individually or in conjunction, execute a set (or sets) of instructions to perform any one or more of the methods described herein.

[0191] In a further aspect, computer system 600 may include a processing device 602, a volatile memory 604 (e.g., random access memory (RAM)), a non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.

[0192] The processing device 602 may be provided by one or more processors, such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of instruction set types), or a special-purpose processor (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).

[0193] Computer system 600 may further include a network interface device 622 (e.g., coupled to a network 674). Computer system 600 may also include a video display unit 610 (e.g., an LCD), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generating device 620.

[0194] In some embodiments, the data storage device 618 may include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable medium) capable of storing instructions 626 encoding any one or more of the methods or functions described herein, including instructions encoding the components of FIG. 1 (e.g., the content providing platform 120, other platforms of the content platform system 102, the communication application 115, the model 190, etc.) and for performing the methods described herein.

[0195] The instructions 626 may reside, completely or partially, within the volatile memory 604 and / or within the processing device 602 during execution thereof by the computer system 600, and thus the volatile memory 604 and the processing device 602 may also constitute machine-readable storage media.

[0196] Although the computer-readable storage medium 624 is shown as a single medium in the illustrated example, the term "computer-readable storage medium" is intended to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" also includes any tangible medium that can store or encode a set of instructions for execution by a computer, which instructions cause the computer to perform any one or more of the methodologies described herein. The term "computer-readable storage medium" includes, but is not limited to, solid-state memory, optical media, and magnetic media.

[0197] The methods, components, and features described herein may be implemented by discrete hardware components or integrated into the functionality of other hardware components, such as an ASIC, FPGA, DSP, or similar device. Furthermore, the methods, components, and features may be implemented by firmware modules or functional circuits within a hardware device. Furthermore, the methods, components, and features may be implemented in any combination of hardware devices and computer program components, or in a computer program.

[0198] Unless otherwise specified, terms such as "receive," "perform," "provide," "acquire," "perform," "access," "determine," "add," "use," "train," "reduce," "generate," "correct," and the like refer to actions and processes performed by or implemented by a computer system that manipulate and transform data represented as physical (electronic) quantities in computer system registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display devices. Also, as used herein, terms such as "first," "second," "third," "fourth," and the like are intended as labels to distinguish different elements and do not necessarily imply an ordering according to their numerical designations.

[0199] The embodiments described herein also relate to apparatus for performing the operations described herein. This apparatus may be specially constructed to perform the methods described herein, or may comprise a general-purpose computer system selectively programmed by a computer program stored on the computer system. Such a computer program may be stored on a computer-readable tangible storage medium.

[0200] The methods and illustrative examples described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized apparatus to perform the methods described herein and / or each of their individual functions, routines, subroutines, or operations. Example structures for a variety of these systems are set forth in the description above.

[0201] The foregoing description is intended to be exemplary, not limiting. While the present disclosure has been described with reference to particular illustrative examples and embodiments, it will be recognized that the present disclosure is not limited to the described examples and embodiments. The scope of the present disclosure should be determined with reference to the following claims, along with the full scope of equivalents to which such claims are entitled.

[0202] References throughout this specification to "one implementation" or "an implementation" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one implementation" or "in an implementation" in various places throughout this specification may, but do not necessarily, refer to the same embodiment, depending on the context. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0203] To the extent that the terms "includes," "including," "has," "contains," variations thereof, and other similar words are used in either the detailed description or the claims, these terms are intended to be inclusive in the same manner as the open transition word "comprising," without excluding additional or other elements.

[0204] As used in this application, terms such as “component,” “module,” or “system” are intended to generally refer to a computer-related entity, either hardware (e.g., circuitry), software, a combination of hardware and software, or an entity related to an operable machine having one or more specific functionalities. For example, a component can be, but is not limited to, a process running on a processor (e.g., a digital signal processor), a processor, an object, an executable file, a thread of execution, a program, and / or a computer. By way of example, both an application running on a controller and the controller can be a component. One or more components may reside within a process and / or thread of execution, and components may be localized on one computer and / or distributed among two or more computers. Furthermore, a “device” can be in the form of specially designed hardware, generalized hardware specialized by software executing on the hardware that enables the hardware to perform specific functions (e.g., generating points of interest and / or descriptions), software on a computer-readable medium, or a combination thereof.

[0205] The aforementioned systems, circuits, modules, etc. have been described with respect to interactions between multiple components and / or blocks. It will be understood that such systems, circuits, components, blocks, etc. may include those components or designated subcomponents, portions of the designated components or subcomponents, and / or additional components, according to various permutations and combinations of the foregoing. Subcomponents may also be implemented as components communicatively coupled to other components rather than being included in a parent component (hierarchical). In addition, it should be noted that one or more components may be combined into a single component providing integrated functionality or may be divided into several separate subcomponents, and that any one or more intermediate layers, such as a management layer, may be provided to communicatively couple such subcomponents to provide integrated functionality. Any component described herein may also interact with one or more other components not specifically described herein but known to those skilled in the art.

[0206] Moreover, the word "embodiment" or "exemplary" is used herein to mean serving as an example, example, or illustration. Any aspect or design described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other aspects or designs. Rather, use of the word "embodiment" or "exemplary" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X uses A or B" is intended to mean any natural inclusive permutation. That is, if X uses A, if X uses B, or if X uses both A and B, then "X uses A or B" is satisfied in each of the foregoing cases. Additionally, the articles "a" and "an," as used in this application and the appended claims, should generally be construed to mean "one or more" unless otherwise specified or clear from the context to refer to the singular form.

Claims

1. presenting a user interface (UI) including one or more graphical representations of one or more videos, each graphical representation of a respective video being selectable to initiate playback of the respective video and displayed with a collapsed UI element, the UI element including information identifying a plurality of products associated with the respective video; continuing to present a graphical representation of the respective video while changing the display of the UI element from the collapsed state to an expanded state in response to a user interaction with the UI element in the collapsed state, the UI element in the expanded state comprising a plurality of visual components each associated with one of the plurality of products; and in response to a user selection of one of the plurality of visual components of the UI element in the expanded state, starting playback of the respective video covering a product associated with the selected visual component.

2. 10. The method of claim 1, further comprising: in response to the user selection of the one of the plurality of visual components of the UI element in the expanded state, changing the presentation of the UI element from the expanded state to a product-focus state, wherein the UI element in the product-focus state includes detailed information about the product associated with the selected visual component.

3. providing the search query to a content platform; The method of claim 1 , further comprising: receiving the one or more graphical representations of the one or more videos from the content platform.

4. providing a user request to present a first video to a content platform; The method of claim 1 , further comprising: receiving, from the content platform, the one or more graphical representations of the one or more videos as recommended videos related to the first video.

5. The UI element in the expanded state is a first tab associated with the plurality of products; a second tab associated with each of the plurality of portions of the video, the second tab presenting the content of the first tab by default; The method of claim 1 further comprising:

6. 10. The method of claim 1, further comprising: in response to a user selection of a component of the UI element associated with the one of the plurality of products, changing the presentation of the UI element to a transaction state, the UI element in the transaction state facilitating a transaction associated with the one of the plurality of products.

7. The method of claim 1 , wherein the UI element in the collapsed state is overlaid on the graphical representation of the respective video or one or more of the respective videos.

8. The method of claim 1 , wherein the user interaction with the UI element in the collapsed state includes hovering over the UI element in the collapsed state.

9. presenting a user interface (UI) including one or more graphical representations of one or more content items for presentation on a user device, each graphical representation of a respective content item being selectable to initiate presentation of the respective content item and displayed with a collapsed UI element, the UI element including information identifying a plurality of products associated with the respective content item; in response to receiving notification of a user interaction with the UI element in the collapsed state, changing a presentation of the UI element from the collapsed state to an expanded state, the UI element in the expanded state comprising a plurality of visual components each associated with one of the plurality of products; and facilitating presentation of a first content item in response to receiving notification of a user selection of one of the plurality of visual components of the UI element in the expanded state.

10. 10. The method of claim 9, further comprising: in response to receiving the notification of the user selection of the one of the plurality of visual components of the UI element in the expanded state, providing instructions to the user device that facilitate changing the presentation of the UI element from the expanded state to a product-focus state, wherein the UI element in the product-focus state includes detailed information about the product associated with the selected visual component.

11. 10. The method of claim 9, wherein the identifiers of the plurality of products associated with the first content item are obtained as output from a trained machine learning model, the trained machine learning model being configured to receive as input data associated with the first content item.

12. 10. The method of claim 9, further comprising: providing instructions to the user device in response to a user selection of a component of the UI element associated with one of the plurality of products, the instructions facilitating changing the presentation of the UI element to a transaction state, the UI element of the transaction state comprising one or more components facilitating a transaction associated with the one of the plurality of products.

13. 10. The method of claim 9, wherein facilitating the presentation of the first content item includes an announcement of a portion of the first content item to present, the portion of the first content item covering one of the plurality of products related to the one of the plurality of visual components.

14. 10. The method of claim 9, further comprising obtaining a user history, the user history including one or more interactions with a content item having an associated product, and wherein providing instructions to display the UI element is performed in response to obtaining the user history.

15. The method of claim 9 , wherein the first content item comprises a live-streamed video.

16. A non-transitory machine-readable storage medium storing instructions that, when executed, cause a processing device to: presenting a user interface (UI) including one or more graphical representations of one or more videos, each graphical representation of a respective video being selectable to initiate playback of the respective video and displayed with a collapsed UI element, the UI element including information identifying a plurality of products associated with the respective video; continuing to present a graphical representation of the respective video while changing the display of the UI element from the collapsed state to an expanded state in response to a user interaction with the UI element in the collapsed state, the UI element in the expanded state comprising a plurality of visual components each associated with one of the plurality of products; and in response to a user selection of one selected visual component of the plurality of visual components of the UI element in the expanded state, starting playback of the respective video covering a product associated with the selected visual component.

17. 17. The non-transitory machine-readable storage medium of claim 16, wherein the operations further include, in response to the user selection of the one of the plurality of visual components of the UI element in the expanded state, changing the presentation of the UI element from the expanded state to a product-focus state, the UI element in the product-focus state including detailed information about the product associated with the selected visual component.

18. 17. The non-transitory machine-readable storage medium of claim 16, wherein the operations further include, in response to a user selection of a component of the UI element associated with the one of the plurality of products, changing the presentation of the UI element to a transactional state, the UI element in the transactional state facilitating a transaction associated with the one of the plurality of products.

19. The UI element in the expanded state is a first tab associated with the plurality of products; 17. The non-transitory machine-readable storage medium of claim 16, further comprising: a second tab associated with each of the plurality of portions of the video, the second tab presenting content of the first tab by default.

20. 17. The non-transitory machine-readable storage medium of claim 16, wherein initiating playback of the respective video includes presenting a portion of the respective video that depicts the product associated with the selected visual component.

Citation Information

Patent Citations

  • Interactive video overlay

    CA3188919A1

  • Generation and presentation of interactive information cards for a video

    US10620801B1

  • Interactive content generation

    US20160042251A1

  • Interactive digital media playback and transaction platform

    US20210127170A1

  • Interactive video overlay

    US20220053233A1