Product Identification within Media Items
The system uses machine learning to analyze media items for associated products, improving user experience by reducing latency and simplifying product discovery and purchase.
Patent Information
- Application Number
- JP2025505866
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-01
- Filing Date
- 2023-07-31
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional systems face difficulties in efficiently identifying and providing information about products associated with media items, such as videos, which is time-consuming, error-prone, and often outdated, leading to increased computing resources and user latency.
A system utilizing machine learning models to analyze metadata and image data of media items to identify associated products, update metadata, and provide user interfaces for seamless product discovery and purchase.
Enhances user experience by accurately and efficiently identifying products within media items, reducing latency and computing resources, and simplifying the product discovery and purchase process.
Smart Images

Figure 2025525169000001_ABST
Abstract
Description
Technical Field
[0001] Aspects and implementations of the present disclosure relate to methods and systems for facilitating the pairing of media items and associated objects, and more particularly, to a system for identifying pairings of media items and products.
Background Art
[0002] A platform (e.g., a content sharing platform) can transmit (e.g., stream) media items to client devices connected to the platform via a network. Different types of client devices can be optimized for different tasks or preferred by users for different tasks, etc. A media item can include a reference to one or more products.
Summary of the Invention
[0003] The following summary is a simplified summary of the present disclosure to provide a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the present disclosure. It is not intended to identify key or critical elements of the present disclosure or to define the scope of any particular implementation of the present disclosure or the scope of any claims. Its sole purpose is to present some concepts of the present disclosure in a simplified form as a prelude to a more detailed description that will be presented later.
[0004] A system and method are disclosed for facilitating the identification of one or more products associated with a media item. In some implementations, the method includes obtaining first data. The first data includes a first identifier of a first product determined in relation to a content item based on first metadata of the content item. The first data further includes a first confidence value associated with the first product and the content item. The method further includes obtaining second data. The second data includes a second identifier of the first product determined in relation to the content item based on first image data of the content item. The second data further includes a second confidence value associated with the first product and the content item. The method further includes providing the first data and the second data to a trained machine learning model. The method further includes obtaining a third confidence value associated with the first product from the trained machine learning model. The method further includes adjusting second metadata associated with the content item in consideration of the third confidence value.
[0005] In some embodiments, the method further includes providing the first metadata of the content item as an input to a second model. The method may include obtaining the first data as an output of the second model. In some embodiments, the method further includes providing the first image data of the content item as an input to the second mode. The method further may include obtaining first dimensionally reduced data as an output of the second model. The method further may include obtaining second dimensionally reduced data associated with the first product from a data store. The second dimensionally reduced data may be obtained from the data store in response to obtaining the first data. The second data may be generated based at least on the first dimensionally reduced data and the second dimensionally reduced data.
[0006] The method may further include providing the second image data to a third model. The method may further include obtaining, from the third model, a third identifier of the first product. The second dimensionally reduced data may be obtained from a data store in response to obtaining the third identifier of the first product. The second data may be generated based at least on the first dimensionally reduced data and the second dimensionally reduced data.
[0007] In some embodiments, the content item is a video. The first data may further include an indication of the timestamp of one or more frames of the video associated with the product. Adjusting the second metadata may include including an indication of the first product and an indication of the timestamp in the second metadata.
[0008] In some embodiments, the metadata may include a title of the content item. The metadata may include a description of the content item. The metadata may include a caption associated with the content item. Adjusting the second metadata may include adjusting the caption associated with the product.
[0009] In some embodiments, the method further includes training a machine learning model to generate a trained machine learning model. Training the machine learning model can include receiving image-based product data associated with a plurality of content items. The image-based product data can include one or more products detected in the image and an indication of one or more product image confidence values. Training can further include receiving metadata-based product data associated with a plurality of content items. The metadata-based product data can include one or more products detected in the text and an indication of one or more confidence values. Training can further include receiving data indicating products included in a plurality of content items. Training can further include providing the image-based product data and the metadata-based product data to the machine learning model as training inputs. Training can further include providing data indicating products included in a plurality of content items to the machine learning model as target outputs.
[0010] In some embodiments, the method further includes receiving third data including a third identifier of a first product category associated with a content item. The method can further include providing the third data to the trained machine learning model. The third confidence value can be based on the first data, the second data, and the third data.
[0011] In another aspect, a method includes obtaining first metadata associated with a content item. The method further includes providing the first metadata to a first model. The method further includes obtaining a first product identifier as an output of the first model based on the first metadata and a first confidence value associated with the first product identifier. The method further includes obtaining image data of the content item. The method further includes providing the image data to a second model. The method further includes obtaining a second product identifier as an output of the second model based on the image data and a second confidence value associated with the second product identifier. The method further includes providing data as input to a third model, the data including the first product identifier, the first confidence value, the second product identifier, and the second confidence value. The method further includes obtaining a third product identifier and a third confidence value as an output of the third model. The method further includes adjusting the metadata associated with the content item in consideration of the third product identifier and the third confidence value.
[0012] In some embodiments, generating the second confidence value includes reducing the dimensionality of the image data to generate first dimensionally reduced data. Generating the second confidence value may further include retrieving the second dimensionally reduced data from a data store. The second dimensionally reduced data may be associated with a product identified by a second product identifier. Generating the second confidence value may further include performing one or more operations to generate the second confidence value. The second confidence value may be based on one or more differences between the first dimensionally reduced data and the second dimensionally reduced data.
[0013] In some embodiments, the first product identifier, the second product identifier, and the third product identifier each identify a first product.
[0014] In some embodiments, the method further includes obtaining the image data and a timestamp associated with the content item. Adjusting the second metadata associated with the content item may include adjusting the second metadata to include an indication that a product identified by the second product identifier is associated with the timestamp and the content item.
[0015] In some embodiments, the second metadata includes a machine-generated caption. A first product identifier may be associated with the product. Language associated with the product may have been inaccurately transcribed during generation of the machine-generated caption. Updating the second metadata associated with the content item may include replacing a portion of the machine-generated caption associated with the product with a text identifier for the product.
[0016] In some embodiments, the method further includes providing a fourth product identifier and the fourth confidence value to a third model. The method may further include obtaining a fifth product identifier as an output of the third model. The third product identifier may be associated with the first product and the fifth product identifier may be associated with the second product.
[0017] In another aspect, the non-transitory machine-readable storage medium stores instructions that, when executed, cause the processing device to perform operations including obtaining first data. The first data includes a first identifier of a first product determined in relation to a content item based on first metadata of the content item. The first data further includes a first confidence value associated with the first product and the content item. The operations further include obtaining second data including a second identifier of the first product determined in relation to the content item based on first image data of the content item. The second data further includes a second confidence value associated with the first product and the content item. The operations further include providing the first data and the second data to a trained machine learning model. The operations further include obtaining a third confidence value associated with the first product from the trained machine learning model. The operations further include adjusting second metadata associated with the content item in consideration of the third confidence value.
[0018] In some embodiments, the operations further include providing first image data of the content item as an input to a second model. The operations may further include obtaining dimensionally reduced first data as an output of the second model. The operations may further include obtaining second dimensionally reduced data associated with the first product from a data store. The data from the data store may be obtained in response to obtaining the first data. The second data may be based at least on the dimensionally reduced first data and the second dimensionally reduced data.
[0019] In some embodiments, the content item is a video. The first data may further include an indication of a timestamp of one or more frames of the video associated with the product. Adjusting the second metadata may include including an indication of the first product and an indication of the timestamp in the second metadata.
[0020] In some embodiments, the operation further includes receiving third data including a third identifier of a first product category associated with the content item. The operation may further include providing the third data to a trained machine learning model. The third confidence value may be generated based on the first data, the second data, and the third data.
[0021] Optional features of one aspect may be combined with other aspects as needed.
[0022] Aspects and embodiments of the present disclosure will be more fully understood from the following detailed description of various aspects and embodiments of the present disclosure and from the accompanying drawings, which should not be construed as limiting the present disclosure to specific aspects or embodiments, but are for illustrative and understanding purposes only. BRIEF DESCRIPTION OF THE DRAWINGS
[0023]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 5F
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0024] Aspects of the present disclosure relate to methods and systems for facilitating the pairing of media items (e.g., content items, content, etc.) and associated products. A platform (e.g., a content sharing platform, etc.) can enable a user to access media items (e.g., video items, audio items, etc.) hosted by the platform (e.g., via a client device connected to the platform). The platform can provide access to the media item to the user's client device via a network (e.g., the Internet) (e.g., by transmitting the media item to the user's client device, etc.). The media / content item can have one or more additional associated activities, and the one or more additional associated activities can deepen the user's engagement with the content item. For example, the engagement activities can include a comment section for the content item, a live chat associated with the content item, etc. Some content items can be related to one or more products. For example, the product can be displayed as part of the content item, the product can be reviewed within the content item, or the content item can be associated with one or more products (e.g., sponsored by the company associated therewith), etc.
[0025] In conventional systems, it can be difficult, inconvenient, time-consuming, etc. to identify products associated with content items. A content platform (e.g., a content delivery platform for presentation to users) can have difficulty in identifying products associated with content items. In some systems, a content creator can identify one or more products, a content channel, or a list of content items associated with a content item. In some systems, a content creator can include information regarding one or more products within a content item, such as a photo of an item, an item appearing within a video, etc. In some systems, a content creator can include information regarding one or more products within one or more fields associated with a content item, such as a content item title, a content item description, a content item comment (e.g., a fixed comment), etc.
[0026] In conventional systems, a user may not be able to easily confirm whether a content item has one or more associated products. A content item may not have an indicator of associated products. A user may need to be presented with the content item (e.g., watch a video) to confirm whether the content item has one or more associated products.
[0027] In conventional systems, a user may have difficulty identifying products associated with content items. There may be no direct indication that a content item has an associated product (e.g., while selecting a content item for consumption, viewing, etc.). For example, it can be difficult to discover product information spread across different fields such as a content item title, content item description, comment section, etc. Product information may be included in the content item; for example, a video may include audio that describes one or more products, an image item may include an image of one or more products, and so on. Extracting this information from the content item can be difficult, time-consuming, error-prone, and so on.
[0028] In conventional systems, it can be difficult and time-consuming to receive additional information about products characterized within a content item, and so on. In some systems, a product may appear within a content item (e.g., within a video). A user may not be given additional information (e.g., product name, shipper name, etc.) and may perform a search independently of the content platform to learn more about the product. A name and / or shipper associated with the product may be provided (e.g., within the description or title of the content item). A user may perform a search to learn more about the product, such as product variations, related products, or availability and price (e.g., independently of the content platform). Instructions may be provided to facilitate learning more about the product (e.g., instructions to the user, instructions to the processing device in the form of a link to the shipper website, etc.). A user may receive additional information from a separate source (e.g., a website) independently of the content platform.
[0029] In conventional systems, information about products included in content items can become outdated. For example, within a content item or related fields (such as content item title, content item description, content item comment, etc.), a content creator may include additional information about one or more products, such as price information, shipper information, availability information, alternative version or variation information, etc. Some information included in the content item or associated fields may not be updated by changes to this information, or may depend on updates by the content creator, and so on.
[0030] In conventional systems, there can be obstacles such as making it difficult, time-consuming, and cumbersome for a user to purchase one or more items associated with a content item. The user can search for products, search for shippers that stock or sell the products, and so on. In some embodiments, the content item or associated fields may include an instruction to purchase an item (for example, the description of the content item may include one or more links to products associated with the content item). The user may be directed to another platform independent of the content platform to complete the purchase of one or more products.
[0031] It may take a significant amount of time and computing resources for a user to find information about products covered by a content item. For example, a video can be long and may characterize the product of interest towards the end. It may take a significant amount of time for the user to consume the video to obtain accurate information about the product of interest, which as a result increases the use of the computing resources of the client device. In addition, the computing resources of the client device that enable the user to consume a media item may not be available for other processes, which may reduce overall efficiency and may increase the overall latency of the client device.
[0032] Aspects of the present disclosure can address one or more of those drawbacks of conventional methods. In some embodiments, aspects of the present disclosure can enable automated identification of products characterized within and / or included in content items. In some embodiments, aspects of the present disclosure can enable the use of a model to identify products from text associated with a content item. The text can include the title of the content item, the description of the content item, the caption associated with the content item, and the like. The caption can be a machine-generated caption, for example, a caption generated by a speech-to-text model for a video or audio content item. The model can be a machine learning model. The model can output a confidence value indicating the likelihood that the content item includes a product.
[0033] In some embodiments, aspects of the present disclosure can enable the use of a model to identify products from an image of a content item. The content item can be a picture or can include a picture. The content item can be a video or can include a video. One or more images of the content item can be provided to a model configured to identify products from the images. The model can reduce the dimensions of the images. The model can search for similar images associated with the product. The model can search for images with reduced dimensions in a reduced-dimensional space. The model can be a machine learning model. The model can output a confidence value indicating the likelihood that the content item includes a product.
[0034] In some embodiments, aspects of the present disclosure enable the use of a model (e.g., a fusion model) to confirm whether a product is included in a content item. The fusion model may receive indications of one or more products detected by a model that receives as input text associated with the content item. The fusion model may receive an indication of a confidence value that one or more products are included in the text. The fusion model may receive indications of one or more products detected by a model that receives as input an image of the content item. The fusion model may receive an indication of a confidence value that one or more products are included in the image. The fusion model may determine a confidence level that one or more products appear within the content item and associated data (e.g., title, description, etc.). The fusion model may be a machine learning model.
[0035] In some embodiments, product detection may be utilized to improve content or associated information (e.g., description, caption). The model may be provided with data associated with the content item. For example, the model may be provided with a machine-generated caption of the content item. The model may identify one or more captions that may be false representations of product names (e.g., a machine-generated caption may include the closest English equivalent of the spoken product name). The model may be provided with one or more images of the content item. The model may determine the likelihood that one or more products are included in the image. Information associated with the content item (e.g., metadata, description, caption, etc.) may be updated considering the one or more detected products.
[0036] In some embodiments, aspects of the present disclosure enable an indicator of a content item having one or more associated products. A list of content items may include one or more indicators that one or more of the content items in the list include an associated product. The indicator may include a visual indicator (e.g., a "shopping" symbol or text indicating one or more products displayed in association with the content item), or an additional field (e.g., a panel containing product information), etc. In some embodiments, the list of content items may be presented to a user via a user interface (UI). The UI elements may be associated with a content platform (e.g., presented via an application associated with the content providing platform). The UI may include elements indicating that the content item is associated with one or more products. In some embodiments, user interaction with the UI elements may result in the presentation of additional UI elements having additional information, e.g., a list of products associated with the content item.
[0037] In some embodiments, aspects of the present disclosure enable identification by a user of products associated with a content item. In some embodiments, a list of content items for presentation to a user may include a list of products associated with one or more content items. For example, the list of content items may be displayed to the user via a UI. The UI may include one or more UI elements presenting to the user one or more products associated with the content item. The UI elements may be provided via an application associated with the content providing platform of the content item. In some embodiments, in response to detecting user interaction with a UI element indicating that a content item has an associated product, a UI element listing the products associated with the content item may be presented. In some embodiments, user interaction with a product in the list of products may cause the UI to present additional information to the user.
[0038] In some embodiments, aspects of the present disclosure enable providing additional information regarding one or more products associated with a content item to a user. A UI element for providing additional information regarding the product may be provided to the user. The UI element may provide information such as product variations (e.g., color variations, size variations, etc.), related products, availability and / or price (e.g., associated with one or more vendors). The UI element may be provided via an application, a content delivery platform, etc. associated with the content item. In response to user interaction with other UI elements, such as selecting a content item for viewing, selecting a UI element indicating one or more associated products, etc., a UI element for displaying additional product information may be presented. In some embodiments, upon user interaction with the UI element, additional UI elements may be provided, for example, to facilitate purchasing the product.
[0039] In some embodiments, aspects of the present disclosure enable automatic updating of information related to one or more products associated with a content item. In some embodiments, a content delivery platform associated with the content item may include one or more memory devices containing product data, may communicate with one or more memory devices, may be connected to one or more memory devices, etc. For example, the content platform may maintain and update a database of information associated with the product, and changes made to the database may be reflected within the UI elements presented to the user.
[0040] In some embodiments, aspects of the present disclosure enable simplified purchasing procedures for users. When an indication of the user's intention to purchase a product is received (e.g., upon a user interaction with a UI element associated with the product), UI elements for facilitating the purchase of the product may be presented to the user. In some embodiments, the user may purchase the product via an application associated with the content item, a content delivery platform, etc. In some embodiments, the user may be directed to one or more external shippers (e.g., shipper applications, shipper web pages, etc.) to purchase the product.
[0041] In some embodiments, in response to various operations of an application (e.g., an application associated with a content platform), UI elements associated with the product may be provided. The UI elements associated with the product may be provided as part of a list of content items adapted to the user, e.g., adapted to the user account associated with the user. The UI elements associated with the product may be provided as part of a list of content times related to the previously presented content items, e.g., a list of what to watch next, a list of recommended videos, etc. The UI elements associated with the product may be provided as part of a list of content items generated in response to a user search. The UI elements associated with the product may be provided as part of a product-specific list, e.g., part of the shopping section of an application associated with a content platform. Inclusion of the UI elements associated with the product, the content items associated with the product, the number of UI elements or content items presented in association with the product, etc. may be based on several metrics. The metrics may include the user viewing history, the user search history, etc.
[0042] In some embodiments, aspects of the present disclosure may enable immediate access to one or more portions of a content item related to a product. For example, a UI element may include a list of products associated with the content item. One or more of the list of products may be associated with a portion of the content item, such as a timestamp of a video content item. Upon interaction with a portion of the UI element associated with the product, the content item may present the associated portion of the content item (e.g., the video may begin playing the portion of the video associated with the content item). In some embodiments, the list of products associated with the content item may be updated as the content item is presented. For example, the list of products associated with a video may be reorganized as the video is played. For example, the product currently highlighted by the video may be at the top of the list of products, and the products currently on the screen may be grouped together within the product presentation UI element, and so on.
[0043] Aspects of the present disclosure may provide technical advantages over previous solutions. Aspects of the present disclosure may enable automatic product detection within content items. This may improve the experience of content creators, for example, by automatically associating one or more products with a content item (e.g., removing the burden of associating products with content items from the creator), by automatically associating a product with a content item. Due to the use of multiple sources (e.g., text and image searches of products), fusion models, etc., the automatically generated product associations may be associations with improved accuracy. Model-based product detection may be utilized to improve content items, associated data, etc. For example, utilizing object detection may improve the machine-generated captions or descriptions of content items. More accurate captions may improve the user's experience when consuming content items. Those improvements may reduce the time required by content creators to generate accurate content, reduce the time spent by users to discover content items and / or products of interest to them, increase the accuracy of machine-generated information associated with content items, etc. Thus, computing resources in content creators, content viewers, and / or client devices associated with the platform are reduced and available for other processes, which increases the overall efficiency of the system and reduces the overall latency of the system.
[0044] Aspects of the present disclosure can improve the user's experience during activities such as searching for content items, browsing content items, and scrolling through a list of content items. For example, a user may be able to identify content items having associated products from a list of content items without the content items being presented. The user may be provided with a seamless way to increase engagement with products associated with content items. For example, an interaction with a UI element indicating the presence of a product associated with a content item may result in the presentation of a content item having further information about one or more products, and further interaction may facilitate purchasing one or more products, and so on. The presentation of one or more UI elements can streamline the user's experience. For example, a user may be able to easily retrieve additional information within an application associated with the content platform / content item. A user may be able to more easily purchase products associated with content items. A user may be directed to relevant portions of a content item based on an expressed interest within a product associated with the content item. A user may be able to easily view price information, availability information, product variations, related products, etc. within the context of a single application. Such implementations can reduce the user's time and frustration, simplify the shopping and / or purchasing process, or simplify the product research, review, and / or selection process, and so on.
[0045] Figure 1 illustrates a system architecture 100 of an example for providing content and associated product information according to some embodiments. The system architecture 100 includes client devices 110, one or more networks 105, a content platform system 102, and a product identification system 175. The content platform system 102 includes one or more server machines 106 and one or more data stores 140, and may include various platforms aimed at executing tasks (for example, tasks associated with content delivery). The platforms of the content platform system 102 may be hosted by one or more server machines 106. The platforms of the content platform system 102 may include one or more computing devices (rack-mounted servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.) and one or more data stores (for example, hard disks, memories, and databases), and / or may be hosted thereon and may be coupled to one or more networks 105. In some embodiments, components of the content platform system 102 (for example, server machines 106, data stores 140, hardware associated with one or more platforms, etc.) may be directly connected to one or more networks 105. In some embodiments, one or more components of the content platform system 102 may access the network 105 via other devices, such as hubs, switches, etc. In some embodiments, one or more components of the content platform system 102 may communicate directly with components of the product identification system 175, such as other components represented in Figure 1, for example, server machines 170 and / or 180, etc. The data store 140 may be included in one or more server machines 106, including external data storage, etc.The platform of the content platform system 102 may include an advertising platform 165, a social network platform 160, a recommendation platform 157, a search platform 145, and a content providing platform 120. The product identification system 175 includes a set of a server machine 170, a server machine 180, and a model 190. The product identification system 175 may include additional devices, such as a data store, additional servers, and the like. Multiple operations of the product identification system 175 may be executed by a single physical or virtual device.
[0046] One or more networks 105 may include one or more public networks (e.g., the Internet), one or more private networks (e.g., a local area network (LAN), a wide area network (WAN), one or more wired networks (e.g., an Ethernet network), one or more wireless networks (e.g., an 802.11 network), one or more cellular networks (e.g., a long term evolution (LTE) network), routers, hubs, switches, server computers, and / or combinations thereof. In one implementation, some components of the architecture 100 are not directly connected to each other. In one implementation, the system architecture 100 includes separate networks 105.
[0047] One or more data stores 140 may be present in a memory (e.g., random access memory), cache, drive (e.g., hard drive), flash drive, etc., and may be part of one or more database systems, one or more file systems, or other types of components or devices having the ability to store data. One or more data stores 140 may also span multiple computing devices (e.g., multiple server computers) and may include multiple storage components (e.g., multiple drives or multiple databases). A data store may be persistent storage having the ability to store data. Persistent storage may be a local storage unit or a remote storage unit, an electronic storage unit (e.g., main memory), or a similar storage unit. Persistent storage may be a monolithic device or a distributed set of devices.
[0048] Content items 121A - C (e.g., media content items) may be stored in one or more data stores. The data store may be part of one or more platforms. Examples of content items 121 may include, but are not limited to, digital videos, digital movies, animated images, digital photos, digital music, digital audio, digital video games, co - created media content presentations, website content, social media updates, e - books, e - journals, digital audio books, web blogs, software applications, etc. Content items 121A - C may also be referred to as media items. Content items 121A - C may be pre - recorded or may be live - streamed. For clarity and simplicity, video may be used throughout this specification as an example of a content item 121 (e.g., content item 121A). Video may include pre - recorded video, live - streamed video, short - form video, etc.
[0049] Content items 121A - C can be provided by a content provider. The content provider can be a user, a company, an organization, etc. The content provider can provide content item 121 which is a video (e.g., content item 121A). The content provider can provide content item 121 including live - streamed content. For example, content item 121 can include a live - streamed video, a live chat associated with the video, etc.
[0050] The client device 110 can include devices such as a television, a smartphone, a mobile information terminal, a portable media player, a laptop computer, an e - book reader, a tablet computer, a desktop computer, a game console, or a set - top box.
[0051] The client device 110 may include a communication application 115. The content item 121 (e.g., content item 121A) may be consumed by a user via the communication application 115. For example, the communication application 115 may access one or more networks 105 (e.g., the Internet) via the hardware of the client device 110 to provide the content item 121 to the user. As used herein, "media", "media item", "online media item", "digital media", "digital media item", "content", "media content item", and "content item" may include an electronic file that can be executed or loaded using software, firmware, and / or hardware configured to present the content item. In one implementation, the communication application 115 may be an application that enables a user to compose, transmit, and receive a content item 121 (e.g., a video) through a platform (e.g., content providing platform 120, recommendation platform 157, social network platform 160, and / or search platform 145), and / or a combination of platforms and / or networks.
[0052] In some embodiments, communication application 115 can be a social networking application, a video sharing application, a video streaming application, a video game streaming application, a photo sharing application, a chat application, or a combination of such applications (or include aspects thereof). The communication application 115 associated with the client device 110 can render, display, present, and / or play one or more content items 121 to one or more users. For example, the communication application 115 can provide a user interface 116 (e.g., a graphical user interface) that is to be displayed on the endpoint device 110 for receiving and / or playing video content. In some embodiments, the communication application 115 is associated with (and managed by) the content platform 102 or the content delivery platform 120.
[0053] In some embodiments, the communication application 115 may include a content viewer 113 and related product components 114. The user interface 116 (UI) may display the content viewer 113 and related product components 114. The related product components 114 may be used to display UI elements that display information about one or more products (e.g., one or more products associated with a content item). The related product components 114 may display UI elements to notify the user that the content item has one or more associated products (e.g., the related product components 114 may cause the display of a "shopping" symbol displayed proximate to the content viewer 113, the related product components 114 may cause elements to be displayed proximate to an element for selecting a content item for presentation from a list of content items, which may indicate that the content item has one or more associated products, etc.). By the related product components 114, the UI may display information about one or more products associated with the content item (e.g., the related product components 114 may include a list of products associated with the content item, may include an image of a product associated with the content item, may include information connecting the product to the content item such as price information, or a time stamp of a portion of a video related to the product, etc., which may cause the display of the UI). By the related product components 114, the UI elements may display additional information about one or more products, such as product variations (e.g., color variations, size variations, etc.), related products, recommended products, etc. By the related product components 114, the UI elements may display one or more options for purchasing one or more products. In some embodiments, the user may be able to navigate between different views provided by the related product components 114. Further descriptions of UI elements of examples associated with products related to a content item may be found in relation to FIGS. 4A - E.In some embodiments, the related product component 114 may display more than one UI element, for example, when more than one content item having an associated product is displayed. In some embodiments, the related product component 114 may not display UI elements indicating related products, for example, based on user settings, user preferences, user history, and the like. In some embodiments, the plurality of content viewers 113 and / or related product components 114 may be associated with one user interface, communication application, client device, and the like. For example, a plurality of content items may be immediately displayed to the user. In some implementations, the communication application 115 can access content (such as web pages, e.g., Hypertext Markup Language (HTML) pages, digital media items), retrieve the content, present the content, and / or navigate the content, and can be a web browser that includes the related product component 114 and the content viewer 113, and the related product component 114 and the content viewer 113 can be an embedded media player embedded within a user interface 116 (such as a web page associated with the content being viewed) provided by the content delivery platform 120. Alternatively, the application 115 is not a web browser, but a stand-alone application (such as a mobile application, desktop application, game console application, television application, etc.) downloaded from a platform (such as the content delivery platform 120, recommendation platform 157, social network platform 160, or search platform 145) or pre-installed on the client device 110. The stand-alone application 115 can provide a user interface 116 that includes the content viewer 113 (such as an embedded media player) and the related product component 114.
[0054] In some embodiments, the content platform system 102 may include a product information platform 161 (e.g., hosted by the server machine 106). The product information platform 161 may store, retrieve, provide, receive, etc., data related to one or more products associated with one or more content items. The content platform system 102 may provide data to the associated product component 114 of the client device 110. The product information platform 161 may include information provided by a content creator. For example, the content creator may provide a list of products associated with the content item that the content creator provided to the content platform system 102. The product information platform 161 may include information provided by one or more users. For example, one or more users may identify a product associated with a content item, for example, in response to being presented with the content item. The product information platform 161 may include information provided by a product identification system 175. For example, one or more machine learning models may be utilized to identify products characterized within a content item, and provide an indication of the associated products and content items to the product information platform 161.
[0055] In some embodiments, the communication application 115 installed on the client device 110 may be associated with a user account. For example, the user may sign into the account on the client device 110. In some embodiments, multiple client devices 110 may be associated with the same client account. In some embodiments, providing information about the association(s) of a product with one or more content items may be performed conditional on a user account, e.g., account settings, account history (e.g., history involving UI elements including associated product information), etc.
[0056] In some embodiments, client device 110 may include one or more data stores. The data store may include commands (e.g., instructions that, when executed by a processing device, cause an operation to occur) for rendering a UI (e.g., user interface 116). The instructions may include instructions for rendering interactive components, e.g., UI elements with which a user may interact to present additional information about one or more products associated with a content item. In some embodiments, the instructions may cause the processing device to render UI elements that present information about one or more products associated with one or more content items (e.g., the number of videos reviewing a product of interest to the user may be presented by UI elements that present more information about the product).
[0057] In some embodiments, one or more server machines 106 may include computing devices such as rack-mounted servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc., and may be coupled to one or more networks 105. Server machine 106 may be a stand-alone device or may be part of any of a platform (e.g., content delivery platform 120, social network platform 160, etc.).
[0058] The social network platform 160 may provide an online social networking service. The social networking platform 160 may provide a communication application 115 for users to create profiles and perform activities with those profiles. Activities may include updating profiles, exchanging messages with other users, evaluating status updates, photos, videos, etc. (e.g., liking, commenting, sharing, recommending), and receiving notifications associated with the activities of other users. In some embodiments, additional product information (such as provided by the product information platform 161) may be shared by a user with one or more additional users via the social network platform 160.
[0059] The recommendation platform 157 may be used to generate and provide content recommendations (e.g., articles, videos, posts, news, games, etc.). The recommendations may be based on search history, content consumption history, followed / subscribed channel content, linked profiles (e.g., friend lists), popular content, etc. The recommendation platform 157 may be utilized, for example, to generate a user home feed, a user watch list, a user play list, etc. One or more UI elements that indicate associated products, display a list of associated products, present product information, and present one or more options to purchase a product may be presented as part of, in relation to, or accessible from, etc., a home feed, a watch list, a play list, a next-to-watch list, etc., in combination with, and related to, those. The presentation of one or more UI elements may be performed based on user history, user settings, data indicating user preferences (e.g., demographic data), etc.
[0060] The search platform 145 can be used to enable a user to query one or more data stores 140 and / or one or more platforms and receive query results. The search platform 145 can be utilized by a user to search for content items or to search for topics. For example, the search platform 145 can be utilized by a user to search for content items having one or more associated products. The search platform 145 can be utilized to search for content items related to a product type (e.g., a headphone review video). The search platform 145 can be utilized to search for content items related to a certain product (e.g., headphones of a particular manufacture and / or model). One or more UI elements can be displayed to the user in response to receiving a search query. The type, style, etc. of the UI elements displayed can be based on the content of the user's search. For example, in response to a search for content related to a product type (e.g., headphones), a UI element indicating that the content items suggested in view of the search have one or more associated products can be displayed. As a further example, in response to a search for content items related to a more specific product (e.g., "the best headphones for podcasts"), a different UI element providing information about the products associated with the content items can be displayed. As a further example, in response to a search for content items related to a specific product (e.g., a particular manufacture and / or model), a different UI element providing specific information about the searched product indicating that the product is associated with the content items recommended in view of the search can be displayed.
[0061] The content providing platform 120 can be used to provide access to content item 121 to one or more users and / or provide content item 121 to one or more users. For example, the content providing platform 120 can enable a user to consume, upload, download, and / or search for content item 121. In other embodiments, the content providing platform 120 can enable a user to rate, recommend, share, evaluate, and / or comment on content item 121, such as approving it ("liking"), disapproving it, etc. In other embodiments, the content providing platform 120 can enable a user to edit content item 121. The content providing platform 120 can also include a website (e.g., one or more web pages) and / or one or more applications (e.g., communication application 115) that can be used to provide access to content item 121 to one or more users. For example, communication application 115 can be used by client device 110 to access content item 121. The content providing platform 120 can include any type of content delivery network that provides access to content item 121.
[0062] The content providing platform 120 may include a plurality of channels (e.g., Channel A 125, Channel B 126, etc.). The channels can be a set of content available from a common source, a set of content having a common topic or theme, etc. The data content can be digital content selected by a user, digital content made available by a user, digital content uploaded by a user, digital content selected by a content provider, digital content selected by a broadcaster, etc. For example, Channel A 125 may include two videos (e.g., Content Items 121A - B). The channels can be associated with an owner, and the owner can be a user who can perform actions on the channel. The content can be one or more content items 121. The data content of the channel can be pre - recorded content, live content, etc. Although the channel is described as one implementation of the content providing platform, the disclosed implementations are not limited to a content sharing platform that provides content items 121 via a channel model.
[0063] The product identification system 175, the server machine 170, and the server machine 180 can each include one or more computing devices such as a rack - mount server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an application - specific integrated circuit (ASIC) accelerator (e.g., a Tensor Processing Unit (TPU)). Operations such as those of the prediction server 112, the server machine 170, the server machine 180, the data store 140, etc. can be performed by cloud computing services, cloud data storage services, etc.
[0064] The product identification system 175 may include one or more models 190. The models 190 included in the product identification system 175 may perform tasks related to identifying one or more products from content items. One or more of the models 190 may be trained machine learning models. Operations for generating a trained machine learning model, including training, validating, and testing the model, are described in relation to FIGS. 3A and 5C.
[0065] The model 190 may include one or more text analysis models 191. The text analysis model 191 may be configured to receive text as input and generate, as output, one or more indications of products associated with the text. For example, a first model of the text analysis model 191 may be configured to predict an associated product from the title of a content item, a second model of the text analysis model 191 may be configured to predict an associated product from the description (e.g., written) of the content item, and a third model may be configured to predict an associated product from the caption of the content item (e.g., an automatically generated caption, a machine-generated caption, a user-provided caption, etc.), and so on. In some embodiments, all operations of the text analysis model 191 may be performed by a single model. In some embodiments, the models of the text analysis model 191 may be configured to generate, as output, product context information, for example, information indicating that the content item is associated with one or more products. For example, the product context information may indicate that the content item includes a product category, for example, a group of different products (e.g., a product type, a product brand, or a product category such as "electronic devices").
[0066] Model 190 may include one or more image analysis models 192. The image analysis model 192 may be configured to identify products from one or more images. The image analysis model 192 may include one or more models that are instructed to identify that an image contains a product, a model configured to separate a product image from a content item image (e.g., remove background elements, etc.), a model configured to determine the identity of a product within a content item image, and the like. The operations of the image analysis model 192 may be performed by a single model. The image analysis model 192 may include one or more models configured to provide an image to a product identification model. For example, the image analysis model 192 may include a model configured to extract a portion of a still image of a content item, a model configured to extract one or more frames of a video content item, and the like. The models of the image analysis model 192 may be provided with one or more frames of a video, one or more portions of one or more frames of a video, etc., and may generate as output one or more products and one or more confidence values associated with the one or more products. For example, the models of the image analysis model 192 may receive as input one or more frames of a video content item and may generate as output a list of products having confidence values indicating the likelihood that the products are included in the image of the content item. The image analysis model 192 may include one or more models that determine which images to utilize from a content item. For example, the image analysis model 192 may include one or more models that select frames from a video content item for image identification.
[0067] The image analysis model 192 may include one or more models configured to reduce the dimensions of an image. For example, the image may be reduced to a vector of values. In some embodiments, one or more models of the image analysis model 192 may be configured to reduce the dimensions of the image such that similar images (e.g., images of the same product) may be represented similarly (e.g., by similar vectors) after the dimension reduction. One or more models of the image analysis model 192 may be configured to compare the reduced-dimension image from the content item (e.g., a vector of values generated from one or more frames of the content item video) with the reduced-dimension images of known products (e.g., via the product information platform 161).
[0068] Model 190 may include a text correction model 193. The text correction model 193 may be configured to provide corrections to text associated with a content item. The text correction model 193 may be configured to adjust the text associated with the content item to include one or more products that are referenced in the content item. One or more models of the text correction model 193 may be configured to adjust computer-generated text, machine-generated text, automatically-generated text, etc. associated with the content item, and one or more models of the text correction model 193 may be configured to update the captions (e.g., inaccurate captions) of a video to include one or more products. In some embodiments, machine-generated text (e.g., captions) associated with a content item may be inaccurate. For example, the name of a product may be approximated and replaced in the generation of a caption (e.g., the name of the product may not be a word in the language of the caption, the name of the product may be a word in a language different from the language of the caption, etc.). The models of the text correction model 193 may be configured to identify portions of text that may be inaccurate, recommend corrections, execute corrections, alert a user or other system, etc. For example, the models of the text correction model 193 may receive machine-generated captions of a video, identify portions of the captions that may inaccurately replace words in the language of the caption for the product name, and provide data indicating potentially inaccurate text to other models, systems, users, etc.
[0069] Model 190 may include a fusion model 194. The fusion model 194 may receive, as input, one or more indications of products associated with a content item. In some embodiments, the fusion model 194 receives, as input, the output from one or more other models (e.g., a text analysis model 191, an image analysis model 192, etc.). The fusion model 194 may receive one or more indications of products, and one or more indications of confidence values associated with the content item. For example, the fusion model 194 may receive an indication of a confidence value associated with one or more products detected within the title of one or more products and the content item. The fusion model 194 may further receive an indication of a confidence value associated with one or more products detected within the description of one or more products and the content item. The fusion model 194 may further receive an indication of a confidence value associated with one or more products detected within the caption of one or more products and the content item. The fusion model 19 may further receive an indication of a confidence value associated with one or more products and one or more products detected within the image of the content item. The fusion model 195 may generate, as output, one or more products detected in relation to the content item. The fusion model 195 may further generate a confidence value associated with, for example, the confidence level at which a product appears within the content item, the confidence level associated with the content item, etc. Based on the output of the fusion model 195 (e.g., product identity information and confidence values), further operations may be performed (e.g., UI elements that describe the products associated with the content item may be presented).
[0070] One type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. An artificial neural network generally includes a feature representation component having a classifier or regression layer that maps features to a desired output space. A convolutional neural network (CNN) hosts, for example, multiple layers of convolutional filters. Pooling is performed, non-linearity can be addressed in the lower layers, and generally a multi-layer perceptron is added at the top of the lower layers to map the features of the top layer extracted by the convolutional layers to a decision (e.g., a classification output).
[0071] A recurrent neural network (RNN) is another type of machine learning model. A recurrent neural network model is designed to interpret a series of inputs where the inputs are inherently related to each other, such as time trace data, sequential data, etc. The output of the perceptron in the RNN is fed back as an input to the perceptron to generate the next output.
[0072] Deep learning is a class of machine learning algorithms that uses a cascade of multiple layers of non-linear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) modes. A deep neural network includes a hierarchy of layers, where different layers learn different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a slightly higher level of abstraction and composite representation. In an image recognition application, for example, the raw input can be a matrix of pixels, the first representation layer can abstract the pixels and encode edges, the second layer can compose and encode the arrangement of edges, the third layer can encode higher-level shapes (e.g., teeth, lips, gums, etc.), and the fourth layer can recognize the roll being scanned. In particular, the deep learning process can learn which features to optimally place at which level of itself. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, a deep learning system has a significant credit assignment path (CAP) depth. The CAP is a chain of input-to-output transformations. The CAP describes the potential causal relationship between the input and the output. For a feedforward neural network, the CAP depth can be the depth of the network, which can be the number of hidden layers plus one. For a recurrent neural network where a signal can propagate through a layer more than once, the CAP depth is potentially unbounded.
[0073] In some embodiments, the product identification system 175 further includes server machines 170 and 180. Server machine 170 includes a dataset generator 172 having the ability to generate datasets (e.g., a set of data inputs, a set of target outputs) for training, validating, and / or testing models 190, including one or more machine learning models. Some operations of dataset generator 172 are described in detail below with respect to FIGS. 2 and 5A. In some embodiments, dataset generator 172 can partition historical data (e.g., pre-existing content item data, content items having one or more designed associated products, content items having product assignments provided by one or more users, etc.) into a training set (e.g., 60 percent of the historical data), a validation set (e.g., 20 percent of the historical data), and a test set (e.g., 20 percent of the historical data).
[0074] In some embodiments, components of the product identification system 175 can generate multiple sets of features. For example, features can be rearrangements of input data, combinations of input data, dimensionality reduction of input data, or subsets of input data, etc. One or more datasets can be generated based on one or more features of the input data.
[0075] Server machine 180 includes a training engine 182, a verification engine 184, a selection engine 185, and / or a test engine 186. The engines (e.g., training engine 182, verification engine 184, selection engine 185, and test engine 186) can refer to hardware (e.g., circuits, dedicated logic, programmable logic, microcode, processing devices, etc.), software (instructions running on a processing device, a general-purpose computer system, or a dedicated machine, etc.), firmware, microcode, or combinations thereof. The training engine 182 can have the ability to train one or more models 190 using one or more sets of features associated with the training set from the dataset generator 172. The training engine 182 can generate multiple trained models 190, and each trained model 190 corresponds to a separate set of features of the training set. The dataset generator 172 can receive the output of the trained model (e.g., the fusion model 194 can be trained based on the output of the text analysis model 191 and / or the image analysis model 192), collect the data into a training dataset, a verification dataset, and a test dataset, and use the datasets to train a second model (e.g., the fusion model 194).
[0076] The verification engine 184 may have the ability to verify the trained model 190 using a corresponding set of features of the verification set from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of features of the training set may be verified using a first set of features of the verification set. The verification engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of features of the verification set. The verification engine 184 may discard a trained model 190 that has an accuracy that does not meet a threshold accuracy. In some embodiments, the selection engine 185 may have the ability to select one or more trained models 190 that have an accuracy that meets the threshold accuracy. In some embodiments, the selection engine 185 may have the ability to select the trained model 190 that has the highest accuracy among the trained models 190.
[0077] The test engine 186 may have the ability to test the trained model 190 using a corresponding set of features of the test set from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of features of the training set may be tested using a first set of features of the test set. The test engine 186 may determine the trained model 190 that has the highest accuracy among all of the trained models based on the test set.
[0078] In the case of a machine learning model, model 190 may refer to model artifacts created by training engine 182 using a training set that includes data inputs and corresponding target outputs (the correct answers for each training input). Patterns within the data set that map the data inputs to the target outputs (correct answers) can be found, and machine learning model 190 is provided with mappings that capture those patterns. Machine learning model 190 may use one or more of support vector machines (SVMs), radial basis functions (RBFs), clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, k-nearest neighbor algorithms (k-NN), linear regression, random forests, decision forests, neural networks (e.g., artificial neural networks, regression neural networks), linear models, function-based models (e.g., NG3 models), etc. Synthetic data generator 174 may include one or more machine learning models, and the one or more machine learning models may include one or more of the same type of model (e.g., artificial neural networks).
[0079] Automated (e.g., model-based) detection of products from content items and associated data provides significant technical advantages over other methods. In some embodiments, content items that characterize a product (e.g., a product review video reviewing the product) may become linked to, or associated with, the product without the attention, action, time, etc. of the content creator (e.g., data may be generated that links the product to the content item). In some embodiments, content items that advertise a product may be linked to, or associated with, the product (e.g., the content item may be supported and may promote one or more products). Model-based detection of products within a content item may generate a product association for a product that is not specifically characterized within the content item but that is present within the content item (e.g., products that a purchased user may be interested in may be on screen within a video content item). Model-based detection of products within a content item may generate a product association for a model advertised within the content item. A user may be directed to a product present within a content item based on model-based detection, for example, by providing the user with a UI element indicating that the product is associated with the content item.
[0080] One or more models 190 may be operated on an input to generate one or more outputs. The model may determine (e.g., extract) confidence data from an output that indicates a level of confidence that the output of the model is an accurate description of the content item. For example, the model may determine that a first product is associated with the content item and may determine the confidence that the first product was accurately discovered by a model within the content item. One or more components of the product identification system 175 may use the confidence data to determine whether to update data associated with the content item, e.g., whether to associate one or more products with the content item, whether to update one or more captions of the content item, etc.
[0081] Reliability data may include, or indicate, a level of confidence that the output of a model (e.g., one or more products) is an accurate indication of a product associated with a content item. For example, the level of confidence output by a model (in relation to one product identified within a content item) may be a real number from 0 to 1 (inclusive). A 0 may indicate no confidence that the predicted product is associated with the content item, and a 1 may indicate absolute confidence that the predicted product is associated with the content item. In response to reliability data indicating a level of confidence below a threshold level for a predetermined number of instances (e.g., percentage of instances, frequency of instances, total number of instances, etc.), one or more trained models 190 may be retrained by product identification system 175 (e.g., based on updated data and / or new data for training, validation, testing, etc.). Retraining may include generating one or more data sets (e.g., via data set generator 172).
[0082] For illustrative purposes and not limitation, aspects of the disclosure train one or more machine learning models 190 using historical data to determine an output indicative of a content item - product association, and input current data (e.g., newly updated content items, content items not previously associated with a product, etc.) into one or more trained machine learning models. In other embodiments, a heuristic model, a physics - based model, or a rule - based model is used (e.g., without using a trained machine learning model) to determine that one or more products are associated with a content item. In some embodiments, such models can be trained using historical data. In some embodiments, those models can be retrained using historical data. Any of the information described with respect to data input 210 of FIG. 2 can be monitored or otherwise used in a heuristic model, a physics - based model, or a rule - based model.
[0083] In some embodiments, the functions of client device 110, product identification system 175, content platform system 102, server machines 170, and server machines 180, server machine 106 can be provided by a fewer number of machines. For example, in some embodiments, server machines 170 and 180 can be integrated into a single machine, and in some other embodiments, server machine 170, server machine 180, and server machine 106 can be integrated into a single machine. In some embodiments, client device 110 and server machine 106 can be integrated into a single machine. In some embodiments, the functions of client device 110, server machine 106, server machine 170, server machine 180, and data store 140 can be performed by cloud - based services.
[0084] Generally, the functionality described in one embodiment as being performed by client device 110, server machine 106, server machine 170, and server machine 180 can also be performed on server machine 106 in other embodiments where appropriate. Additionally, the functionality attributed to a particular component can be performed by different components or multiple components that operate together. For example, in some embodiments, product identification system 175 can determine an association between a product and a content item. In other examples, content platform system 102 can determine an association between a content item and one or more products.
[0085] Additionally, the functionality of a particular component can be performed by different components or multiple components that operate together. One or more of server machine 106, server machine 170, or server machine 180 can be accessed as a service provided to other systems or devices through an appropriate application programming interface (API).
[0086] In the disclosed implementations, a "user" can be represented as a single individual. However, other implementations of the disclosure include a "user" that is an entity controlled by a set of users and / or automated sources. For example, a set of individual users associated as a community within a social network can be considered a "user". In other examples, an automated consumer can be an automated ingestion pipeline such as one or more platforms, topic channels such as one or more content items, etc. In addition to the above description, a user can be provided with control that enables the user to make selections both when the systems, programs, or features described herein can and do enable the collection of user information (e.g., information regarding the user's social network, social actions or activities, occupation, user preferences, or the user's current location) and when the user transmits content or communications from the server. Additionally, certain data can be handled in one or more ways before it is stored or used so that personally identifiable information is removed. For example, the user's identity can be trained so that personally identifiable information cannot be determined about the user, or the geographical location of the user from which location information is obtained can be generalized (e.g., to the city, ZIP code, or state level) so that the user's specific location cannot be determined. Thus, a user can have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0087] FIG. 2 is a block diagram of a system 200 that includes a dataset generator 272 for creating a dataset for one or more models, according to some embodiments. The dataset generator 272 may use historical data to create a dataset (e.g., data input 210, target output 220). A dataset generator similar to the dataset generator 272 may be utilized to train an unsupervised machine learning model. For example, the target output 220 may not be generated by the dataset generator 272. A dataset generator similar to the dataset generator 272 may be utilized to train a semi-supervised machine learning model. For example, the target output 220 corresponding to a subset of the data input 210 may be generated by the dataset generator 272.
[0088] The dataset generator 272 may generate a dataset for training, testing, and validating a model. The dataset generator 272 may generate a dataset for a machine learning model. The system 200 may generate a dataset for training, testing, and / or validating a fusion model, e.g., for determining the likelihood that one or more products appear within a content item. A system similar to the system 200 may generate a dataset for training, testing, and / or validating a model having different functions by making corresponding changes to the input data and / or output data included in the dataset. Models that parse text (e.g., extract one or more references to a product from text associated with a content item), analyze images (e.g., extract one or more references to a product from an image associated with a content item), correct text (e.g., include one or more references to a product in text generated by a machine associated with a content item), etc., may have a dataset for training, testing, and / or validating a model generated by a dataset generator similar to the dataset generator 272.
[0089] In some embodiments, a dataset generator, such as dataset generator 272, can be associated with two or more distinct models (e.g., the dataset can be used to train an ensemble model). For example, an input dataset can be provided to a first model, the output of the first model can be provided to a second model, and a target output can be provided to the second model to train, test, and / or validate the first and second models (e.g., the ensemble model).
[0090] Dataset generator 272 can generate one or more datasets for providing to a model, for example, during training, validation, and / or testing operations. A machine learning model can be provided with a set of historical data. A machine learning model (e.g., a fusion model) can be provided with a set of historical text analysis data 264A - 264Z as data inputs. The text analysis data can be provided by a machine learning model and can include, for example, one or more products and confidence values identified by the machine learning model in text associated with a content item. A machine learning model can be provided with a set of historical image analysis data 265A - 265Z as data inputs. The image analysis data can be provided by a trained machine learning model and can include, for example, one or more products and confidence values identified by the machine learning model in one or more images associated with a content item.
[0091] In some embodiments, the dataset generator 272 may be configured to generate datasets for training, textifying, validating, etc. a fusion model. A dataset generator similar to the dataset generator 272 may generate a set of text data (e.g., content item title text, content item description text, content item caption text, etc.) as data input for training a machine learning model to determine one or more products associated with a content item. A dataset generator similar to the dataset generator 272 may generate a set of image data (e.g., one or more frames from a video content item, portions of one or more frames from a video content item, etc.) as data input for training a machine learning model to determine one or more products associated with a content item.
[0092] In some embodiments, the dataset generator 272 generates a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210 (e.g., training input, validation input, test input). The data input 210 may be provided to the training engine 182, the validation engine 184, or the test engine 186 of FIG. 1. The dataset may be used to train, validate, or test a model (e.g., a fusion model, a text analysis model, an image analysis model, etc.). The dataset generator 272 may generate a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210. The data input 210 may be referred to as a "feature", "attribute", "vector", or "information".
[0093] In some embodiments, the dataset generator 272 may generate a first data input corresponding to a first set of historical text analysis data 264A and / or a first set of historical image analysis data 265A for training, validating, or testing a first machine learning model. The dataset generator 272 may generate a second data input corresponding to a second set of historical metrology data 264B and / or a second set of design rule data 265B for training, validating, or testing a second machine learning model. Some embodiments for generating training sets, test sets, validation sets, etc. are further described with respect to FIG. 5A.
[0094] In some embodiments, the dataset generator 272 may generate a target output 220 provided for training, testing, validating, etc. one or more machine learning models. The dataset generator 272 may generate product association data 268 as the target output 220. The product association data 268 may include identifiers of one or more products associated with a content item (e.g., human-labeled product associations). The product association data 268 may include an input-output mapping. For example, a set of historical text analysis data 264A may be associated with a first set of product association data 268, and so on. A machine learning model may be updated (e.g., trained) by providing input data, generating an output, and comparing it to a provided target output (e.g., the "correct answer"). Various weights, biases, etc. of the model are then updated to drive the model towards better alignment with the training data. This process may be repeated multiple times to generate a model that provides accurate outputs for threshold portions of the provided inputs. The target output 220 may share one or more features of the data input 210. For example, the target output 220 may be organized into attributes or vectors, the target output 220 may be organized into sets A-Z, and so on.
[0095] In some embodiments, a dataset generator similar to dataset generator 272 may be utilized in relation to a text analysis model configured to determine one or more product associations with content items. The product associations may include context associations such as the brand of the product, the type of the product, the class of the product, etc. The dataset generator may be associated with a content item and generate, as a target output, a list of products, a type of product, or a brand of product, etc., that is associated with the content item and associated with the text of the content item. A dataset generator similar to dataset generator 272 may be utilized in relation to an image analysis model configured to determine one or more product associations with content items.The dataset generator may generate, as a target output, a list of products, a class of product, a category of product, etc., that is associated with the image input. A dataset generator similar to dataset generator 272 may be utilized in relation to a text correction model. The text correction model may be configured to recognize machine-generated text that inaccurately provides one or more words in the target language instead of a product name. The text correction model may be provided with one or more sets of machine-generated text as input data and, as a target output, a product associated with the text (e.g., a product not accurately captured by the machine generation of the text).
[0096] In some embodiments, following generating a dataset and using the dataset to train, validate, or test a machine learning model, the model may be further trained, validated, tested, or adjusted (e.g., adjusting weights or parameters associated with the input data of the model, such as connection weights within a neural network). The model may be adjusted and / or retrained based on data different from the original training operation, e.g., data generated after training, validation, and / or testing of the model.
[0097] FIG. 3A is a block diagram illustrating a system 300A that generates output data (e.g., product / content item association data) according to some embodiments. System 300A can be used with a fusion model to generate product / content item association data and confidence data based on potential product / content item associations generated by other models (e.g., a text-based model that detects products in text associated with content items, an image-based model that detects products in images associated with content items, etc.). Systems similar to system 300A can be utilized to generate output data from other types of models, such as text analysis models, image analysis models, text correction models, etc.
[0098] In block 310, system 300A (e.g., a component of the product identification system 175 of FIG. 1) performs data partitioning of data that will be used when training, validating, and / or testing a machine learning model (e.g., via the dataset generator 272 of FIG. 2). In some embodiments, the training data 364 includes historical data such as historical associations between text-based products and content items, historical associations between image-based products and content items, etc. In some embodiments, for example, when system 300A is directed to generate an output from a fusion model, the training data 364 may include data generated by one or more trained machine learning models, such as models configured to detect products in text or images associated with content items. The training data 364 can undergo data partitioning in block 310 to generate a training set 302, a validation set 304, and a test set 306. For example, the training set can be 60% of the training data, the validation set can be 20% of the training data, and the test set can be 20% of the training data.
[0099] The generation of the training set 302, the validation set 304, and the test set 306 can be adapted to a specific application. For example, the training set can be 60% of the training data, the validation set can be 20% of the training data, and the test set can be 20% of the training data. The system 300A can generate multiple sets of features for each of the training set, the validation set, and the test set. For example, if the training data 364 includes product associations extracted from text data from more than one text source (e.g., a title associated with a content item and a description associated with the content item), the input training data can be split into a first set of features including products identified from the text from the first source and a second set of features including products identified within the text from the second source. Either the target input, the target output, both can be split into sets, or neither can be split into sets. Multiple models can be trained on different sets of data.
[0100] In block 312, the system 300A uses the training set 302 to perform model training (e.g., via the training engine 182 of FIG. 1). The training of a machine learning model can be achieved in a supervised learning manner, which involves providing a training data set including labeled inputs through the model, observing its output, defining an error (by measuring the difference between the output and the label value), and using techniques such as deep gradient descent and backpropagation to tune the weights of the model so that the error is minimized. In many applications, by repeating this process over many labeled inputs in the training data set, a model can be obtained that can produce accurate outputs when presented with inputs different from those present in the training data set. In some embodiments, the training of a machine learning model can be achieved in an unsupervised manner, for example, where labels or classifications may not be supplied during training. Unsupervised models can be configured to perform anomaly detection, result clustering, etc.
[0101] For each training data item in the training data set, the training data item can be input into a model (e.g., a machine learning model). The model can then process the input training data item (e.g., one or more products detected in relation to a content item, and indications such as associated confidence values) to generate an output. The output can include a list of products that can be associated with the content item and corresponding confidence values. The output can be compared with the label of the training data item (e.g., a human-labeled set of products associated with the content item).
[0102] The processing logic can then compare the generated output (e.g., the predicted product / content item association) with the label that was included in the training data item (e.g., a human-generated list of product / content item associations). The processing logic determines an error (i.e., a classification error) based on the difference between the output and the label(s). The processing logic adjusts one or more weights and / or values of the model based on the error.
[0103] In the case of training a neural network, an error term or delta can be determined for each node within the artificial neural network. Based on this error, the artificial neural network adjusts one or more of its parameters (weights for one or more inputs to the node) for one or more of its nodes. The parameters can be updated in a backpropagation manner such that the nodes in the topmost layer are updated first, followed by the nodes in the next layer, and so on. The artificial neural network includes multiple layers of "neurons", and each layer receives input values from the neurons in the previous layer. The parameters for each neuron include weights associated with the values received from each of the neurons in the previous layer. Thus, adjusting the parameters can include adjusting the weights assigned to each of the inputs to one or more neurons in one or more layers within the artificial neural network.
[0104] System 300A can train multiple models using multiple sets of features of the training set 302 (e.g., the first set of features of the training set 302, the second set of features of the training set 302, etc.). For example, System 300A can use the first set of features within the training set (e.g., a subset of the training data 364 such as data associated only with a subset of models configured to generate product / content item associations, etc.) to generate a first trained model and train the model to generate a second trained model using the second set of features within the training set. In some embodiments, the first trained model and the second trained model can be combined to generate a third trained model (e.g., which may be better than the first or second trained model for itself). In some embodiments, the sets of features used when comparing models can overlap (e.g., the first set of features is a product based on the content item title, description, and some images, and the second set of features is a product detected based on a different set of images of the content item, the description of the content item, and the detected context of the content item (e.g., the type of product associated with the content item)). In some embodiments, hundreds of models can be generated, including models with various substitutions of features and combinations of models.
[0105] In block 314, system 300A performs model verification using the validation set 304 (e.g., via the validation engine 184 of FIG. 1). System 300A may verify each of the trained models using a corresponding set of features of the validation set 304. For example, system 300A may verify the first trained model using the first set of features in the validation set and may verify the second trained model using the second set of features in the validation set. In some embodiments, system 300A may verify hundreds of models (e.g., models having various substitutions of features, combinations of models, etc.) generated in block 312. In block 314, system 300A may determine the accuracy for each of one or more trained models (e.g., via model verification) and may determine whether one or more of the trained models have an accuracy that meets a threshold accuracy. In response to determining that none of the trained models have an accuracy that meets the threshold accuracy, the flow returns to block 312, where system 300A performs model training using a different set of features of the training set, or an updated or expanded training set provided by the dataset generator, etc. In response to determining that one or more of the trained models have an accuracy that meets the threshold accuracy, the flow proceeds to block 316. System 300A may discard the trained models that have an accuracy less than the threshold accuracy (e.g., based on the validation set).
[0106] In block 316, system 300A performs model selection (e.g., via the selection engine 185 of FIG. 1) to determine which of one or more trained models (e.g., the selected model 308 based on the verification in block 314) that meet the threshold accuracy have the highest accuracy. In response to determining that two or more of the trained models that meet the threshold accuracy have the same accuracy, the flow may return to block 312, where system 300A performs model training using a more refined training set corresponding to a more refined set of features for determining the trained model with the highest accuracy.
[0107] In block 318, system 300A performs a model test using test set 306 to test the selected model 308 (e.g., via test engine 186 of FIG. 1). System 300A may test the first trained model to determine that the first trained model meets a threshold accuracy using a first set of features within the test set (e.g., based on the first set of features of test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 is overfit to the training set 302 and / or validation set 304 and not applicable to other data sets such as test set 306), the flow proceeds to block 312, where system 300A performs model training (e.g., retraining) using a different training set corresponding to a different set of features, or different content items, etc. In response to determining that the selected model 308 has an accuracy that meets the threshold accuracy based on test set 306, the flow proceeds to block 320. At least in block 312, the model may learn patterns in the training data to make predictions, and in block 318, system 300A may apply the model to the remaining data (e.g., test set 306) to test the predictions.
[0108] In block 320, system 300A uses a trained model (e.g., selected model 308) to receive current data 322 (e.g., newly uploaded content items, newly created content items, content items not included in the training, test, or validation sets of the selected model 308, etc.), and determines (e.g., extracts) output data 324 (e.g., product / content item associations and corresponding confidence values) from the output of the trained model. Correction actions associated with the content item and / or the associated data may be performed in consideration of the output data 324. For example, the instructions may be updated to include presenting UI elements with content items that specify that the content item includes one or more associated products, and the instructions may be updated to include presenting UI elements with content items that include additional information about the associated products, etc. The instructions may depend on additional factors, such as user preferences for the content item (e.g., search page, home page, etc.) or the presentation environment. In some embodiments, the current data 322 may correspond to the same type of features in the historical data used to train the machine learning model. In some embodiments, the current data 322 corresponds to a subset of the types of features in the historical data used to train the selected model 308 (e.g., the machine learning model may be trained using product associations and / or context information and confidence values from several sources such as text-based sources and image-based sources, and a subset of this data may be provided as the current data 322).
[0109] In some embodiments, the performance of a machine learning model (e.g., the selected model 308) can be adjusted, improved, and / or updated over time. For example, additional training data can be provided to the model to improve its ability to accurately classify product associations with content items. In some embodiments, a portion of the current data 322 can be provided to retrain the model (e.g., via the training engine 182 of FIG. 1). A portion of the current data 322 can be labeled (e.g., labeled by a human), and the labels can be provided to retrain the model (e.g., as the current target output data 346). The current data 322 and the current target output data 346 can be utilized to update and / or improve the selected model 308 periodically, continuously, etc.
[0110] In some embodiments, one or more of acts 310-320 can be performed in various orders and / or in conjunction with other acts not presented and described herein. In some embodiments, one or more of acts 310-320 may not be performed. For example, in some embodiments, one or more of the data partitioning of block 310, the model verification of block 314, the model selection of block 316, or the model testing of block 318 may not be performed.
[0111] System 300A has been described with respect to the fusion model. The fusion model receives one or more indications of a product detected in relation to a content item (e.g., from text associated with the content item, from an image associated with the content item, etc.) and a confidence value (e.g., the degree of confidence that the product is indeed referenced within the content item), and based on the multiple inputs, generates as output the overall likelihood that the product is referenced within the content item. A system similar to System 300A can be utilized to perform other machine learning-based tasks, such as text or image analysis models, where its output is provided as input to the fusion model, and can operate in a manner similar to that described in relation to System 300A with appropriate substitutions for input data, output data, etc.
[0112] FIG. 3B is a block diagram of an example system 300B for generating an association between a content item and one or more products, according to some embodiments. System 300B may include a plurality of modules, such as, for example, image identification 330, image inspection 340, text identification 350, fusion 360, etc. In some embodiments, the plurality of modules of System 300B may operate together to identify products from a content item. For example, products detected in an image of the content item and metadata of the content item (e.g., title, description, caption, etc.) may be provided to a fusion model to determine the likelihood that one or more products are associated with the content item, based on the multiple input channels. In response to the likelihood that a product appears within the content item as determined by the fusion model, an action (e.g., updating the metadata of the content item to include an indication of the product) may be taken.
[0113] The image recognition module 330 can be utilized to identify one or more products from an image associated with a model. For example, an image from a content item can be compared with images in a database of products (e.g., thousands of products) to identify the products present in the video. The products that exist can include a specific subject of the content item (e.g., the products reviewed within the content item), or products included in the content item (e.g., products that appear incidentally, products that appear without being specifically highlighted, etc.). The text recognition module 350 can identify one or more products associated with a content item from metadata / text data associated with the content item, such as text including the title of the content item, the description of the content item, the caption associated with the content item, etc. The image inspection module 340 can use one or more images to inspect the products identified within a content item. For example, the image inspection module 340 can operate similarly to the image recognition module 330 but can be used to confirm the presence of one or more products identified by a separate module (e.g., by comparing an image of a potential product with a more restricted range of product images provided by another module). The fusion module 360 can receive candidate products and associated confidence values included in a content item and can determine the likelihood that one or more products appear within the content item based on various inputs.
[0114] Image recognition 330 can be used to determine content items having visual components, such as products associated with videos. Image recognition 330 may include frame selection 332. Frame selection 332 can be used to select one or more frames of a video that search for an image of a product. Frame selection 332 can be performed via random sampling, regular sampling, intelligent sampling methods, etc. For example, a content item (e.g., a video) can be provided to a machine learning model, and the machine learning model can be trained to predict frames of a video that are likely to include one or more products.
[0115] One or more frames can be provided to object detection model 334. Object detection 334 can extract predicted objects from one or more frames. For example, object detection 334 can separate potential products from people, animals, backgrounds, etc. in the image data of the content item. Object detection 334 can be a machine learning model or can include a machine learning model.
[0116] Images of the detected objects can be supplied to embedding 336. Embedding 336 can include converting one or more images to a lower dimension. Embedding 336 can include providing one or more images to a dimensionality reduction model. The dimensionality reduction model can be a machine learning model. The dimensionality reduction model can be configured to reduce the dimensions of similar images in a similar manner. For example, embedding 336 can receive an image as input and generate a vector of values as output. Embedding 336 can be configured, trained, etc. such that similar images (e.g., images of the same or similar products) are represented similarly within the reduced dimensional vector space (e.g., by Euclidean distance, by cosine distance, by other distance metrics, etc.). Embedding 336 can generate dimensionally reduced data.
[0117] The reduced-dimensional image data can be provided to product identification 338. Product identification 338 can identify one or more products associated with the reduced-dimensional representation provided by embedding 336. Product identification 338 can compare the reduced-dimensional image data (e.g., provided by embedding 336) with the reduced-dimensional image data of products included in product image index 339 (e.g., generated from product images by the same machine learning model used by embedding 336). Product image index 339 can be stored as part of a data store. Product image index 339 can include, for example, many products (e.g., hundreds of products, thousands of products, or more). Product image index 339 can include associations between the stored image data (e.g., reduced-dimensional image data) and product identifiers, product indicators, etc. Product image index 339 can be segmented. For example, the stored dimensionally reduced data can be classified into one or more categories, classes, etc. For example, product identification 338 can compare the data received from embedding 336 with products of a specific category, type, classification, etc. In some embodiments, the category, type, classification, etc. can be provided by one or more users, or one or more content creators can be automatically detected (e.g., by one or more machine learning models), etc. A content item or one or more products associated with the content item can be associated with a category (e.g., a general category such as electronic devices, a more restricted category such as screen devices, a classification of a product such as a tablet, a manufacturer or brand, or a model, etc.). Product identification 338 can generate one or more indications of one or more products detected in the image of the content item (e.g., a list of products that may match the products represented in product image index 339) and one or more indications of a confidence value (e.g., the confidence level that each of the list of products was accurately detected).The output of the image identification 330 can be utilized to update the metadata of the content item (e.g., to include one or more product identifiers or indicators so as to include an association with one or more products). The output of the image identification 330 can be provided to the image inspection 340, for example, to inspect the presence within the content item of the image identified by the image identification 330. The output of the image identification 330 can be provided to the fusion 360, for example, to generate an overall and / or multi-input determination of the products included in the content item via the fusion model 366. The output of the image identification 330 can be provided to the text identification 350 (not shown), for example, to limit the space of the products such that they are queried, searched, compared, etc. by the text identification module 350. In some embodiments, the image identification 330 is utilized to identify products found within one or more frames of a video content item. For example, the image identification module 330 can be configured to generate a list of all products detected within any selected frame and provide a confidence value for each product within the selected frame. The image identification module 330 can generate image-based product data, e.g., one or more identifiers of the product, where the product is identified based on an image of the content item.
[0118] The image inspection module 340 can be configured to inspect for the presence of an identified product of a content item using one or more images of the content item. For example, the image inspection module 340 can include a model configured to confirm the presence of a product identified by other models. The image inspection module 340 can include a secondary identification 345. The secondary identification 345 can include components similar to the image identification 330. In some embodiments, the image identification module 330 can communicate directly with the product candidate image index 344 instead of or in addition to the secondary identification 345 that communicates with the product candidate image index 344. In some embodiments, the secondary identification 345 can perform a role similar to the image identification 330, but can include, for example, a different model, a model trained using different training data, a model configured to select different frames, or a model configured to detect different objects.
[0119] The image inspection 340 may include a synthesis model 341. The synthesis model 341 may receive indications of products identified by, for example, an image identification module 330, a text identification module 350, secondary identification 345 (data flow not shown), etc. The synthesis model 341 may select an object detection model 342 that provides data (for example, an object detection model specially configured for product category or classification). The synthesis model 341 may include image selection. For example, the synthesis model 341 may provide one or more images to the object detection 342 and may select one or more frames to provide to the object detection 334, etc. For example, the synthesis model 341 may provide one or more frames highly likely to contain a product to the object detection 342 based on the data received from the image identification 330 and the text identification 350. The object detection model 342 may perform functions similar to those of the object detection model 334 modified by the functions of the synthesis model 341, for example. The embedding 343 may perform functions similar to those of the embedding 336, for example, to reduce the dimensions of the detected product image. In some embodiments, the product candidate image index 344 may include reduced-dimension image data (for example, a vector of values) detected by other modules (for example, the image identification 330, the text identification 350, etc.). The secondary identification 345 may compare the reduced-dimension image data (for example, the embedded image data) with the candidate data of the product candidate image index 344 (for example, products identified by modules other than the image inspection module 340) to inspect for the presence of products within the content item.
[0120] Text identifier 350 may be configured to identify one or more products from text data (e.g., metadata) associated with a content item. Text identifier 350 may generate metadata-based product data, e.g., one or more product identifiers, based on the metadata of the content item. Text identifier 350 may generate text-based product data, e.g., one or more product identifiers, based on the text data associated with the content item. Text identifier 350 may identify a product from one or more of the content item title, content item description, caption associated with the content item (e.g., a machine-generated caption), comment associated with the content item, and / or other text data or metadata of the content item. The text data (e.g., metadata) associated with the content item may be provided to text analysis model 352. Text analysis model 352 may be a machine learning model. Text analysis model 352 may be configured to detect or predict a product from the text data associated with the content item. Text analysis model 352 may be configured to detect one or more products having product identifiers stored in product identifier 354. Text analysis model 352 may provide an output (e.g., a list of detected candidate products, associated confidence values, etc.) to image inspection module 340. Text analysis model 352 may provide an output to synthesis model 341. Text analysis model 352 may provide an output that affects the products of the image inspection. For example, due to the output of text analysis model 352, the products detected by text identification module 350 may be added to product candidate image index 344. Image inspection module 340 may query an index (e.g., product candidate image index 344) that includes products detected by other modules, such as image identification module 330, text identification module 350, etc.
[0121] The fusion module 360 may receive output data (e.g., detected products, associated confidence values) from one or more sources (e.g., the image identification module 330, the image inspection module 340, the text identification module 350, etc.). The fusion module 360 may further receive data from other sources, such as context term extraction 362, or additional feature extraction 363. The context term extraction 362 may provide context to potential products of the content item, for example, may detect categories or topics associated with some products. The context term extraction 362 may be performed by one or more machine learning models. The context term extraction 362 may detect context information from text associated with the content item, metadata associated with the content item, etc. The additional feature extraction 363 may provide additional details that may be used to determine whether one or more products appear within the content item. The additional features may include video embeddings. The additional features may include other metadata of the content item, such as the date the content item was uploaded to the content providing platform (e.g., compared to the release date of the product), the classification of the content item (e.g., shopping or product review videos may be more likely to contain products than other types of videos), etc.
[0122] Data from multiple sources may be provided to the fusion model 366. The fusion model 366 may be configured to receive data including one or more products with confidence values, for example, and determine one or more products with confidence values indicating the likelihood that the products appear within the content item. In some embodiments, the content item may be a video. In some embodiments, the content item may be a live streaming feed, such as a live streaming video feed (e.g., a product review stream, an unboxing stream, etc.). The content item may be a short form video.
[0123] Figures 4A - E represent example UIs presented on a user device that include UI elements associated with related products according to several embodiments. Figures 4A - E may include UIs provided as part of an application of devices 400A - E, for example, a web browser application, a mobile application, etc., associated with / provided by a content platform. User interactions with various elements of the UIs of Figures 4A - E may result in changes to the presented UI elements. For example, an interaction with a UI element indicating that a content item has an associated product may cause a second UI element to be displayed that presents additional information (e.g., regarding the associated product) (e.g., replacing the first UI element, expanding the first UI element, etc.). UI elements may include elements that, when interacted with, cause less information regarding the associated product to be displayed in the UI (e.g., collapsing a panel that describes one or more associated products). Interacting with a UI element associated with one or more products may result in different effects, such as causing a transition to a UI environment for presenting a content item. The new UI environment may include one or more UI elements associated with one or more products of the content item. Interacting with a UI element associated with a product may cause the display of UI elements that facilitate a transaction (e.g., a purchase) of the product. Various interactions are possible between the UIs and UI elements presented in Figures 4A - E (e.g., interacting with an element of a first UI layout may cause a transition to a second UI layout), and any transition between sample UIs, similar UIs, inclusion of similar UI elements, etc. is within the scope of the present disclosure. Figures 4A - E are described in relation to video content items, and other types of content items (e.g., image content, text content, audio content, etc.) may be presented within similar UIs. Any optional features, elements, etc. presented in relation to one or more of Figures 4A - E may be included in systems similar to those figures as necessary.
[0124] FIG. 4A represents a device 400A presenting a UI 402 of an example including a UI element 404 showing one or more associated products, according to some embodiments. A UI including elements of UI 402 and / or similar UI elements may be presented by device 400A as part of an operation to present one or more content items for user selection. UI element 404 is represented in FIG. 4A in a folded state, e.g., a folded default state.
[0125] UI 402 includes a first content item selector 406 (e.g., a video thumbnail) and a second content item selector 408. In some embodiments, more or fewer content items may be selectable, and UI 402 may be scrolled to view additional content items and the like. UI element 404 is associated with the content item indicated by content item selector 406. UI element 404 (and other UI elements of FIGS. 4A - D related to the product) may be presented above, adjacent to, overlapping, within, below, etc. the content item selector for the content item or the content item associated with UI element 404. In some embodiments, product information (e.g., product / content item association) may be provided by the content creator. Product information may be provided by one or more users. Product information may be provided by an administrator. Product information may be retrieved from content items, e.g., via one or more machine learning models or via a system such as system 300B.
[0126] In some embodiments, the user may interact with UI element 404, which will be presented along with replaced UI elements, updated UI elements, etc. For example, the user may interact with expansion element 410 to display more information about a product associated with a content item. In some embodiments, expanding UI element 404 may open a panel that includes additional information about one or more products associated with the content item. Expanding UI element 404, interacting with UI element 404, etc. may adjust the presentation of UI 402, for example, to include elements represented in FIGS. 4B - D.
[0127] UI element 404 may include, for example, expansion element 410, an indication of associated products (e.g., how many products are associated with the content item), or a visual indication of product 412 (e.g., a visual indication that a transaction or purchase is available, a link to the shipper for the product is available, etc.). User interaction with one or more components of UI element 404 may cause device 400A to modify the presentation of UI element 404. For example, user selection of expansion element 410 may cause the presentation of UI element 404 to be modified to an expanded state.
[0128] The UI including UI element 404 can be presented in response to the device 400A sending a request for a content item to a content delivery platform (e.g., the content delivery platform 120 in FIG. 1). The UI element 404 can be presented within a home feed (e.g., a list of suggested content items for a user or user account). The UI element 404 can be presented within a suggested feed (e.g., a list of suggested content items based on one or more most recently presented content items). The UI element 404 can be presented as part of a play list (e.g., a list of content items to be presented, such as populated by a user or a creator). The UI element 404 can be presented within a search feed (e.g., in response to a search query generated by a user and sent by the device 400A to the content delivery platform). For example, one or more UI elements associated with a product can be presented based on product inclusion, product category, product brand, etc. included in the search query. The UI element 404 can be presented within a product focus feed (e.g., a shopping content feed). The UI element 404 can respond to a user selection of a content item, for example, can be displayed when there is a user selection to view an associated video, can be presented together with the associated content item, etc. The UI element 404 can be presented based on a detected interest of the user within a content item. For example, the user can dwell on a thumbnail for a video (e.g., the user can place a cursor over the thumbnail, the user can pause scrolling with the presented thumbnail, etc.). When dwelling meets one or more conditions (e.g., a duration condition, a thumbnail position condition, etc.), the UI element 404 can be presented. The UI element 404 can be presented in response to additional data, such as user account history, user settings, user preferences, etc. The UI element 404 can be presented together with a list of content items to be presented.The UI element 404 can be presented while presenting a content item (for example, while playing a video associated with the product of the UI element 404), and so on.
[0129] FIG. 4B depicts a device 400B presenting a UI 420 of an example including a UI element 422 that shows associated products, according to some embodiments. The UI element 422 includes information about one or more products associated with a content item. The UI element 422 is presented in FIG. 4B in an expanded state. The UI element 422 may include several components. For example, the first component may include information about the first product (e.g., picture, product name, price, timestamp, etc.), and the second component may include information about the second product, and so on. In some embodiments, the UI element 422 may be scrollable, for example, to access information about additional products. The UI element 422 may include multiple tabs (e.g., the UI element 422 may be associated with product information and one or more other types of information). For example, the UI element 422 may include a product tab 424 and a chapter tab 426. In some embodiments, the product tab 424 may be opened by default (e.g., the content of the product tab 424 may be presented by default). In some embodiments, the chapter tab 426 may be opened by default (e.g., the content of the chapter tab 426 may be presented by default). In some embodiments, other tabs may be opened by default. For example, the chapter tab 426 may be opened by default excluding the content item of the associated product, and the tab opened by default may be selected based on the user history (e.g., interacting with elements such as the chapter tab, product tab, etc.) or based on a search query (e.g., the search may include a product name, related terms, or a phrase such as "product review", etc.). The UI element 422 may include additional elements, for example, elements that the user may utilize to control the presentation of the UI 420, such as the UI element 422.For example, the UI element 422 may include a "closed" element for presenting the UI 420 without the UI element 422, a "collapse" element for displaying less information (e.g., folding the panel as similar to the UI element 404 in FIG. 4A and modifying the presentation of the UI element 422 to a collapsed state), one or more elements for displaying more information (e.g., one or more listed products, listed product icons, etc. may be selected to present additional information about the product or to promote the purchase of the product), and so on.
[0130] The product tab 424 of the UI element 422 may include one or more pictures of the product, information about the product (e.g., the name of the product, the description of the product, etc.), the price of one or more products, the timestamp of the content item related to the product, and so on. In some embodiments, the product information may be provided by a content creator, one or more users, or a system administrator, etc. In some embodiments, the product information may be retrieved by one or more models, such as a machine learning model. For example, the presence of a product in a content item, the association of the product in the content item, or the timestamp or location at which the product appears in the content item may be determined by one or more machine learning models. A system such as the system 300B in FIG. 3B may be utilized to determine one or more products associated with a content item (e.g., a video). The part of the content item related to the content item (e.g., the timestamp) may be determined via, for example, the timing of the caption associated with the product, or the timing of the display of the image or frame of the video including the product. In some embodiments, selecting a product may result in the presentation of the part of the content item associated with the product, the presentation of the content item starting at the time indicated by the timestamp associated with the product, and so on.
[0131] In some embodiments, portions of the UI element 422 associated with a particular product (e.g., visual components) may be displayed by default, may be displayed differently (e.g., highlighted), etc. For example, in response to receiving a search query from a user that includes the name of a product, the UI element 422 including the display associated with the searched product may be displayed.
[0132] In some embodiments, one or more associations between content items and products may be stored, for example, as metadata associated with the content items (the metadata associated with the content items may further include a content item title, description, presentation history, caption associated with the content item, etc.). In response to the device (e.g., device 400B) executing instructions such as displaying a list of content items for presentation, presenting a content item to the user, or presenting a UI element (e.g., UI element 422) including information about one or more products, the device may retrieve information about the products based on the metadata associating the products with the content items. Information about the products (e.g., associated products such as images, color variations, availability, price, etc.) may be retrieved from a data store. The data store may include information about the products, may be updated, for example, as information such as the price of a product change, and the UI may retrieve the updated information based on the content item / product association and may display the updated information.
[0133] The UI element 422 can be presented as part of, for example, a home feed (e.g., a list of suggested content items for a user or user account), a suggested feed (e.g., a list of suggested content items based on one or more recently presented content items), a playback list, a list of search results, a shopping content page, etc. The UI element 422 can be presented when there is a selection of a content item for presentation, when there is a presentation of a content item, etc. The UI element 422 can be presented during ideation (e.g., pausing a scroll on a content thumbnail, pausing a scroll on a less detailed element such as the UI element 404 in FIG. 4A). In some embodiments, the UI element 422 can be removed or replaced upon a user action, e.g., upon a scroll, and the UI element 422 can be collapsed in a manner similar to the UI element 404 (e.g., facilitating selection of a content item from a list of content items, simplifying scrolling through a list of content items). The UI element 422 can be displayed while presenting a list of content items, while presenting a single content item (e.g., while playing a video associated with the product of the UI element 422), etc.
[0134] Figure 4C represents a device 400C presenting an example UI 430 that includes a UI element 432 presenting a content item and a UI element 434 presenting information about an associated product, according to some embodiments. The UI element 434 may include more detailed information about the product associated with the content item than the UI element 422 of Figure 4B. The UI element 434 may be the product in focus and may function, for example, to display product information to the user. The UI element 434 may include one or more components, such as a component associated with a first product, a component associated with a second product, and the like. In some embodiments, the UI element 434 may include a list of products associated with the content item. The UI element 434 may be navigable, scrollable, and the like. The UI element 434 may include one or more control elements, such as a back button to return to the previous view, a close button to close the UI element 434 and view a different set of UI elements (e.g., not related to the product), and the like. In some embodiments, the UI element 434 may be removed from the UI 430 in response to other user actions, such as the user scrolling past the associated content item. In some embodiments, a user selection of a product presented via the UI 434 may drive the presentation of UI elements that promote the purchase of the product. In some embodiments, the UI element 434 may be displayed in response to a determination of the user's interest in one or more of the products associated with the content item (e.g., based on the user history, based on one or more terms in the user search query, based on the user browsing and / or selecting to be presented with shopping content items, based on a user selection of a product or a UI element associated with the product, and the like).
[0135] The UI element 434 may include one or more pictures and / or additional information about one or more products associated with a content item (e.g., a content item presented via the UI element 432). The pictures and / or information may be provided by a content creator, one or more users, retrieved from a database (e.g., based on product / content item association metadata), and so on. In some embodiments, the UI element 432 may scroll automatically. For example, the UI element 432 may scroll as the content item is presented such that a product associated with a portion of the currently presented content item is visible. The UI element 432 may be presented in response to a user selection of a content item to be presented. The UI element 432 may be presented in response to other factors, such as a user history. The UI element 432 may be presented when a user contemplates a related content item, related UI element, etc.
[0136] In some embodiments, the UI element 434 may present information about a single product, e.g., a product selected by a user (e.g., via the UI element 422 of FIG. 4B). The UI element 434 may be the product in focus, may have a single product in focus, may display product variations (e.g., as described in relation to the UI element 444 of FIG. 4D), or may include other elements, components, and / or information described in relation to other UI elements described herein.
[0137] FIG. 4D depicts a device 400D presenting an example UI 440 that includes a UI element 442 presenting a content item and a UI element 444 facilitating a transaction associated with a product. The UI element 444 may provide one or more fields associated with a user performing a transaction, e.g., purchasing a product. For example, the UI element 444 may include an alternative product panel 446 that may include information, pictures, prices, etc., regarding alternative products (e.g., related products associated with one or more products associated with the content item such as color variations, size variations, variations of a bundle of products, similar products of other brands, etc.).
[0138] The UI element 444 may include a transaction element 448. In some embodiments, the transaction element 448 may facilitate a transaction (e.g., a purchase) within the application providing the UI 440. In some embodiments, the transaction element 448 may facilitate a transaction via another application or another website, etc., e.g., interacting with the transaction element 448 may direct the user to a shipper website, may direct the device 400D to open an application associated with purchasing a product, etc.
[0139] The UI element 444 can be navigable, scrollable, etc. The UI element 444 can include one or more control elements, such as a back button to return to the previous view, a close button to close the UI element 444 and display a different set of UI elements via the UI440, etc. In some embodiments, the UI element 444 can be displayed as part of a list of content items for user selection, as part of the UI440 that presents content items, etc. The UI element 444 can be responsive to a determination of user interest within a transaction (e.g., a purchase) associated with one or more products included in a content item, such as a selection of a product from a UI element such as the UI element 434 in FIG. 4C, a selection of a content item, a consideration for a content item, a product or product-related term included in a search query, navigation by the user to a shopping focus list of content items, etc. The UI element 444 can receive information regarding one or more products from a database, for example, based on metadata associating the content item with the one or more products.
[0140] The UI elements represented in FIGS. 4A - D can be integrated in various configurations. For example, a UI element such as the UI element 404 can be presented, and upon a user interaction with the UI element 404, an element such as the UI element 422 can be presented, and upon a user interaction with the UI element 422, the UI element 434 can be presented, and upon a user interaction with the UI element 434, the UI element 444 can be presented, etc. One or more UI elements can include navigation elements for instructing the device to display different UI elements. For example, interacting with the expansion element 410 can result in the presentation of a UI element such as the UI element 422, and interacting with different elements of the UI element 404 can result in the presentation of a UI element such as the UI element 434 or the UI element 444.
[0141] Other connections between UI elements are possible. For example, interacting with a UI element such as UI element 404 can cause the display of UI elements such as UI element 422, UI element 434, UI element 444, and so on. Interaction with a UI element such as UI element 422 or a portion thereof can cause the presentation by UI elements such as UI element 404, UI element 434, UI element 444, and so on. User interaction with a UI element such as UI element 434 or a portion thereof can cause the presentation of UI elements such as UI element 404, UI element 422, UI element 444, and so on. User interaction with a UI element such as UI element 444 or a portion thereof can cause the display of UI elements such as UI element 404, UI element 434, UI element 444, and so on. The default UI elements presented can depend on the environment in which the UI elements are presented (e.g., a list of content items presented in response to a search, home feed, shopping feed, viewing feed, etc., or an environment that includes the presented content items). For example, the selection of the form of the UI element can be based on several factors. In some embodiments, including a product name, category, or the like in a search query can change the default UI elements, for example, the UI can be made to demonstrate by default UI elements that include information about the product or purchase options for the product. The determination of the form of the UI elements to be displayed can be based on user history, user account history, user actions (e.g., opening the home feed, presenting the viewing feed, sending a search query, selecting shopping, etc.). The transition between the forms of the UI elements associated with the product can be determined by additional data similar to the data used to determine the form of the presented UI elements.
[0142] Figure 4E depicts device 400E of an example having UI elements overlaid on content presentation element 452, according to some embodiments. Device 400E includes UI 450. UI 450 may be provided by an application, for example, an application associated with a content delivery platform. Presentation element 452 may present a content item (e.g., a video). UI 450 may present a list of additional content items, additional information associated with the presented content item (e.g., title, description, comments, live chat, etc.), additional UI elements associated with the product (e.g., UI elements such as UI elements 404, 422, 434, 444, or variations thereof), and the like.
[0143] The hint element 452 can be overlaid on one or more UI elements. The UI element 454 can indicate a product included in a content item (e.g., shown within a video). The UI element 454 can perform functions similar to other UI elements in FIGS. 4A - D, for example, present information about the product, enable display of further information about the product, facilitate purchasing the product, and so on. The placement of the overlaid elements can be determined by one or more users, content creators, models, etc. (e.g., a machine learning model configured to detect products or objects within an image can be utilized to avoid areas of the content item display that contain the object). The UI element 454 can include a visual indicator of where one or more related products are located within the content item (e.g., where within the video). The UI element 454 can be presented during the presentation of the content item, for example, while the video is playing. The UI element 454 can be displayed in response to the presence of related products within the content item and / or removed, for example, it can be displayed while the content item is within the video. The UI element 454 can indicate multiple products, identify products (e.g., name one or more products), display information about the products, and so on. Multiple UI elements such as the UI element 454 can be displayed, for example, on the thumbnail of the video, throughout the presentation of the video, simultaneously during the video, and so on.
[0144] Overlaid UI elements such as the UI element 454 can be presented in combination with other UI elements associated with the product. For example, the UI element 456 can open a panel containing information about multiple products associated with the content item, and the UI element 454 can cause the display of a UI element containing information about the pictured product, and so on.
[0145] The overlaid UI element 454 can be displayed on the visual representation of the content item (e.g., in front of it, having visual precedence over it, etc.). For example, the UI element 454 can be overlaid on a video thumbnail. The UI element 454 can be displayed on the content item. For example, the UI element 454 can be overlaid on a playing video. The presentation of the UI element 454 can be executed in response to a user action. For example, when the user determines that they are interested in one or more products (e.g., via a search query, via interaction with product-related UI elements, via the user history, etc.), the UI element 454 can be displayed and overlaid on other UI elements. In some embodiments, the content item can be a live-streamed video. In some embodiments, the content item can be a short-form video.
[0146] In some embodiments, the UI element 454 can execute similarly to the execution of the UI element 404. For example, it can notify the user that one or more products are associated with the content item. The UI element 454 can respond to user interaction similarly to the UI element 404. For example, it can open a panel containing product information, or expand, or modify the presentation of the UI element 404 to display more or different information, or expand the UI element to include more information, or initiate the presentation of the content item, etc. The UI element 454 can respond to user interaction similarly to the UI element 434. For example, it can open or expand a panel that facilitates a transaction.
[0147] Figures 5A - F are flowcharts of methods 500A - F related to content items having associated products, according to some embodiments. Methods 500A - F may be executed by processing logic that may include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, processing devices, etc.), software (e.g., instructions running on a processing device, general - purpose computer system, or dedicated machine), firmware, microcode, or combinations thereof. In some embodiments, methods 500A - F may be partially executed by the content platform system 102, product identification system 175, and / or client device 110 of FIG. 1. Method 500A may be partially executed by a product identification system 175 (e.g., server machine 170 and dataset generator 172 of FIG. 1, dataset generator 272 of FIG. 2). The product identification system 175 may use method 500A to generate a dataset for at least one of training, validating, or testing a machine - learning model, according to the disclosed embodiments. Methods 500B - D may be executed by a product identification system 175 (e.g., system 300B of FIG. 3B) and / or server machine 180 (e.g., training, validation, and testing operations may be executed by server machine 180). Method 500E may be executed by client device 110. Method 500E may be utilized by client device 110, for example, to display one or more UI elements associated with a product that facilitate user identification of the product included in a content item. Method 500F may be executed by content platform system 102 and may be executed, for example, by the processing logic of content - providing platform 120 to facilitate the presentation of one or more UI elements associated with a product by client device 110. In some embodiments, a non - transitory machine - readable storage medium stores instructions that, when executed by a processing device (e.g., of product identification system 175, server machine 180, etc.), cause the processing device to execute one or more of methods 500A - F.
[0148] For simplicity of explanation, methods 500A - F are presented and described as a series of operations. However, the operations according to the present disclosure can be performed in various orders and / or concurrently with other operations not presented and described herein. Further, not all of the illustrated operations need to be performed so as to implement methods 500A - F by the disclosed subject matter. Additionally, those skilled in the art will understand and recognize that methods 500A - F can alternatively be represented as a series of interrelated states via a state diagram or events.
[0149] FIG. 5A is a flowchart of a method 500A for generating a dataset for a machine learning model, according to some embodiments. Referring to FIG. 5A, in some embodiments, at block 401, the processing logic implementing method 500A initializes the training set T to an empty set.
[0150] At block 502, the processing logic generates a first data input (e.g., a first training input, a first validation input) that can include one or more of product data, image data, metadata, text data, reliability data, etc. In some embodiments, the first data input can include a first set of features for the type of data, and the second data input can include a second set of features for the type of data (e.g., as described with respect to FIG. 3A). The input data can include historical data in some embodiments.
[0151] In some embodiments, at block 503, the processing logic optionally generates a first target output for one or more of the data inputs (e.g., the first data input). In some embodiments, the input includes one or more predicted products and an associated confidence interval detected within the content item, and the target output may include a label for the product included in the content item. In some embodiments, the input includes one or more sets of data associated with the content item (e.g., image data such as a video frame or a portion of a video frame, metadata such as title text or caption text, etc.), and the target output is a list of products included in the content item. In some embodiments, the first target output is predictive data. In some embodiments, the input data may be in the form of caption text data, and the target output may be a list of possible corrections to the caption to include product names / references for a machine learning model configured to correct the caption by including product information. In some embodiments, no target output is generated (e.g., an unsupervised machine learning model does not require a target output to be provided, but rather has the ability to group or find correlations in the input data).
[0152] At block 504, the processing logic optionally generates mapping data indicating an input / output mapping. The input / output mapping (or mapping data) may refer to a data input (e.g., one or more of the data inputs described herein), a target output for the data input, and the association between the data input(s) and the target output. In some embodiments, block 504 may not be executed, such as in relation to a machine learning model that does not provide a target output.
[0153] At block 505, in some embodiments, the processing logic adds the mapping data generated at block 504 to the data set T.
[0154] In block 506, the processing logic branches based on whether the data set T is sufficient for at least one of training, validating, and / or testing a machine learning model, such as one of the models 190 of FIG. 1. If so, execution proceeds to block 507; otherwise, execution returns following block 502. In some embodiments, whether the data set T is sufficient can be determined simply based on the number of inputs mapped to the output in the data set in some embodiments, and in some other embodiments, whether the data set T is sufficient can be determined based on one or more other criteria (e.g., a measure of the diversity of the data examples, accuracy, etc.) in addition to or instead of the number of inputs.
[0155] In block 507, the processing logic provides a dataset T (e.g., to server machine 180 of FIG. 1) for training, validating, and / or testing the machine learning model 190. In some embodiments, dataset T is a training set and is provided to the training engine 182 of server machine 180 to perform training. In some embodiments, dataset T is a validation set and is provided to the validation engine 184 of server machine 180 to perform validation. In some embodiments, dataset T is a test set and is provided to the test engine 186 of server machine 180 to perform testing. In the case of a neural network, for example, the input values of a given input / output mapping (e.g., the numerical values associated with data input 210) are input into the neural network, and the output values of the input / output mapping (e.g., the numerical values associated with target output 220) are stored at the output nodes of the neural network. The connection weights within the neural network are then adjusted by a learning algorithm (e.g., backpropagation, etc.), and the procedure is repeated for other input / output mappings within dataset T. After block 507, the model (e.g., model 190) can be at least one of trained using the training engine 182 of server machine 180, validated using the validation engine 184 of server machine 180, or tested using the test engine 186 of server machine 180. The trained model can be implemented by product identification system 175 to generate output data, for example, for use by product information platform 161 to provide product data to a user, to be provided to a fusion model, and to be utilized to update the metadata of content items to include one or more product associations.
[0156] FIG. 5B is a flowchart of a method 500B for updating metadata of a content item, according to some embodiments. At block 510, processing logic (e.g., a processing device, a computer processor, etc.) receives first data. The first data includes a first identifier of a first product (e.g., an indicator, a pointer to further data, a code identifying the product, etc.) determined in relation to the content item based on the metadata of the content item. The content item can be, or can include, visual content, audio content, text content, video content, etc. The first product may have been determined by providing the metadata of the content item to one or more trained machine learning models (e.g., the text identification module 350 of FIG. 3B). The metadata can include text data, such as a content item title, description, caption, comment, live chat, etc. The first data may further include a first confidence value associated with the first product and the content item. The first confidence value can indicate the likelihood that the first product is associated with the first content item, e.g., the likelihood that the first product appears in or is referenced within the metadata of the content item. The first data may further include an identifier of a second product and a second confidence value associated with the second product. The first data can include a list of products and associated confidence values.
[0157] In block 512, the processing logic receives second data that includes a second identifier of a first product. The second identifier is determined to be associated with the content item based on image data of the content item (e.g., one or more frames of a video, portions of one or more images, etc.). The second data also includes a second confidence value associated with the first product and the content item. The confidence value and the identifier can be generated by one or more machine learning models (e.g., the image identification module 330 of FIG. 3B). The machine learning model can include a system configured to reduce the dimensionality of the image data. One or more candidate product images can be dimensionally reduced (e.g., an image transformed through a trained machine learning model into a vector of values). The machine learning model can perform operations including comparing the reduced dimensionality image data from the content item with the dimensionally reduced product images in the data store to determine, for example, the likelihood that the image includes the product. The second data can include a list of products (e.g., product identifiers) and a list of confidence values that include at least the first product. The confidence value can indicate the likelihood that the associated product appears within the content item and is referenced by the content item, etc.
[0158] In some embodiments, one or more images of the content item are analyzed for potential products included therein. The presence of potential products can be examined, for example, by providing images of the content item (e.g., different frames of a video content item, additional frames, etc.) for further product image detection analysis, such as through text or metadata verification. For example, further analysis can be performed after a candidate product is found that is designated for verification of the candidate product, such as a search can be performed for other evidence of the identified product. In some embodiments, text data and / or metadata associated with the content item can be analyzed for potential / candidate products. The presence of potential products can be examined, for example, by image-based verification, text verification, etc.
[0159] In some embodiments, the processing logic may further provide one or more timestamps, such as the timestamp of a frame of a video having a detected candidate product, the timestamp of a caption associated with a video or audio content in which the detected candidate product appears, and the like. The processing logic may utilize the timestamps for further analysis, such as to adjust the metadata of the content item and to generate UI elements associated with the presentation of the content item.
[0160] In block 514, the processing logic provides first data and second data to a trained machine learning model. The trained machine learning model may be a fusion model. The trained machine learning model may provide one or more lists of products with associated confidence values.
[0161] In block 516, the processing logic receives a third confidence value associated with a first product from the trained machine learning model. In some embodiments, the processing logic may receive a list of confidence values associated with a list of products, including the first product.
[0162] In block 518, the processing logic adjusts the metadata associated with the content item in consideration of the third confidence value. In some embodiments, adjusting the metadata may include adding one or more connections between the content item and the product to the metadata. For example, adjusting the metadata may include adding indications such as that a particular product is associated with the content item, characterized within the content item, included within the content item, advertised by the content item, and the like. Adjusting the metadata may include, for example, adjusting the caption of the content item to include one or more references to a product that was mis-transcribed during caption generation.
[0163] FIG. 5C is a flowchart of a method 500C for training a machine learning model associated with content item-product pairings, according to some embodiments. In some embodiments, the machine learning model trained using method 500C can be a fusion model. Similar methods can be utilized to train different models connected to media item-product pairings, such as, for example, an image identification model, an image inspection model, a text identification model, a caption update model, and the like.
[0164] At block 520, the processing logic receives product image data associated with a plurality of content items. The product image data can include data associating the product with the content item, and the association can be derived from one or more images, such as frames of a video. The product image data includes indications of one or more products (e.g., potential products, candidate products) detected (e.g., determined) within the image and one or more product image confidence values.
[0165] At block 522, the processing logic receives product text data associated with a plurality of content items. The product text data can include data associating the product with the content item. The association can be derived from text associated with the content item, such as metadata associated with the content item. The product text data includes indications of one or more products (associated with the content item) detected within the text and one or more product text confidence values.
[0166] The data received (or, in some embodiments, obtained) by the processing logic at blocks 520 and 522 can be used as training inputs for training a fusion model. Training a machine learning model to perform different functions can include the processing logic receiving different data as training inputs.
[0167] In block 524, the processing logic receives data indicating products included in a plurality of content items. For example, each of the plurality of content items used to train a model (e.g., the data associated with the content item can be used to train the model) can include a list of associated products, such as being labeled by one or more users, labeled by a content creator, etc. The data received by the processing logic in block 524 can be used as a target output for training a fusion model. Training a machine learning model to perform different functions can include the processing logic receiving different data as the target output.
[0168] In block 526, the processing logic provides product image data and product text data to the machine learning model as training inputs. The processing logic can provide different types of data to train different machine learning models. In some embodiments, a machine learning model for frame selection can be trained by providing frames of a video to the model as training inputs. In some embodiments, a machine learning model for object detection can be trained by providing an image (optionally including a product) to the machine learning model as a training input. A machine learning model for embedding can be trained by providing one or more images of an object (e.g., a product) to the model as training inputs. In some embodiments, a text analysis model can be trained by providing text (e.g., metadata) associated with a content item as a training input. In some embodiments, a model configured to correct captions can be provided with machine-generated captions as training inputs.
[0169] In block 528, the processing logic provides data indicating products included in a plurality of content items (e.g., a list of products included in each content item of the plurality of content items) to the machine learning model as a target output. The processing logic may provide different types of data for training different machine learning models. In some embodiments, a machine learning model for frame selection may be trained by providing, as a target output, data indicating which frames of one or more videos contain products. In some embodiments, a machine learning model for object detection may be trained by providing, as a target output, labels of objects within an image provided to the model. In some embodiments, a text analysis model may be trained by providing, as a target output, content items referenced by the text of a content item. In some embodiments, a model configured to correct captions may be provided with a corrected caption (e.g., including one or more products) as a target output. In some embodiments, the target output is not provided for training a machine learning model (e.g., an unsupervised machine learning model).
[0170] FIG. 5D is a flowchart of a method 500D for adjusting metadata associated with a content item, according to some embodiments. In block 530, the processing logic obtains first metadata associated with the content item. The metadata may include text data. The metadata may include a content item title, description, caption, comment, live chat, and the like. In block 531, the processing logic provides the first metadata to a first model. In some embodiments, the model is a trained machine learning model. In some embodiments, the model is a product detection model, and for example, the model is configured to receive the metadata and generate an indication of a product associated with the content item (e.g., considering the metadata).
[0171] In block 532, the processing logic obtains, as the output of the first model, a first product identifier based on the first metadata and a first confidence value associated with the first product identifier. The product identifier can be an ID number, an indicator, a product name, or any data that (uniquely) distinguishes the product. The first product identifier can identify the first product. In some embodiments, the processing logic can obtain a list of products (e.g., candidate products, potential products) and associated confidence values.
[0172] In block 533, the processing logic obtains image data of the content item. In some embodiments, the image data can include one or more frames of a video or can be extracted from one or more frames. In some embodiments, the image data can be obtained from an object detection model. In some embodiments, the image data can include one or more products associated with the content item.
[0173] In block 534, the processing logic provides the image data to the second model. In some embodiments, the second model is a machine learning model. In some embodiments, the second model is a model configured to identify products from images. In some embodiments, the second model is a model configured to inspect for the presence of products identified from images. In some embodiments, the second model can reduce the dimensions of the provided image data. In some embodiments, the second model can compare the reduced-dimension image data with second reduced-dimension image data (e.g., retrieved from a data store, output by a machine learning model, etc.).
[0174] In block 535, the processing logic obtains, as the output of the second model, a second product identifier based on the image data and a second confidence value associated with the second product identifier. In some embodiments, the second product identifier indicates a second product. In some embodiments, the second product is the same as the first product. In some embodiments, the processing logic may obtain a list of products and associated confidence values.
[0175] In block 536, the processing logic provides, as input to the third model, data including the first product identifier, the first confidence value, the second product identifier, and the second confidence value. The third model may be a fusion model.
[0176] In block 537, the processing logic obtains, as the output of the third model, a third product identifier and a third confidence value. In some embodiments, the third confidence value may indicate the likelihood that the product indicated by the third product identifier is associated with (e.g., present within) the content item. In some embodiments, the third model may output a list of products and associated confidence values. In some embodiments, the third product identifier identifies a third product. In some embodiments, the third product is the same as the second product. In some embodiments, the third product is the same as the first product. In some embodiments, the first, second, and third products are all the same product.
[0177] In block 538, the processing logic adjusts the second metadata associated with the content item, taking into account the third product identifier and the third confidence value. Adjusting the metadata can include one or more product associations, for example, supplementing the metadata with indications of associated products. Adjusting the metadata can include, for example, updating the caption to include products that may have been inaccurately transcribed (e.g., inaccurately transcribed by a machine-generated caption model). In some embodiments, the processing logic may further receive one or more timestamps associated with the content item and one or more products (e.g., the time of the video at which the product was detected within the video image). Updating the metadata can include adding an indication of the time at which the product was found within the content item to the metadata.
[0178] Figure 5E is a flowchart of a method 500E for presenting UI elements associated with one or more products, according to some embodiments. At block 540, processing logic (e.g., of a user device, a client device, etc.) presents a UI. The UI includes one or more graphical representations (e.g., videos) of one or more content items. The graphical representation of a content item (e.g., a video thumbnail) can be selected to initiate presentation of the associated content item. One or more graphical representations of the content item can be displayed by one or more UI elements associated with the one or more products. Each graphical representation of each content item can be displayed by one or more UI elements associated with the one or more products. The UI element(s) can be presented / displayed in a folded state, e.g., a folded default state. The UI element includes information identifying a plurality of products covered by each content item. The UI element can identify that one or more products are associated with the content item. The UI element can identify one or more products associated with the content item (e.g., via a name, a picture, etc.) (e.g., covered within a video).
[0179] The UI can present selectable graphical representations of the content items. The presented content items can be part of a home feed, can be provided in response to a search, can be part of a viewing list, can be part of a play list, can be part of a shopping feed, etc. In some embodiments, the UI element can be overlaid on top of and / or in front of one or more other elements of the UI. For example, the UI element (e.g., in a folded state) can be overlaid on the graphical representation of the content item, can be overlaid on the content item (e.g., while the content item is being presented), etc.
[0180] In block 542, in response to a user interaction with a UI element in a folded state, the processing logic continues to facilitate the presentation of the graphical representation of each video and modifies the presentation of the UI element from the folded state to an expanded state. The interaction with the UI element may include selecting the UI element. The interaction with the UI element may include contemplating the UI element (e.g., placing a cursor on the UI element, scrolling on the UI element, and pausing the scroll, etc.). The UI element in the expanded state may include a plurality of visual components. Each visual component may be associated with one of a plurality of products. The visual component may include pictures, descriptions, prices, timestamps, etc. associated with various products.
[0181] In some embodiments, the UI element (e.g., in an expanded state) may include a plurality of tabs. For example, the UI element may include tabs for products, chapters, or parts of content items. The UI element may display / open the tab for the product by default for a content item having an associated product. The UI element may display the tab for the product by default in response to a user action and / or history.
[0182] In block 544, in response to a user selection of one of the plurality of visual components of the UI element in the expanded state, the processing logic initiates the presentation of each content item covering the product associated with the selected visual component. The processing logic may initiate the playback of a video covering the product associated with the selected visual component. The processing logic may initiate the presentation of a portion of the content item associated with the product of the selected visual component (e.g., initiate the playback of a portion of the video).
[0183] In some embodiments, interacting with a UI element may cause a modification of the UI element to a product focus state. The product focus state may present additional information, detailed information, etc. regarding one or more products. Interacting with a UI element in a collapsed state may cause a modification of the UI element to a product focus state. Interacting with a UI element in an expanded state (e.g., interacting with a visual component of a UI element associated with a product) may cause a modification of the UI element to a product focus state.
[0184] In some embodiments, interacting with a UI element may cause a modification of the presentation of the UI element to a transaction state. Interacting with a UI element in a collapsed state may cause a modification of the presentation of the UI element to a transaction state. Interacting with a UI element in an expanded state (e.g., selecting a component associated with a product) may cause a modification of the presentation of the UI element to a transaction state. Interacting with a UI element in a product focus state may cause a modification of the presentation of the UI element to a transaction state. Determining whether a selection or interaction, such as with a UI element or a UI element component, causes a transition to a transaction state may be performed based on user history, user preferences, content items, or a content item feed (e.g., search results, viewing feeds, etc.).
[0185] In some embodiments, the UI element can be overlaid on the presented content item. For example, a UI element that identifies one or more products can be displayed and overlaid on a video while the video is playing, while the video shows one or more products, etc. Selection of the overlaid UI element can cause additional UI elements to be displayed, the overlaid UI element to be modified to a different state, a separate UI element to be modified to a different state, and so on. The overlaid UI element can be in a collapsed state, an expanded state, a product focus state, a transaction state, etc. The presence and / or position of the overlaid UI element can be determined by one or more trained machine learning models, such as one or more models configured to detect products.
[0186] Figure 5F is a flowchart of a method 500F for instructing a device to present one or more UI elements associated with a product, according to some embodiments. At block 550, the processing logic provides the device with a UI that includes one or more graphical representations of one or more content items. The graphical representations are provided for display / presentation by the device's UI. Each graphical representation of each content item is selectable to initiate the presentation of each content item. The content items may include videos. The content items may include live-streamed videos. The graphical representations may be provided in response to a request by the device. The graphical representations may include a home feed, a viewing feed, a playlist, a search results list, a shopping feed, and the like. Commands sent to the device (e.g., commands associated with any step of method 500F), the UI sent to the device, the UI elements sent to the device, etc. may be determined / selected based on obtaining the user's history. The user's history may include content item history interactions and / or selections, including content items having the associated product. The user's history may include history interactions and / or selections of UI elements or components of UI elements associated with the product. The user's history may include one or more searches by the user, e.g., searches including the product name. The commands may be provided to the device in response to the processing logic receiving the user's history.
[0187] One or more of the graphical representations of the content item are displayed by a UI element in a folded state. In some embodiments, each graphical representation is displayed by a UI element in a folded state. In some embodiments, a subset of the graphical representations is displayed by a UI element in a folded state. The UI element in the folded state is presented / displayed by a first graphical representation of a first content item. The UI element includes information identifying a plurality of products covered by the first content item. The UI element can identify how many products are associated with the content item, can identify one or more products by name, can identify the category or classification of the products covered by the content item, and so on. In some embodiments, the plurality of products are obtained as output from one or more trained machine learning models. The trained machine learning model can be similar to those described in relation to Figure 3B.
[0188] In block 554, in response to receiving an indication of user interaction with the UI element in the folded state, the device, by the processing logic, modifies the presentation of the UI element. The presentation of the UI element can be modified from the folded state to an expanded state. The UI element in the expanded state can include a plurality of visual components, each of which is associated with one of the plurality of products. The visual components can include pictures, names, descriptions, prices, timestamps, and so on.
[0189] In block 556, in response to receiving an indication of a user selection of one of a plurality of visual components of a UI element in an expanded state, the processing logic facilitates the presentation of a first content item. The processing logic may provide instructions to facilitate the presentation of a portion of the first content item associated with a first product, e.g., a product associated with one of the plurality of visual components. The processing logic may provide instructions to display a portion of a video associated with the product (e.g., to start playing the video from a selected point within the video based on a timestamp associated with the product).
[0190] In some embodiments, the processing logic may further provide instructions to the device to modify the presentation of the UI element to a product focus state. For example, upon selection of a visual component of a UI element in an expanded state, the UI element may be modified to a product focus state. The product focus state may encompass, be included in, or be associated with a content item and may include additional details regarding one or more products related to the content item.
[0191] In some embodiments, the processing logic may further provide instructions to the device to modify the presentation of the UI element to a transaction state. The transaction state may facilitate a user to initiate a transaction associated with a product (e.g., purchasing a product). The transaction state may be presented in response to a user action, user history, user selection of one or more UI elements, etc. A UI element in a transaction state may include one or more components that facilitate a transaction associated with one or more products.
[0192] FIG. 6 is a block diagram illustrating a computer system 600 according to some embodiments. In some embodiments, computer system 600 may be connected to other computer systems (e.g., via a network such as a local area network (LAN), intranet, extranet, or the Internet). Computer system 600 may operate as a server or client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. Computer system 600 may be provided by a device having the ability to execute a set of instructions (sequential or otherwise) that specify actions to be taken by a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, web appliance, server, network router, switch or bridge, or any device. Further, the term “computer” shall be construed to include any collection of computers that individually or jointly execute a set (or sets) of instructions to perform any one or more of the methods described herein.
[0193] In a further aspect, computer system 600 may include a processing device 602, volatile memory 604 (e.g., random access memory (RAM)), non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.
[0194] The processing device 602 can be provided by one or more processors such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of types of instruction sets), or a special processor (e.g., an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor, etc.).
[0195] The computer system 600 can further include a network interface device 622 (e.g., coupled to the network 674). The computer system 600 can also include a video display unit 610 (e.g., an LCD), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generation device 620.
[0196] In some embodiments, the data storage device 618 can include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable medium) that encodes the components of FIG. 1 (e.g., the content providing platform 120, other platforms of the content platform system 102, the communication application 115, the model 190, etc.) and stores instructions 626 that encode any one or more of the methods or functions described herein, including instructions for implementing the methods described herein.
[0197] The instructions 626 can also be fully or partially present in the volatile memory 604 and / or within the processing device 602 during its execution by the computer system 600, and thus, the volatile memory 604 and the processing device 602 can also constitute a machine-readable storage medium.
[0198] A computer-readable storage medium 624 is illustrated as a single medium in an exemplary embodiment, and the term "computer-readable storage medium" includes a single medium or multiple media (e.g., a centralized or distributed database, and / or an associated cache and server) that store one or more sets of executable instructions. The term "computer-readable storage medium" also includes any tangible medium that has the ability to store or encode a set of instructions for causing a computer to execute any one or more of the methods described herein. The term "computer-readable storage medium" includes, but is not limited to, solid state memory, optical media, and magnetic media.
[0199] The methods, components, and features described herein can be implemented by discrete hardware components or integrated in the functionality of other hardware components such as ASICs, FPGAs, DSPs, or similar devices. Additionally, the methods, components, and features can be implemented by firmware modules or functional circuits within a hardware device. Further, the methods, components, and features can be implemented in any combination of a hardware device and computer program components or in a computer program.
[0200] Unless otherwise specified, terms such as "receiving", "executing", "providing", "obtaining", "causing to occur", "accessing", "determining", "adding", "using", "training", "reducing", "generating", or "correcting" refer to actions and processes performed or implemented by a computer system that manipulates and transforms data represented as a physical (electronic) quantity within a computer system register and memory into other data similarly represented as a physical quantity within a computer system memory or register, or other such information storage, transmission, or display device. Also, the terms "first", "second", "third", "fourth", etc. used herein are intended as labels to distinguish different elements and may not necessarily have the meaning of ordinal numbers following their numerical designations.
[0201] The embodiments described herein also relate to an apparatus for executing the methods described herein. This apparatus can be specially constructed to execute the methods described herein, or it can include a general-purpose computer system selectively programmed by a computer program stored in a computing system. Such a computer program can be stored in a computer-readable tangible storage medium.
[0202] The methods and illustrative embodiments described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems can be used in accordance with the teachings described herein, or it may be convenient to construct more specialized apparatus for executing each of the methods and / or their individual functions, routines, subroutines, or operations described herein. Examples of structural embodiments for various such systems are shown in the above description.
[0203] The above description is intended to be illustrative and not restrictive. Although the present disclosure has been described with reference to specific exemplary embodiments and implementations, it will be recognized that the present disclosure is not limited to the described embodiments and implementations. The scope of the present disclosure should be determined with reference to the following claims, along with the full scope of equivalents to which the claims are entitled.
[0204] References throughout this specification to "one implementation" or "an implementation" mean that a particular feature, structure, or characteristic described in connection with the implementation is included in at least one implementation. Thus, the appearances of the phrases "in one implementation" or "in an implementation" in various places in this specification are not necessarily all referring to the same implementation. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more implementations.
[0205] As used in any detailed description or claims, the terms "includes", "including", "has", "contains", their variations, and other similar words are intended to be as inclusive as the term "comprising" as an open transitional word, without excluding additional elements or other elements.
[0206] As used in this application, terms such as "component", "module", or "system" generally intend to refer to a computer-related entity, either hardware (e.g., circuitry), software, a combination of hardware and software, or an entity related to an operable machine with one or more specific functionalities. For example, a component can be, but is not limited to, a process running on a processor (e.g., a digital signal processor), a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a controller and the controller can be components. One or more components may be present within a process and / or thread of execution, and a component may be localized on one computer and / or distributed between two or more computers. Further, a "device" can be in the form of specially designed hardware, generalized hardware specialized by the execution of software thereon enabling the hardware to perform a specific function (e.g., generating a point of interest and / or an explanation), software on a computer-readable medium, or a combination thereof.
[0207] The foregoing systems, circuits, modules, etc. have been described with respect to the interactions between several components and / or blocks. Such systems, circuits, components, blocks, etc. may include those components or designated sub-components, a part of the designated components or sub-components, and / or additional components, and it will be understood that they follow the various substitutions and combinations described above. Sub-components may not be included in the parent component (hierarchical type), but can also be implemented as components communicably coupled to other components. Furthermore, it should be noted that one or more components can be combined into a single component that provides an integrated function, or can be divided into several separate sub-components, and that any one or more intermediate layers, such as a management layer, can be provided to communicably couple with such sub-components to provide an integrated function. Any component described herein may also interact with one or more other components that are not specifically described herein but are known to those skilled in the art.
[0208] Moreover, the words "example" or "exemplary" are used herein to mean serving as an example, instance, or illustration. Any aspect or design described as "exemplary" herein should not necessarily be construed as more preferred or advantageous than other aspects or designs. Rather, the use of the words "example" or "exemplary" is intended to present concepts in a concrete fashion. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" is intended to mean any of the natural inclusive permutations. That is, "X uses A or B" is satisfied in any of the following cases: when X uses A, when X uses B, or when X uses both A and B. Additionally, the articles "a" and "an" used in this application and the appended claims should generally be construed to mean "one or more" unless otherwise specified or it is clear from the context that the singular form is being referred to.
Claims
1. obtaining, by a processing device, first data including (i) a first identifier of a first product determined in relation to the content item based on first metadata of the content item and (ii) a first confidence value associated with the first product and the content item; obtaining, by the processing device, second data including (i) a second identifier of the first product determined in relation to the content item based on first image data of the content item and (ii) a second confidence value associated with the first product and the content item; providing, by the processing device, the first data and the second data to a trained machine learning model; obtaining, from the trained machine learning model, a third confidence value associated with the first product; adjusting second metadata associated with the content item in consideration of the third confidence value; A method comprising the above.
2. providing, to a second model, the first metadata of the content item as an input; obtaining, as an output of the second model, the first data; The method according to claim 1, further comprising the above.
3. The first metadata includes a title of the content item, a description of the content item, or a caption associated with the content item, The method according to claim 2, including at least one of the above.
4. providing, to a second model, the first image data of the content item as an input; obtaining, as an output of the second model, data reduced in a first dimension; obtaining, from a data store, data reduced in a second dimension associated with the first product; The method according to claim 1, further comprising the above.
5. The data reduced in the second dimension is obtained from the data store in response to obtaining the first data, and the second data is generated based on at least the data reduced in the first dimension and the data reduced in the second dimension. The method according to claim 4.
6. providing second image data to a third model; Obtaining a third identifier of the first product from the third model, wherein the data reduced in the second dimension is obtained from the data store in response to obtaining the third identifier of the first product, and the second data is generated based on at least the data reduced in the first dimension and the data reduced in the second dimension; The method according to claim 4, further comprising. **Claim 7** The content item is a video, the first data further includes an indication of the time stamps of one or more frames of the video associated with the product, and adjusting the second metadata includes including the indication of the first product and the indication of the time stamps in the second metadata. The method according to claim 1. **Claim 8** Adjusting the second metadata includes adjusting a caption associated with the product. The method according to claim 1. **Claim 9** Further comprising training a machine learning model to generate the trained machine learning model, and training the machine learning model includes Receiving image-based product data associated with a plurality of content items, the image-based product data including one or more products detected in the image and an indication of one or more product image confidence values; Receiving metadata-based product data associated with the plurality of content items, the metadata-based product data including one or more products detected in the text and an indication of one or more confidence values; Receiving data indicating products included in the plurality of content items; Providing the image-based product data and the metadata-based product data as training inputs to the machine learning model; Providing the data indicating the products included in the plurality of content items as target outputs to the machine learning model; The method according to claim 1, including. **Claim 10** Receiving third data including a third identifier of a first product category associated with the content item; Providing the third data to the trained machine learning model, wherein the third confidence value is generated based on the first data, the second data, and the third data; The method according to claim 1, further comprising.
11. Obtaining, by a processing device, first metadata associated with a content item; Providing the first metadata to a first model; Obtaining the first product identifier as an output of the first model based on the first metadata and a first confidence value associated with a first product identifier; Obtaining image data of the content item; Providing the image data to a second model; Obtaining the second product identifier as an output of the second model based on the image data and a second confidence value associated with a second product identifier; To a third model, The first product identifier, The first confidence value, The second product identifier, The second confidence value, Providing, as an input, data including; Obtaining a third product identifier and a third confidence value as outputs of the third model; Adjusting second metadata associated with the content item in consideration of the third product identifier and the third confidence value; A method comprising
12. Generating the second confidence value comprises: Reducing the dimension of the image data so as to generate first dimensionally reduced data; Obtaining, from a data store, second dimensionally reduced data, wherein the second dimensionally reduced data is associated with a product indicated by the second product identifier; Performing one or more operations so as to generate the second confidence value, wherein the second confidence value is based on one or more differences between the first dimensionally reduced data and the second dimensionally reduced data; The method according to claim 11, comprising
13. The method according to claim 11, wherein the first product identifier, the second product identifier, and the third product identifier each identify a first product.
14. Further comprising obtaining a time stamp associated with the image data and the content item, and adjusting the second metadata associated with the content item includes adjusting the second metadata such that it includes an indication that the product identified by the second product identifier is associated with the time stamp and the content item, the method of claim 11.
15. The second metadata includes a machine-generated caption, the first product identifier is associated with a product, the language associated with the product was inaccurately transcribed during generation of the machine-generated caption, and updating the second metadata associated with the content item includes replacing a portion of the machine-generated caption associated with the product with a text identifier of the product, the method of claim 11.
16. Providing a fourth product identifier and a fourth confidence value to the third model; Obtaining a fifth product identifier as an output of the third model, wherein the third product identifier is associated with a first product and the fifth product identifier is associated with a second product; The method of claim 11, further comprising.
17. When executed, the processing device Obtaining first data including (i) a first identifier of a first product determined in relation to the content item based on first metadata of the content item and (ii) a first confidence value associated with the first product and the content item; Obtaining second data including (i) a second identifier of the first product determined in relation to the content item based on first image data of the content item and (ii) a second confidence value associated with the first product and the content item; Providing the first data and the second data to a trained machine learning model; Obtaining a third confidence value associated with the first product from the trained machine learning model; Adjusting second metadata associated with the content item in consideration of the third confidence value; A non-transitory machine-readable storage medium storing instructions for causing an operation to be performed, the operation including **Claim 18** wherein the operation further includes providing, as an input, the first image data of the content item to a second model; obtaining, as an output of the second model, data dimensionally reduced in a first dimension; in response to obtaining the first data from a data store, obtaining second dimensionally reduced data associated with the first product, the second data being based at least on the first dimensionally reduced data and the second dimensionally reduced data; The non-transitory machine-readable storage medium according to claim 17, further comprising **Claim 19** wherein the content item is a video, the first data further includes an indication of a time stamp of one or more frames of the video associated with the product, and adjusting the second metadata includes including an indication of the first product and the indication of the time stamp in the second metadata. The non-transitory machine-readable storage medium according to claim 17. **Claim 20** wherein the operation further includes receiving third data including a third identifier of a first product category associated with the content item; providing the third data to the trained machine learning model, the third confidence value being generated based on the first data, the second data, and the third data; The non-transitory machine-readable storage medium according to claim 17, further comprising
Citation Information
Patent Citations
Method and system for identifying relevant media content
JP2018078576A
Desired video information notification system
JP2019193023A
Object detection through visual search queries
JP2019531547A
Video search engine using joint categorization of video clips and queries based on multiple modalities
US20070255755A1
Machine-based object recognition of video content
WO2018094201A1