Product identification in media items

Through the machine learning model combining image and text data, identifying and adjusting the metadata associated with content items, the problem of difficult to identify product related to media items in the prior art is solved, efficient product automatic recognition and association is achieved, and user experience and system efficiency are improved.

CN119948475APending Publication Date: 2025-05-06GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202380057656.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-01
Filing Date
2023-07-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and automatically identify products associated with media items, making it difficult for users to find relevant product information, and the information is not updated in time, affecting the user's purchasing experience.

Method used

The machine learning model is used to combine image and text data, and the dimensional reduction data is generated by training the model, and the metadata associated with the content items is identified and adjusted to achieve automatic identification and association of the product.

Benefits of technology

It improves the accuracy of automated identification of products in content items, reduces the time and computing resources for users to find product information, and improves user experience and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948475A_ABST
    Figure CN119948475A_ABST
Patent Text Reader

Abstract

A method includes obtaining first data including a first identifier of a first product determined in association with a content item based on first metadata of the content item. The method further includes obtaining a first confidence value associated with the first product and the content item. The method further includes obtaining second data including a second identifier of the first product and a second confidence value. The method further includes providing the first data and the second data to the trained machine learning model. The method further includes obtaining a third confidence value associated with the first product from the trained machine learning model. The method further includes adjusting second metadata of the content item in view of a third confidence value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects and implementations of the present disclosure relate to methods and systems for facilitating pairing of media items and associated objects, and more particularly to systems for identifying pairings of media items and products. Background Art

[0002] A platform (e.g., a content sharing platform) can transmit (e.g., stream) media items to client devices connected to the platform via a network. Different types of client devices can be optimized for different tasks, preferred by users for different tasks, etc. A media item can contain references to one or more products. Summary of the invention

[0003] The following summary of the invention is a simplified summary of the present disclosure in order to provide a basic understanding of some aspects of the present disclosure. This summary of the invention is not a comprehensive review of the present disclosure. It is neither intended to identify the key or important elements of the present disclosure, nor to indicate any scope of a specific implementation of the present disclosure or any scope of the claims. Its only purpose is to present some concepts of the present disclosure in a simplified form as a preface to a more detailed description presented later.

[0004] Systems and methods for facilitating identification of one or more products associated with a media item are disclosed. In some implementations, a method includes obtaining first data. The first data includes a first identifier of a first product determined in association with a content item based on first metadata of the content item. The first data further includes a first confidence value associated with the first product and the content item. The method further includes obtaining second data. The second data includes a second identifier of the first product determined in association with the content item based on first image data of the content item. The second data further includes a second confidence value associated with the first product and the content item. The method further includes providing the first data and the second data to a trained machine learning model. The method further includes obtaining a third confidence value associated with the first product from the trained machine learning model. The method further includes adjusting second metadata associated with the content item in view of the third confidence value.

[0005] In some embodiments, the method further includes providing first metadata of the content item as input to the second model. The method may include obtaining first data as output of the second model. In some embodiments, the method further includes providing first image data of the content item as input to the second model. The method may further include obtaining first dimensionally reduced data as output of the second model. The method may further include obtaining second dimensionally reduced data associated with the first product from a data repository. The second dimensionally reduced data may be obtained from the data repository in response to obtaining the first data. The second data may be generated based on at least the first dimensionally reduced data and the second dimensionally reduced data.

[0006] The method may further include providing the second image data to a third model. The method may further include obtaining a third identifier of the first product from the third model. The second dimensionally reduced data may be obtained from a data repository in response to obtaining the third identifier of the first product. The second data may be generated based on at least the first dimensionally reduced data and the second dimensionally reduced data.

[0007] In some embodiments, the content item is a video. The first data may further include an indication of a timestamp of one or more frames of the video associated with the product. Adjusting the second metadata may include including an indication of the first product and an indication of the timestamp in the second metadata.

[0008] In some embodiments, the metadata may include a title of the content item. The metadata may include a description of the content item. The metadata may include a narration associated with the content item. Adjusting the second metadata may include adjusting a narration associated with the product.

[0009] In some embodiments, the method further includes training the machine learning model to generate a trained machine learning model. Training the machine learning model may include receiving image-based product data associated with a plurality of content items. The image-based product data may include indications of one or more products detected in the image and one or more product image confidence values. Training may further include receiving metadata-based product data associated with a plurality of content items. The metadata-based product data may include indications of one or more products detected in the text and one or more confidence values. Training may further include receiving data indicating products included in the plurality of content items. Training may further include providing the image-based product data and the metadata-based product data as training inputs to the machine learning model. Training may further include providing data indicating products included in the plurality of content items as target outputs to the machine learning model.

[0010] In some embodiments, the method further includes receiving third data, the third data including a third identifier of the first product category associated with the content item. The method may further include providing the third data to a trained machine learning model. The third confidence value may be based on the first data, the second data, and the third data.

[0011] On the other hand, a method includes obtaining first metadata associated with a content item. The method further includes providing the first metadata to a first model. The method further includes obtaining, as an output of the first model, a first product identifier based on the first metadata and a first confidence value associated with the first product identifier. The method further includes obtaining image data of the content item. The method further includes providing the image data to a second model. The method further includes obtaining, as an output of the second model, a second product identifier based on the image data and a second confidence value associated with the second product identifier. The method further includes providing, as input to a third model, data including the first product identifier, the first confidence value, the second product identifier, and the second confidence value. The method further includes obtaining, as an output of the third model, a third product identifier and a third confidence value. The method further includes adjusting the metadata associated with the content item in view of the third product identifier and the third confidence value.

[0012] In some embodiments, generating the second confidence value includes reducing the dimension of the image data to generate the first dimensionally reduced data. Generating the second confidence value may further include obtaining the second dimensionally reduced data from a data repository. The second dimensionally reduced data may be associated with the product indicated by the second product identifier. Generating the second confidence value may further include performing one or more operations to generate the second confidence value. The second confidence value may be based on one or more differences between the first dimensionally reduced data and the second dimensionally reduced data.

[0013] In some embodiments, each of the first product identifier, the second product identifier, and the third product identifier identifies a first product.

[0014] In some embodiments, the method further comprises obtaining a timestamp associated with the image data and the content item.Adjusting the second metadata associated with the content item may comprise adjusting the second metadata to include an indication that the product identified by the second product identifier is associated with the timestamp and the content item.

[0015] In some embodiments, the second metadata includes a machine-generated narration. The first product identifier may be associated with the product. The language associated with the product may have been incorrectly transcribed when the machine-generated narration was generated. Updating the second metadata associated with the content item may include replacing a portion of the machine-generated narration associated with the product with a text identifier of the product.

[0016] In some embodiments, the method further includes providing a fourth product identifier and a fourth confidence value to the third model. The method may further include obtaining a fifth product identifier as an output of the third model. The third product identifier is associated with the first product, and the fifth product identifier may be associated with the second product.

[0017] On the other hand, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform an operation, the operation comprising obtaining first data. The first data includes a first identifier of a first product determined in association with a content item based on first metadata of the content item. The first data further includes a first confidence value associated with the first product and the content item. The operation further includes obtaining second data, the second data including a second identifier of the first product determined in association with the content item based on first image data of the content item. The second data further includes a second confidence value associated with the first product and the content item. The operation further includes providing the first data and the second data to a trained machine learning model. The operation further includes obtaining a third confidence value associated with the first product from the trained machine learning model. The operation further includes adjusting the second metadata associated with the content item in view of the third confidence value.

[0018] In some embodiments, the operation further includes providing the first image data of the content item as an input to the second model. The operation may further include obtaining first dimensionally reduced data as an output of the second model. The operation may further include obtaining second dimensionally reduced data associated with the first product from a data repository. The data from the data repository may be obtained in response to obtaining the first data. The second data may be based at least on the first dimensionally reduced data and the second dimensionally reduced data.

[0019] In some embodiments, the content item is a video. The first data may further include an indication of a timestamp of one or more frames of the video associated with the product. Adjusting the second metadata may include including an indication of the first product and an indication of the timestamp in the second metadata.

[0020] In some embodiments, the operation further includes receiving third data, the third data including a third identifier of the first product category associated with the content item. The operation may further include providing the third data to a trained machine learning model. A third confidence value may be generated based on the first data, the second data, and the third data.

[0021] Where appropriate, optional features of one aspect may be combined with other aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Aspects and implementations of the present disclosure will be more fully understood from the detailed description given below and from the drawings of various aspects and implementations of the present disclosure, however, they should not be considered to limit the present disclosure to specific aspects or implementations, but are only for explanation and understanding.

[0023] Figure 1 An example system architecture for providing relevant and associated product information in accordance with some embodiments is shown.

[0024] Figure 2 is a block diagram of a system 200 including a dataset generator for creating datasets for one or more models, according to some embodiments.

[0025] Figure 3A is a block diagram illustrating a system for generating output data, such as associations between products and content items, according to some embodiments.

[0026] Figure 3B is a block diagram of an example system for generating data describing an association between a content item and one or more products, according to some embodiments.

[0027] Figure 4A Depicted is an apparatus presenting an example user interface (UI) including a UI element indicating a product associated with a content item in accordance with some embodiments.

[0028] Figure 4B Depicted is an apparatus presenting an example UI including a UI element indicating a product associated with a content item in accordance with some embodiments.

[0029] Figure 4C Depicted is an apparatus presenting an example UI including a UI element presenting a content item and a UI element presenting information about a product associated with the content item in accordance with some embodiments.

[0030] Figure 4D Depicted is an apparatus presenting an example UI including UI elements that present content items and UI elements that facilitate transactions associated with products in accordance with some embodiments.

[0031] Figure 4E Depicted is an example device having a UI element overlaid on a content presentation element in accordance with some embodiments.

[0032] Figure 5A is a flowchart of a method for generating a dataset for a machine learning model according to some embodiments.

[0033] Figure 5B is a flow diagram of a method for updating metadata for a content item according to some embodiments.

[0034] Figure 5C is a flow diagram of a method for training a machine learning model associated with content item and product pairings according to some embodiments.

[0035] Figure 5D is a flow diagram of a method for adjusting metadata associated with a content item according to some embodiments.

[0036] Figure 5E is a flow diagram of a method for presenting UI elements associated with one or more products according to some embodiments.

[0037] Fig. 5F is a flow chart of a method for instructing a device to present one or more UI elements associated with a product according to some embodiments.

[0038] Figure 6 is a block diagram illustrating an exemplary computer system according to an implementation of the present disclosure. DETAILED DESCRIPTION

[0039] Aspects of the present disclosure relate to methods and systems for facilitating the pairing of media items (e.g., content items, content, etc.) with associated products. A platform (e.g., a content sharing platform, etc.) may enable a user to access (e.g., via a client device connected to the platform) media items (e.g., video items, audio items, etc.) hosted by the platform. The platform may provide access to media items to a user's client device via a network (e.g., the Internet) (e.g., by transmitting the media items to the user's client device, etc.). The media / content items may have one or more additional associated activities that may deepen the user's interaction with the content items. For example, an interactive activity may include a comment area for a content item, a real-time chat associated with a content item, etc. Some content items may be associated with one or more products. For example, a product may be displayed as part of a content item, a product may be evaluated in a content item, a content item may be associated with one or more products (e.g., sponsored by a company connected to the one or more products), etc.

[0040] In conventional systems, identifying products associated with a content item may be difficult, inconvenient, time consuming, etc. It may be difficult for a content platform (e.g., a platform that provides content for presentation to a user) to identify products associated with a content item. In some systems, a content creator may identify one or more products associated with a content item, a content channel, a list of content items, etc. In some systems, a content creator may include information about one or more products in a content item, such as a photo of the item, an item that appears in a video, etc. In some systems, a content creator may include information about one or more products in one or more fields associated with a content item, such as a content item title, a content item description, a content item comment (e.g., a pinned comment), etc.

[0041] In conventional systems, a user may not be able to easily determine whether a content item has one or more associated products. A content item may not have an indicator of the associated products. A user may need to be presented with a content item (e.g., watch a video) to determine whether a content item has one or more associated products.

[0042] In conventional systems, it may be difficult for a user to identify products associated with a content item. There may be no direct indication that a content item has an associated product (e.g., when selecting a content item to consume, to view, etc.). Product information may be difficult to find, e.g., scattered among different fields such as a content item title, a content item description, a comment area, etc. Product information may be included in a content item, e.g., a video may include audio describing one or more products, an image item may include an image of one or more products, etc. Extracting this information from a content item may be difficult, time consuming, error prone, etc.

[0043] In conventional systems, receiving additional information about a product that features in a content item may be difficult, time-consuming, etc. In some systems, a product may appear in a content item (e.g., may appear in a video). A user may not be given additional information (e.g., product name, merchant name, etc.), and may perform a search independently of the content platform to learn more about the product. A name and / or merchant associated with the product may be provided (e.g., in the description or title of the content item). A user may perform a search (e.g., independent of the content platform) to learn more about the product, such as product variations, related products, availability, and price, etc. A guide for promoting a better understanding of the product (e.g., a guide to the user, a guide to the processing device in the form of a link to the merchant's website, etc.) may be provided. A user may receive additional information from a separate source (e.g., a website) independent of the content platform.

[0044] In conventional systems, information about products included in content items may become outdated. For example, within a content item or associated fields (e.g., content item title, content item description, content item reviews, etc.), a content creator may include additional information about one or more products, such as price information, merchant information, availability information, alternate versions or variations information, etc. Some information included in a content item or associated fields may not be updated with changes to that information, may be updated depending on the content creator, etc.

[0045] In conventional systems, there may be obstacles that make it difficult, time-consuming, cumbersome, etc. for a user to purchase one or more items associated with a content item. A user may search for a product, search for a merchant that stocks or sells a product, etc. In some embodiments, a content item or an associated domain may include instructions for purchasing an item (e.g., a description of a content item may include one or more links to products associated with the content item). A user may be directed to another platform independent of the content platform to complete the purchase of one or more products.

[0046] The user may spend a lot of time and computing resources to find information about the products covered by the content item. For example, the video may be very long, and the product of interest may not be shown until the end. The user may spend a lot of time consuming the video to obtain accurate information about the product of interest, resulting in increased use of the computing resources of the client device. In addition, the computing resources of the client device that enable the user to consume the media item may not be available for other processes, which can reduce overall efficiency and increase the overall latency of the client device.

[0047] Aspects of the present disclosure can solve one or more of these shortcomings of conventional methods. In certain embodiments, aspects of the present disclosure can realize the automatic identification of products displayed and / or included in content items. In certain embodiments, aspects of the present disclosure can realize the use of models to identify products from text associated with content items. Text can include the title of content items, the description of content items, the commentary associated with content items, etc. The commentary can be machine-generated commentary, such as the commentary generated by a speech-to-text model for video or audio content items. The model can be a machine learning model. The model can output a confidence value indicating that a content item includes the possibility of a product.

[0048] In some embodiments, aspects of the present disclosure may implement the use of a model to identify a product from an image of a content item. The content item may be a picture or include a picture. The content item may be or include a video. One or more images of the content item may be provided to a model configured to identify a product from an image. The model may reduce the dimensionality of the image. The model may search for similar images associated with the product. The model may search for similar dimensionality-reduced images in a reduced dimensional space. The model may be a machine learning model. The model may output a confidence value indicating the likelihood that the content item includes a product.

[0049] In some embodiments, aspects of the present disclosure implement the use of a model (e.g., a fusion model) to determine whether a product is included in the content item. The fusion model may receive an indication of one or more products detected by a model that receives text associated with the content item as input. The fusion model may receive an indication of a confidence value that one or more products are included in the text. The fusion model may receive an indication of one or more products detected by a model that receives an image of the content item as input. The fusion model may receive an indication of a confidence value that one or more products are included in the image. The fusion model may determine the confidence that one or more products appear in the content item and associated data (e.g., title, description, etc.). The fusion model may be a machine learning model.

[0050] In some embodiments, product detection can be used to improve content or associated information (e.g., descriptions, captions). The model can be provided with data associated with a content item. For example, the model can be provided with a machine-generated caption for a content item. The model can identify one or more captions that may be misrepresentations of a product name (e.g., the machine-generated caption may include the closest English equivalent to the spoken product name). The model can be provided with one or more images of the content item. The model can determine the likelihood that one or more products are included in the image. Information associated with the content item (e.g., metadata, descriptions, captions, etc.) can be updated in view of the one or more detected products.

[0051] In some embodiments, aspects of the present disclosure implement indicators of content items with one or more associated products. A list of content items may include one or more indicators that one or more content items in the content items in the list include associated products. Indicators may include visual indicators (e.g., a "shopping" symbol or text indicating one or more products displayed in association with a content item), additional fields (e.g., a panel including product information), etc. In some embodiments, a list of content items may be presented to a user via a user interface (UI). UI elements may be associated with a content platform (e.g., may be presented via an application associated with a content providing platform). UI may include elements indicating that a content item is associated with one or more products. In some embodiments, user interaction with UI elements may cause the presentation of additional UI elements with additional information, e.g., a list of products associated with a content item.

[0052] In some embodiments, aspects of the present disclosure enable users to identify products associated with content items. In some embodiments, a list of content items for presentation to a user may include a list of products associated with one or more content items. For example, a list of content items may be displayed to a user via a UI. The UI may include one or more UI elements that present one or more products associated with a content item to the user. The UI element may be provided via an application associated with a content providing platform for the content item. In some embodiments, a UI element that lists products associated with a content item may be presented in response to detecting a user interaction with a UI element indicating that a content item has associated products. In some embodiments, a user interaction with a product in the list of products may cause the UI to present additional information to the user.

[0053] In some embodiments, aspects of the present disclosure implement providing additional information about one or more products associated with a content item to a user. A UI element may be provided to the user, which provides additional information about a product. The UI element may provide information such as changes in a product (e.g., color changes, size changes, etc.), related products, availability (e.g., associated with one or more merchants), and / or prices. The UI element may be provided via an application associated with a content item, a content providing platform, etc. The UI element displaying additional product information may be presented in response to a user interaction with another UI element, such as selecting a content item to be viewed, selecting a UI element indicating one or more associated products, etc. In some embodiments, after a user interaction with a UI element, further UI elements may be provided, such as to facilitate the purchase of a product.

[0054] In some embodiments, aspects of the present disclosure implement automatic updating of information connected to one or more products associated with a content item. In some embodiments, a content providing platform associated with a content item may include one or more memory devices containing product data, communicate with one or more memory devices including product data, be communicatively connected to one or more memory devices including product data, etc. For example, the content platform may maintain and update a database of information associated with a product, and changes made to the database may be reflected in a UI element presented to a user.

[0055] In some embodiments, aspects of the present disclosure enable a simplified purchasing process for a user. Upon receiving an indication of a user's intent to purchase a product (e.g., upon user interaction with a UI element associated with a product), a UI element for facilitating the purchase of the product may be presented to the user. In some embodiments, the user may purchase the product via an application associated with a content item, a content providing platform, etc. In some embodiments, the user may be directed to one or more external merchants (e.g., merchant applications, merchant web pages, etc.) to purchase the product.

[0056] In some embodiments, UI elements associated with products may be provided in response to various operations of an application (e.g., an application associated with a content platform). UI elements associated with products may be provided as part of a list of content items customized for a user, e.g., customized for a user account associated with a user. UI elements associated with products may be provided as part of a list of content items related to previously presented content items, e.g., a list of watched next, a list of recommended videos, etc. UI elements associated with products may be provided as part of a list of content items generated in response to a user search. UI elements associated with products may be provided as part of a product-specific list, e.g., a shopping area of ​​an application associated with a content platform. The inclusion of UI elements associated with products, content items associated with products, the number of UI elements or content items associated with products presented, etc. may be based on multiple indicators. Indicators may include user viewing history, user search history, etc.

[0057] In some embodiments, aspects of the present disclosure can enable quick access to one or more parts of content items associated with products. For example, a UI element can include a list of products associated with a content item. One or more products in the list of products can be associated with a part of the content item, for example, a timestamp of a video content item. After interacting with the product-related part of the UI element, the content item can present the relevant part of the content item (for example, a video can start playing the part associated with the content item of the video). In some embodiments, the list of products associated with a content item can be updated when the content item is presented. For example, a list of products associated with a video can be rearranged when the video is played. For example, the product currently highlighted by the video can be at the top of the list of products, and the product currently on the screen can be grouped together in the product presentation UI element, etc.

[0058] Aspects of the present disclosure may provide technical advantages over previous solutions. Aspects of the present disclosure may enable automatic product detection in content items. This may improve the experience of content creators by automatically associating one or more products with content items (e.g., removing the burden of associating products with content items from creators). Automatically generated product associations may have improved accuracy due to the use of multiple sources (e.g., text and image searches for products), fusion models, etc. Model-based product detection may be utilized to improve content items, associated data, etc., for example, machine-generated commentary or descriptions of content items may be improved using object detection. More accurate commentary may improve the user's experience when consuming content items. These improvements may reduce the time required for content creators to generate accurate content, may reduce the time spent by users to find content items and / or products of interest to users, may increase the accuracy of machine-generated information associated with content items, etc. Therefore, computing resources at client devices associated with content creators, content viewers, and / or platforms are reduced and may be used for other processes, which increases overall efficiency and reduces the overall latency of the system.

[0059] Aspects of the present disclosure can improve the user's experience when searching for content items, viewing content items, scrolling through a list of content items, etc. A user may be able to identify the content item from a list of content items, for example, without being presented with a content item with an associated product. A seamless method can be provided to the user to increase interaction with the product associated with the content item, for example, interaction with a UI element indicating the presence of a product associated with the content item may cause the presentation of a content item with further information about one or more products, and further interaction may facilitate the purchase of one or more products, etc. The presentation of one or more UI elements can streamline the user's experience. The user may be able to easily retrieve additional information, for example, within an application associated with a content platform / content item. The user may be able to more easily purchase products associated with a content item. The user may be directed to the relevant part of the content item based on the expressed interest in the product associated with the content item. The user may be able to easily view price information, availability information, product changes, related products, etc. within the context of a single application. Such implementations can save user time and energy, simplify the shopping and / or purchasing process, and simplify product research, evaluation, and / or selection processes, etc.

[0060] Figure 1 An example system architecture 100 for providing content and associated product information according to some embodiments is shown. The system architecture 100 includes a client device 110, one or more networks 105, a content platform system 102, and a product identification system 175. The content platform system 102 includes one or more server machines 106, one or more data repositories 140, and may include various platforms designed to perform tasks (e.g., tasks associated with content delivery). The platform of the content platform system 102 may be hosted by one or more server machines 106. The platform of the content platform system 102 may include and / or be hosted on one or more computing devices (such as rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.) and one or more data repositories (e.g., hard disks, memories, and databases), and may be coupled to one or more networks 105. In some embodiments, components of the content platform system 102 (e.g., server machines 106, data repositories 140, hardware associated with one or more platforms, etc.) may be directly connected to one or more networks 105. In some embodiments, one or more components of the content platform system 102 may access the network 105 via another device, such as a hub, a switch, etc. In some embodiments, one or more components of the content platform system 102 may directly communicate with the network 105. Figure 1106, including external data stores, etc. The platforms of the content platform system 102 may include an advertising platform 165, a social networking platform 160, a recommendation platform 157, a search platform 145, and a content providing platform 120. The product identification system 175 includes server machines 170, server machines 180, and a set of models 190. The product identification system 175 may include further devices such as data repositories, further servers, etc. Multiple operations of the product identification system 175 may be performed by a single physical or virtual device.

[0061] The one or more networks 105 may include one or more public networks (e.g., the Internet), one or more private networks (e.g., a local area network (LAN), a wide area network (WAN), one or more wired networks (e.g., an Ethernet network), one or more wireless networks (e.g., an 802.11 network), one or more cellular networks (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or combinations thereof. In one implementation, some components of the architecture 100 are not directly connected to each other. In one implementation, the system architecture 100 includes separate networks 105.

[0062] One or more data repositories 140 may reside in a memory (e.g., random access memory), a cache, a drive (e.g., a hard drive), a flash drive, etc., and may be part of one or more database systems, one or more file systems, or another type of component or device capable of storing data. One or more data repositories 140 may include multiple storage components (e.g., multiple drives or multiple databases) that may also span multiple computing devices (e.g., multiple server computers). A data repository may be a persistent storage capable of storing data. The persistent storage may be a local storage unit or a remote storage unit, an electronic storage unit (e.g., a main memory), or a similar storage unit. The persistent storage may be a monolithic device or a distributed set of devices.

[0063] Content items 121A-121C (e.g., media content items) can be stored on one or more data repositories. A data repository can be part of one or more platforms. Examples of content items 121 can include, but are not limited to, digital videos, digital movies, animated images, digital photos, digital music, digital audio, digital video games, collaborative media content presentations, website content, social media updates, e-books, e-journals, digital audio books, web blogs, software applications, etc. Content items 121A-121C can also be referred to as media items. Content items 121A-121C can be pre-recorded or live. For the sake of brevity and simplicity, videos can be used throughout this document as examples of content items 121 (e.g., content items 121A). Videos can include pre-recorded videos, live videos, short videos, etc.

[0064] Content items 121A-121C may be provided by a content provider. A content provider may be a user, a company, an organization, etc. A content provider may provide content item 121 (e.g., content item 121A) as a video. A content provider may provide content item 121 including live content, for example, content item 121 may include live video, real-time chat associated with the video, etc.

[0065] Client devices 110 may include devices such as televisions, smart phones, personal digital assistants, portable media players, laptop computers, e-book readers, tablet computers, desktop computers, game consoles, set-top boxes, and the like.

[0066] The client device 110 may include a communication application 115. A user may consume content items 121 (e.g., content item 121A) via the communication application 115. For example, the communication application 115 may access one or more networks 105 (e.g., the Internet) via the hardware of the client device 110 to provide the content items 121 to the user. As used herein, "media," "media items," "online media items," "digital media," "digital media items," "content," "media content items," and "content items" may include electronic files that may be executed or loaded using software, firmware, and / or hardware configured to present content items. In one implementation, the communication application 115 may be an application that allows a user to compose, send, and receive content items 121 (e.g., videos) through a platform (e.g., a content providing platform 120, a recommendation platform 157, a social networking platform 160, and / or a search platform 145) and / or a combination of platforms and / or networks.

[0067] In some embodiments, the communication application 115 may be a social networking application, a video sharing application, a video streaming application, a video game streaming application, a photo sharing application, a chat application, or a combination of such applications (or including aspects thereof). The communication application 115 associated with the client device 110 may render, display, present, and / or play one or more content items 121 to one or more users. For example, the communication application 115 may provide a user interface 116 (e.g., a graphical user interface) to be displayed on the endpoint device 110 for receiving and / or playing video content. In some embodiments, the communication application 115 is associated with (and managed by) the content platform 102 or the content providing platform 120.

[0068] In some embodiments, the communication application 115 may include a content viewer 113 and a related product component 114. A user interface 116 (UI) may display the content viewer 113 and the related product component 114. The related product component 114 may be used to display a UI element that displays information about one or more products (e.g., one or more products associated with a content item). The related product component 114 may display a UI element for notifying a user that a content item has one or more associated products (e.g., the related product component 114 may cause a display of a "shopping" symbol displayed near the content viewer 113, the related product component 114 may cause an element to be displayed near an element for selecting a content item to be presented from a list of content items and indicating that the content item has one or more associated products, etc.). The related product component 114 may cause the UI to display information about one or more products associated with a content item (e.g., the related product component 114 may cause a UI to be displayed that may include a list of products associated with a content item, may include images of products associated with a content item, may include price information or information connecting a product to a content item, such as a timestamp of a portion of a video associated with a product, etc.). Related products component 114 may cause the UI element to display additional information about one or more products, such as variations of the product (e.g., color variations, size variations, etc.), related products, recommended products, etc. Related products component 114 may cause the UI element to display one or more options for purchasing one or more products. In some embodiments, the user may be able to navigate between different views provided by related products component 114. FIG. 4A to FIG. 4EFind further description of example UI elements associated with products related to content items. In some embodiments, for example, if more than one content item with associated products is displayed, the related product component 114 may cause more than one UI element to be displayed. In some embodiments, the related product component 114 may cause the UI element indicating the related product to not be displayed, for example, based on user settings, user preferences, user history, etc. In some embodiments, multiple content viewers 113 and / or related product components 114 may be associated with a user interface, communication application, client device, etc. For example, multiple content items may be displayed for a user at one time. In some implementations, the communication application 115 is a web browser that can access, retrieve, present and / or navigate content (e.g., web pages, such as hypertext markup language (HTML) pages, digital media items, etc.), and may include a related product component 114 and a content viewer 113, which may be an embedded media player embedded in a user interface 116 (e.g., a web page associated with viewing content) provided by a content providing platform 120. Alternatively, the application 115 is not a web browser, but a standalone application (e.g., mobile application, desktop application, game console application, TV application, etc.) downloaded from a platform (e.g., content providing platform 120, recommendation platform 157, social networking platform 160, or search platform 145) or pre-installed on the client device 110. The standalone application 115 can provide a user interface 116 including a content viewer 113 (e.g., an embedded media player) and related product components 114.

[0069] In some embodiments, the content platform system 102 may include a product information platform 161 (e.g., hosted by the server machine 106). The product information platform 161 may store, retrieve, provide, receive, etc. data related to one or more products associated with one or more content items. The content platform system 102 may provide data to the related product component 114 of the client device 110. The product information platform 161 may include information provided by a content creator, for example, a content creator may provide a list of products associated with a content item provided by the content creator to the content platform system 102. The product information platform 161 may include information provided by one or more users, for example, one or more users may identify products associated with a content item, for example, in response to having the content item presented to them. The product information platform 161 may include information provided by a product identification system 175, for example, one or more machine learning models may be utilized to identify products displayed in a content item and provide an indication of the associated products and content item to the product information platform 161.

[0070] In some embodiments, the communication application 115 installed on the client device 110 can be associated with a user account, for example, the user can log into the account on the client device 110. In some embodiments, multiple client devices 110 can be associated with the same client account. In some embodiments, providing information about the association of a product with one or more content items can be performed based on the user account, for example, account settings, account history (e.g., history of interaction with UI elements including associated product information), etc.

[0071] In some embodiments, the client device 110 may include one or more data repositories. The data repository may include commands (e.g., instructions that cause operations when executed by a processing device) for rendering a UI (e.g., user interface 116). The instructions may include instructions for rendering interactive components, such as UI elements with which a user can interact to be presented with additional information about one or more products associated with a content item. In some embodiments, the instructions may cause the processing device to render a UI element that presents information about one or more products associated with one or more content items (e.g., multiple videos reviewing the product may be presented together with a UI element that presents more information about the product of interest to the user).

[0072] In some embodiments, one or more server machines 106 may include computing devices such as rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc., and may be coupled to one or more networks 105. Server machines 106 may be stand-alone devices or part of any of the platforms (e.g., content providing platform 120, social networking platform 160, etc.).

[0073] The social networking platform 160 may provide an online social networking service. The social networking platform 160 may provide a communication application 115 for users to create profiles and perform activities using their profiles. Activities may include updating profiles, exchanging messages with other users, evaluating (e.g., liking, commenting, sharing, recommending) status updates, photos, videos, etc., and receiving notifications associated with activities of other users. In some embodiments, a user may share additional product information (e.g., as provided by the product information platform 161) with one or more additional users via the social networking platform 160.

[0074] The recommendation platform 157 may be used to generate and provide content recommendations (e.g., articles, videos, posts, news, games, etc.). Recommendations may be based on search history, content consumption history, followed / subscribed channel content, linked profiles (e.g., friend lists), popular content, etc. The recommendation platform 157 may be utilized to generate, for example, a user home feed, a user watch list, a user play list, etc. One or more UI elements indicating associated products, displaying a list of associated products, presenting product information, presenting one or more options for purchasing a product, etc. may be presented in conjunction with, as part of, in association with, accessible from, etc., a home feed, a watch list, a play list, a watch next list, etc. The presentation of the one or more UI elements may be performed based on user history, user settings, data indicating the user's preferences (e.g., demographic data), etc.

[0075] The search platform 145 can be used to allow users to query one or more data repositories 140 and / or one or more platforms and receive query results. The user can use the search platform 145 to search for content items, search topics, etc. For example, the user can use the search platform 145 to search for content items with one or more associated products. The search platform 145 can be used to search for content items related to a certain type of product (e.g., headphone review videos). The search platform 145 can be used to search for content items related to a certain product (e.g., a specific brand and / or model of headphones). In response to receiving a search query, one or more UI elements can be displayed to the user. The type, style, etc. of the displayed UI elements can be based on the content of the user's search. For example, in response to a search for content related to a certain type of product (e.g., headphones), a UI element can be displayed to indicate that the content item suggested in view of the search has one or more associated products. As a further example, in response to a search for content items related to a more specific product (e.g., "best headphones for podcasting"), different UI elements can be displayed to provide information about products associated with the content items. As a further example, in response to a search for content items related to a particular product (e.g., a particular brand and / or model), different UI elements may be displayed to provide specific information about the searched product, thereby indicating that the product is associated with content items that are recommended in light of the search.

[0076] The content providing platform 120 may be used to provide access to content items 121 for one or more users and / or provide content items 121 to one or more users. For example, the content providing platform 120 may allow users to consume, upload, download, and / or search for content items 121. In another example, the content providing platform 120 may allow users to evaluate content items 121, such as to approve (“like”), disapprove, recommend, share, rate, and / or comment on content items 121. In another example, the content providing platform 120 may allow users to edit content items 121. The content providing platform 120 may also include a website (e.g., one or more web pages) and / or one or more applications (e.g., communication applications 115) that may be used to provide access to content items 121 for one or more users. For example, the client device 110 may use the communication application 115 to access the content items 121. The content providing platform 120 may include any type of content delivery network that provides access to the content items 121.

[0077] The content providing platform 120 may include a plurality of channels (e.g., channel A 125, channel B 126, etc.). A channel may be a collection of content available from a common source, a collection of content having a common topic or theme, etc. The data content may be digital content selected by a user, digital content provided by a user, digital content uploaded by a user, digital content selected by a content provider, digital content selected by a broadcaster, etc. For example, channel A 125 may include two videos (e.g., content items 121A-121B). A channel may be associated with an owner, who may be a user who can perform actions on the channel. The content may be one or more content items 121. The data content of a channel may be pre-recorded content, live content, etc. Although a channel is described as one implementation of a content providing platform, implementations of the present disclosure are not limited to content sharing platforms that provide content items 121 through a channel model.

[0078] Product identification system 175, server machine 170, and server machine 180 may each include one or more computing devices, such as a rack server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc. The operations of predictive server 112, server machine 170, server machine 180, data repository 140, etc. may be performed by a cloud computing service, a cloud data storage service, etc.

[0079] Product identification system 175 may include one or more models 190. Models 190 included in product identification system 175 may perform tasks related to identifying one or more products from content items. One or more of models 190 may be trained machine learning models. Figure 3A and Figure 5C Describes the operations for producing a trained machine learning model, which includes training, validating, and testing the model.

[0080] Model 190 may include one or more text parsing models 191. Text parsing model 191 may be configured to receive text as input and generate one or more indications of products associated with the text as output. For example, a first model in text parsing model 191 may be configured to predict associated products from the title of a content item, a second model in text parsing model 191 may be configured to predict associated products from a (e.g., written) description of a content item, a third model may be configured to predict associated products from a description of a content item (e.g., an automatically generated description, a machine generated description, a user provided description, etc.), etc. In some embodiments, all operations in text parsing model 191 may be performed by a single model. In some embodiments, a model in text parsing model 191 may be configured to generate product context information, e.g., information indicating that a content item is associated with one or more products, as output. For example, product context information may indicate that a content item includes a category of products, e.g., a group of different products (e.g., a type of product, a trademark of a product, a category of a product, such as "electronic devices", etc.).

[0081] Model 190 may include one or more image parsing models 192. Image parsing model 192 may be configured to identify products from one or more images. Image parsing model 192 may include one or more models for identifying that an image includes a product, a model configured to isolate a product image from a content item image (e.g., remove background elements, etc.), a model configured to determine the identity of a product in a content item image, etc. The operation of image parsing model 192 may be performed by a single model. Image parsing model 192 may include one or more models configured to provide an image to a product recognition model. For example, image parsing model 192 may include a model configured to extract a portion of a still image of a content item, a model configured to extract one or more frames of a video content item, etc. The model in image parsing model 192 may be provided with one or more frames of a video, one or more portions of one or more frames of a video, etc., and generate one or more products and one or more confidence values ​​associated with one or more products as output. For example, a model in image parsing model 192 may receive one or more frames of a video content item as input, and generate a list of products with confidence values ​​as output, the confidence value indicating the likelihood that the product is included in the image of the content item. Image parsing model 192 may include one or more models that determine which images to utilize from a content item. For example, image parsing model 192 may include one or more models that select frames for image recognition from a video content item.

[0082] The image parsing model 192 may include one or more models configured to reduce the dimensionality of an image. For example, an image may be simplified to a value vector. In some embodiments, one or more models in the image parsing model 192 may be configured to reduce the dimensionality of an image in such a manner that similar images (e.g., images of the same product) may be represented similarly (e.g., by similar vectors) after dimensionality reduction. One or more models in the image parsing model 192 may be configured to compare a reduced dimensional image from a content item (e.g., a value vector generated from one or more frames of a content item video) with a reduced dimensional image of a known product (e.g., via the product information platform 161).

[0083] Model 190 may include text correction model 193. Text correction model 193 may be configured to provide corrections to text associated with content items. Text correction model 193 may be configured to adjust text associated with content items to include one or more products referenced in the content items. One or more models in text correction model 193 may be configured to adjust computer-generated text, machine-generated text, automatically generated text, etc. associated with content items. One or more models in text correction model 193 may be configured to update the commentary of a video (e.g., incorrect commentary) to include one or more products. In some embodiments, the machine-generated text (e.g., commentary) associated with a content item may be incorrect. For example, when generating commentary, the name of a product may be replaced with an approximate term (e.g., the name of the product may not be a word in the language of the commentary, the name of the product may be a word in a language different from the language of the commentary, etc.). The model in text correction model 193 may be configured to identify the incorrect part in the text and recommend correction, perform correction, warn the user or another system, etc. For example, a model in text correction model 193 may receive a machine-generated transcript of a video, identify portions of the transcript that may incorrectly substitute words used for a product name in the language of the transcript, and provide data indicating potentially incorrect text to another model, system, user, etc.

[0084] Model 190 may include fusion model 194. Fusion model 194 may receive one or more indications of products associated with a content item as input. In some embodiments, fusion model 194 receives outputs from one or more other models (e.g., text parsing model 191, image parsing model 192, etc.) as input. Fusion model 194 may receive one or more indications of products associated with a content item and one or more indications of confidence values. For example, fusion model 194 may receive indications of one or more products detected in the title of a content item and confidence values ​​associated with the one or more products. Fusion model 194 may further receive indications of one or more products detected in the description of the content item and confidence values ​​associated with the one or more products. Fusion model 194 may further receive indications of one or more products detected in the caption of the content item and confidence values ​​associated with the one or more products. Fusion model 195 may further receive indications of one or more products detected in the image of the content item and confidence values ​​associated with the one or more products. Fusion model 195 may generate one or more products detected in association with the content item as output. Fusion model 195 may further generate a confidence value associated with the confidence that the product appears in the content item, that the product is associated with the content item, etc. Further operations may be performed based on the output of fusion model 195 (e.g., product identity information and confidence value) (e.g., a UI element describing the product associated with the content item may be presented).

[0085] One type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. An artificial neural network typically includes a feature representation component having a classifier or regression layer that maps features to a desired output space. A convolutional neural network (CNN), for example, hosts multiple convolutional filter layers. On top of the lower layers, where pooling is performed and nonlinear problems can be solved, a multi-layer perceptron is typically attached to map the top features extracted by the convolutional layers to a decision (e.g., a classification output).

[0086] Recurrent Neural Networks (RNNs) are another type of machine learning model. Recurrent Neural Network models are designed to interpret a series of inputs where the inputs are intrinsically related to each other, such as time tracking data, sequential data, etc. The output of the perceptron of the RNN is fed back into the perceptron as input to generate the next output.

[0087] Deep learning is a class of machine learning algorithms that use a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Deep neural networks include a hierarchy of layers, where different layers learn different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a slightly more abstract and complex representation. For example, in an image recognition application, the raw input may be a matrix of pixels; the first representation layer may abstract the pixels and encode the edges; the second layer may compile and encode the arrangement of the edges; the third layer may encode higher-level shapes (e.g., teeth, lips, gums, etc.); and the fourth layer may recognize the scanned character. Notably, the deep learning process can learn by itself which features are optimally placed at which level. The "depth" in "deep learning" refers to the number of layers through which the data is transformed. More precisely, deep learning systems have a considerable credit assignment path (CAP) depth. A CAP is a chain of transformations from an input to an output. A CAP describes the underlying causal connection between an input and an output. For feed-forward neural networks, the depth of CAP can be the depth of the network, and can be the number of hidden layers plus 1. For recurrent neural networks where signals can propagate through a layer more than once, the CAP depth is potentially unlimited.

[0088] In some embodiments, product identification system 175 further includes server machine 170 and server machine 180. Server machine 170 includes a dataset generator 172 that can generate a dataset (e.g., a set of data inputs and a set of target outputs) for training, validating, and / or testing a model 190 including one or more machine learning models. Figure 2 and Figure 5A Some operations of the dataset generator 172 are described in detail. In some embodiments, the dataset generator 172 may partition the historical data (e.g., pre-existing content item data, content items with one or more specified associated products, content items with product assignments provided by one or more users, etc.) into a training set (e.g., sixty percent of the historical data), a validation set (e.g., twenty percent of the historical data), and a test set (e.g., twenty percent of the historical data).

[0089] In some embodiments, components of product identification system 175 may generate multiple sets of features. For example, a feature may be a rearrangement of input data, a combination of input data, a dimensionality reduction of input data, a subgroup of input data, etc. One or more data sets may be generated based on one or more features of the input data.

[0090] The server machine 180 includes a training engine 182, a validation engine 184, a selection engine 185, and / or a test engine 186. An engine (e.g., a training engine 182, a validation engine 184, a selection engine 185, and a test engine 186) may refer to hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, a processing device, etc.), software (such as instructions running on a processing device, a general purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. The training engine 182 may be capable of training one or more models 190 using one or more sets of features associated with a training set from the data set generator 172. The training engine 182 may generate a plurality of trained models 190, where each trained model 190 corresponds to a distinct set of features of the training set. The dataset generator 172 can receive the output of a trained model (e.g., the fusion model 194 can be trained based on the output of the text parsing model 191 and / or the image parsing model 192), collect the data into training, validation, and testing datasets, and use the datasets to train a second model (e.g., the fusion model 194).

[0091] The validation engine 184 may be capable of validating the trained models 190 using a corresponding set of features of a validation set from the dataset generator 172. For example, a first trained machine learning model 190 that has been trained using a first set of features of a training set may be validated using a first set of features of a validation set. The validation engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of features of the validation set. The validation engine 184 may discard trained models 190 that have an accuracy that does not meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting one or more trained models 190 that have an accuracy that meets a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting a trained model 190 with the highest accuracy among the trained models 190.

[0092] The testing engine 186 may be capable of testing the trained model 190 using a corresponding set of features of the test set from the dataset generator 172. For example, a first trained machine learning model 190 that has been trained using a first set of features of the training set may be tested using a first set of features of the test set. The testing engine 186 may determine the trained model 190 with the highest accuracy among all trained models based on the test set.

[0093] In the case of a machine learning model, the model 190 may refer to a model artifact created by the training engine 182 using a training set that includes data inputs and corresponding target outputs (correct answers for the corresponding training inputs). Patterns in the data set that map the data inputs to the target outputs (correct answers) may be found, and mappings that capture these patterns are provided to the machine learning model 190. The machine learning model 190 may use one or more of a support vector machine (SVM), a radial basis function (RBF), clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, a k-nearest neighbor algorithm (k-NN), linear regression, a random forest, a decision forest, a neural network (e.g., an artificial neural network, a recurrent neural network), a linear model, a function-based model (e.g., an NG3 model), etc. The synthetic data generator 174 may include one or more machine learning models, which may include one or more of the same type of models (e.g., an artificial neural network).

[0094] Automatic (e.g., model-based) detection of products from content items and associated data provides significant technical advantages over other methods. In some embodiments, the content item (e.g., product review video for reviewing products) showing products can become linked or associated with the product without the attention, action, time, etc. of the content creator (e.g., data linking the product to the content item can be generated). In some embodiments, the content item (e.g., content item can be sponsored or can promote one or more products) of the advertising product can be linked or associated with the product. Model-based detection of products in content items can generate product associations for products that are not particularly displayed in the content items but exist in the content items (e.g., products that users may be interested in purchasing can be in video content items on the screen). Model-based detection of products in content items can generate product associations for products advertised in content items. By providing (e.g., indicating that the product is associated with the content item) UI elements to the user, the user can be guided to the products existing in the content items based on model-based detection.

[0095] One or more models 190 may be run on inputs to generate one or more outputs. The model may determine (e.g., extract) confidence data from the outputs, the confidence data indicating a level of confidence that the output of the model is an accurate depiction of the content item. For example, the model may determine that a first product is associated with a content item, and determine the confidence that the model has correctly discovered the first product in the content item. One or more components of the product identification system 175 may use the confidence data to decide whether to update data associated with the content item, e.g., whether to associate one or more products with the content item, whether to update one or more narrations for the content item, etc.

[0096] The confidence data may include or indicate a confidence level that the output of the model (e.g., one or more products) is an accurate indication of the product associated with the content item. For example, the confidence level output by the model (e.g., associated with a product identified in the content item) may be a real number between 0 and 1 (including 0 and 1). 0 may indicate that the confidence level that the predicted product is associated with the content item is zero, and 1 may indicate an absolute confidence level that the predicted product is associated with the content item. In response to the confidence data indicating that the confidence level is below a threshold level for a predetermined number of instances (e.g., a percentage of instances, a frequency of instances, a total number of instances, etc.), the product identification system 175 may cause one or more trained models 190 (e.g., based on updated and / or new data for training, validation, testing, etc.) to be retrained. Retraining may include generating one or more data sets (e.g., via the data set generator 172).

[0097] For purposes of illustration and not limitation, aspects of the present disclosure describe the training of one or more machine learning models 190 using historical data, and inputting current data (e.g., newly updated content items, content items not previously associated with products, etc.) into one or more trained machine learning models to determine outputs indicating content item-product associations. In other embodiments, heuristic models, physics-based models, or rule-based models are used to determine that one or more products are associated with content items (e.g., without using trained machine learning models). In some embodiments, such models can be trained using historical data. In some embodiments, these models can be retrained using historical data. Information about content items can be monitored or otherwise used in heuristic, physics-based, or rule-based models. Figure 2 The data input 210 describes any information.

[0098] In some embodiments, the functionality of client device 110, product identification system 175, content platform system 102, server machine 170, and server machine 180, server machine 106 may be provided by a smaller number of machines. For example, in some embodiments, server machines 170 and 180 may be integrated into a single machine, while in some other embodiments, server machines 170, server machines 180, and server machine 106 may be integrated into a single machine. In some embodiments, client device 110 and server machine 106 may be integrated into a single machine. In some embodiments, the functionality of client device 110, server machine 106, server machine 170, server machine 180, and data repository 140 may be performed by a cloud-based service.

[0099] In general, functions described in one embodiment as being performed by client device 110, server machine 106, server machine 170, and server machine 180 may also be performed on server machine 106 in other embodiments, if appropriate. In addition, functionality attributed to a particular component may be performed by different or multiple components operating together. For example, in some embodiments, product identification system 175 may determine an association between a product and a content item. In another example, content platform system 102 may determine an association between a content item and one or more products.

[0100] Additionally, the functionality of a particular component may be performed by different or multiple components operating together.One or more of server machine 106, server machine 170, or server machine 180 may be accessed as a service provided to other systems or devices through an appropriate application programming interface (API).

[0101] In the implementation of the present disclosure, a "user" can be represented as a single individual. However, other implementations of the present disclosure cover "users" as entities controlled by a group of users and / or automated sources. For example, a group of individual users united into a community in a social network can be regarded as "users". In another example, an automated consumer can be an automated ingestion pipeline of one or more platforms, one or more content items, etc., such as a topic channel. In addition to the above description, controls can also be provided to users, which allow users to select whether and when the system, program or feature described herein can realize the collection of user information (e.g., information about the user's social network, social actions or activities, occupation, user's preferences or the user's current location) and whether to send content or communication from the server to the user. In addition, before storing or using certain data, the data can be processed in one or more ways so that personally identifiable information is removed. For example, the identity of the user can be processed so that any personally identifiable information of the user cannot be determined, or in the case of obtaining location information, the user's geographic location can be generalized (such as generalized to the city, zip code or state level) so that the specific location of the user cannot be determined. Therefore, the user can control what information is collected about the user, how the information is used, and what information is provided to the user.

[0102] Figure 2is a block diagram of a system 200 including a dataset generator 272 for creating datasets for one or more models according to some embodiments. The dataset generator 272 can use historical data to create a dataset (e.g., data input 210, target output 220). An unsupervised machine learning model can be trained using a dataset generator similar to the dataset generator 272, e.g., the target output 220 may not be generated by the dataset generator 272. A semi-supervised machine learning model can be trained using a dataset generator similar to the dataset generator 272, e.g., the target output 220 corresponding to a subset of the data input 210 may be generated by the dataset generator 272.

[0103] Dataset generator 272 can generate data sets for training, testing and validating models. Dataset generator 272 can generate data sets for machine learning models. System 200 can generate data sets for training, testing and / or validating fusion models, which are used, for example, to determine the likelihood that one or more products appear in a content item. Systems similar to system 200 can generate data sets for training, testing and / or validating models with different functions, with corresponding changes to input data and / or output data included in the data sets. Models for parsing text (e.g., extracting one or more references to a product from text associated with a content item), parsing images (e.g., extracting one or more references to a product from an image associated with a content item), correcting text (e.g., including one or more references to a product in a machine-generated text associated with a content item), etc. can have data sets generated by a data set generator similar to data set generator 272 for training, testing and / or validating models.

[0104] In some embodiments, a dataset generator, such as dataset generator 272, can be associated with two or more separate models (e.g., a dataset can be used to train an integrated model). For example, an input dataset can be provided to a first model, an output of the first model can be provided to a second model, and a target output can be provided to the second model to train, test, and / or validate the first and second models (e.g., an integrated model).

[0105] The data set generator 272 can generate one or more data sets for providing to the model, for example, during training, validation, and / or testing operations. Multiple sets of historical data can be provided to the machine learning model. Multiple sets of historical text parsing data 264A-264Z can be provided as data input to the machine learning model (e.g., a fusion model). The text parsing data can be provided by the machine learning model, and the text parsing data can, for example, include one or more products and confidence values ​​identified by the machine learning model in the text associated with the content item. Multiple sets of historical image parsing data 265A-265Z can be provided as data input to the machine learning model. Image parsing data can be provided by a trained machine learning model, and the image parsing data can, for example, include one or more products and confidence values ​​identified by the machine learning model in one or more images associated with the content item.

[0106] In some embodiments, dataset generator 272 may be configured to generate datasets for training, testing, validating, etc. fusion models. A dataset generator similar to dataset generator 272 may generate multiple sets of text data (e.g., content item title text, content item description text, content item commentary text, etc.) as data input to train a machine learning model for determining one or more products associated with a content item. A dataset generator similar to dataset generator 272 may generate multiple sets of image data (e.g., one or more frames from a video content item, portions of one or more frames from a video content item, etc.) as data input to train a machine learning model for determining one or more products associated with a content item.

[0107] In some embodiments, the data set generator 272 generates a data set (e.g., a training set, a validation set, a test set) including one or more data inputs 210 (e.g., a training input, a validation input, a test input). The data input 210 may be provided to Figure 1 The training engine 182, validation engine 184, or test engine 186 of the present invention may be used to train, validate, or test a model (e.g., a fusion model, a text parsing model, an image parsing model, etc.). The dataset generator 272 may generate a dataset (e.g., a training set, a validation set, a test set) including one or more data inputs 210. The data inputs 210 may be referred to as "features," "attributes," "vectors," or "information."

[0108] In some embodiments, the dataset generator 272 may generate a first data input corresponding to the first set of historical text parsing data 264A and / or the first set of historical image parsing data 265A to train, validate, or test the first machine learning model. The dataset generator 272 may generate a second data input corresponding to the second set of historical metrology data 264B and / or the second set of design rule data 265B to train, validate, or test the second machine learning model. Figure 5A Some embodiments of generating training sets, test sets, validation sets, etc. are further described.

[0109] In some embodiments, the data set generator 272 may generate a target output 220 to be provided to train, test, validate, etc. one or more machine learning models. The data set generator 272 may generate product association data 268 as the target output 220. The product association data 268 may include identifiers of one or more products associated with the content item (e.g., manually labeled product associations). The product association data 268 may include an input-output mapping, for example, a set of historical text parsing data 264A may be associated with a first set of product association data 268, etc. The machine learning model may be updated (e.g., trained) by providing input data, generating output, and comparing the output with the provided target output (e.g., "correct answer"). The various weights, biases, etc. of the model are then updated to better align the model with the training data. This process may be repeated many times to generate a model that provides an accurate output for a threshold portion of the provided input. The target output 220 may share one or more features of the data input 210, for example, the target output 220 may be organized into attributes or vectors, the target output 220 may be organized into groups AZ, etc.

[0110] In some embodiments, a dataset generator similar to dataset generator 272 can be utilized in conjunction with a text parsing model, configured to determine one or more product associations with a content item. Product associations can include contextual associations, such as a product's trademark, a product's type, a product's classification, etc. The dataset generator can generate a list of products associated with the content item, a product's type, a product's trademark, etc. associated with the text of the content item as a target output. A dataset generator similar to dataset generator 272 can be utilized in conjunction with an image parsing model, configured to determine one or more product associations with the content item. The dataset generator can generate a list of products associated with the image input, a product's classification, a product's category, etc. as a target output. A dataset generator similar to dataset generator 272 can be utilized in conjunction with a text correction model. The text correction model can be configured to identify machine-generated text that incorrectly provides one or more words in the target language to replace a product name. One or more sets of machine-generated text can be provided as input data for the text correction model, and products associated with the text (e.g., products that are not correctly captured by the machine generation of the text) are provided as target outputs.

[0111] In some embodiments, after a data set is generated and used to train, validate, or test a machine learning model, the model may be further trained, validated, or tested or adjusted (e.g., adjusting weights or parameters associated with the input data of the model, such as connection weights in a neural network). The model may be adjusted and / or retrained based on data different from the original training operation, such as data generated after training, validation, and / or testing of the model.

[0112] Figure 3A is a block diagram illustrating a system 300A for generating output data (e.g., product / content item association data) according to some embodiments. System 300A can be used in conjunction with a fusion model to generate product / content item association data and confidence data based on potential product / content item associations generated by other models (e.g., a text-based model that detects products in text associated with a content item, an image-based model that detects products in images associated with a content item, etc.). Systems similar to system 300A can be utilized to generate output data from other types of models, such as text parsing models, image parsing models, text correction models, etc.

[0113] At block 310, system 300A (e.g., Figure 1 The product identification system 175 of the present invention may be used to train, validate and / or test the machine learning model (e.g., via Figure 2Data partitioning may be performed by a data set generator 272 of the system 300A. In some embodiments, the training data 364 includes historical data, such as historical associations between products and text-based content items, historical associations between products and image-based content items, and the like. In some embodiments, for example, when the system 300A is directed to generating outputs from a fusion model, the training data 364 may include data generated by one or more trained machine learning models, for example, a model configured to detect products in text or images associated with content items. The training data 364 may undergo data partitioning at box 310 to generate a training set 302, a validation set 304, and a test set 306. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data.

[0114] The generation of training set 302, validation set 304 and test set 306 can be customized for a specific application. For example, the training set can be 60% of the training data, the validation set can be 20% of the training data, and the test set can be 20% of the training data. System 300A can generate multiple sets of features for each of the training set, validation set and test set. For example, if the training data 364 includes product associations extracted from text data from more than one text source (e.g., a title associated with a content item and a description associated with the content item), the input training data can be divided into a first set of features including products identified from text from a first source, and a second set of features including products identified from text from a second source. The target input can be divided into groups, the target output can be divided into groups, both can be divided into groups, or neither can be divided into groups. Multiple models can be trained on different sets of data.

[0115] At block 312, system 300A uses training set 302 (e.g., via Figure 1 The training engine 182 of the training engine 182 is used to perform model training. The training of the machine learning model can be implemented in a supervised learning manner, which involves: providing a training data set including labeled inputs through the model, observing its output, defining errors (by measuring the difference between the output and the labeled value), and using techniques such as deep gradient descent and back propagation to tune the weights of the model so that the error is minimized. In many applications, repeating this process across many labeled inputs in the training data set produces a model that can produce correct outputs when presented with inputs that are different from the inputs present in the training data set. In some embodiments, the training of the machine learning model can be implemented in an unsupervised manner, for example, no labels or classifications may be supplied during training. Unsupervised models can be configured to perform anomaly detection, result clustering, etc.

[0116] For each training data item in the training data set, the training data item can be input into the model (e.g., into a machine learning model). The model can then process the input training data item (e.g., an indication of one or more products detected in conjunction with a content item, and an associated confidence value, etc.) to generate an output. The output can include, for example, a list of products that can be associated with the content item and the corresponding confidence values. The output can be compared to the labeling of the training data item (e.g., a manually labeled set of products associated with the content item).

[0117] The processing logic may then compare the generated output (e.g., predicted product / content item associations) to the labels already included in the training data item (e.g., a manually generated list of product / content item associations). The processing logic determines an error (i.e., a classification error) based on the difference between the output and the labels. The processing logic adjusts one or more weights and / or values ​​of the model based on the error.

[0118] In the case of training a neural network, an error term or increment may be determined for each node in the artificial neural network. Based on this error, the artificial neural network adjusts one or more of its parameters (the weights of one or more inputs of the node) for one or more of its nodes. The parameters may be updated in a back-propagation manner so that the nodes of the highest layer are updated first, followed by the nodes at the next layer, and so on. The artificial neural network contains multiple layers of "neurons", each of which receives as input the values ​​of the neurons from the previous layer. The parameters of each neuron include weights associated with the values ​​received from each neuron in the previous layer of neurons. Therefore, adjusting the parameters may include adjusting the weights of each of the inputs of one or more neurons assigned to one or more layers in the artificial neural network.

[0119] System 300A may train multiple models using multiple sets of features of training set 302 (e.g., a first set of features of training set 302, a second set of features of training set 302, etc.). For example, system 300A may train a model using a first set of features in a training set (e.g., a subset of training data 364, such as data associated only with a subset of models configured to generate product / content item associations, etc.) to generate a first trained model, and train a model using a second set of features in a training set to generate a second trained model. In some embodiments, the first trained model and the second trained model may be combined to generate a third trained model (e.g., the third trained model may be better than the first or second trained models alone). In some embodiments, the multiple sets of features used when comparing models may overlap (e.g., a first set of features is a product based on a content item title, description, and some images, and a second set of features products are detected based on a content item description, a different set of images of the content item, and a detected context of the content item (e.g., the type of product associated with the content item)). In some embodiments, hundreds of models may be generated, including models with various feature arrangements and combinations of models.

[0120] At block 314, system 300A uses validation set 304 (e.g., via Figure 1 The validation engine 184 of the validation set 304) is used to perform model validation. The system 300A can use a corresponding set of features of the validation set 304 to validate each of the trained models. For example, the system 300A can use a first set of features in the validation set to validate a first trained model, and use a second set of features in the validation set to validate a second trained model. In some embodiments, the system 300A can validate hundreds of models generated at block 312 (e.g., models with various feature arrangements, combinations of models, etc.). At block 314, the system 300A can determine (e.g., via model validation) the accuracy of each of the one or more trained models, and can determine whether one or more of the trained models have an accuracy that meets a threshold accuracy. In response to determining that none of the trained models have an accuracy that meets the threshold accuracy, the process returns to block 312, where the system 300A performs model training using a different set of features of the training set, an updated or expanded training set provided by the data set generator, etc. In response to determining that one or more of the trained models have an accuracy that meets the threshold accuracy, the process continues to block 316. The system 300A may discard trained models having an accuracy below a threshold accuracy (eg, based on a validation set).

[0121] At block 316, system 300A (e.g., via Figure 1The selection engine 185 of the system 300A performs model selection to determine which of the one or more trained models that meet the threshold accuracy has the highest accuracy (e.g., the selected model 308 based on the verification of block 314). In response to determining that two or more of the trained models that meet the threshold accuracy have the same accuracy, the process can return to block 312, where the system 300A performs model training using a further refined training set corresponding to the further refined multiple groups of features to determine the trained model with the highest accuracy.

[0122] At block 318, system 300A uses test set 306 (e.g., via Figure 1 The system 300A may use a test engine 186 of the system 300A to perform model testing to test the selected model 308. The system 300A may test the first trained model using the first set of features in the test set to determine that the first trained model meets the threshold accuracy (e.g., based on the first set of features of the test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 is overfit to the training set 302 and / or the validation set 304 and is not applicable to other data sets such as the test set 306), the process continues to block 312, where the system 300A performs model training (e.g., retraining) using a different training set corresponding to a different set of features, different content items, etc. In response to determining that the selected model 308 has an accuracy that meets the threshold accuracy based on the test set 306, the process continues to block 320. In at least block 312, the model may learn patterns in the training data to make predictions, and in block 318, the system 300A may apply the model to the remaining data (e.g., the test set 306) to test the predictions.

[0123] At block 320, the system 300A uses a trained model (e.g., the selected model 308) to receive current data 322 (e.g., newly uploaded content items, newly created content items, content items not included in the training, testing, or validation set of the selected model 308, etc.) and determines (e.g., extracts) output data 324 (e.g., product / content item associations and corresponding confidence values) from the output of the trained model. Corrective actions associated with the content item and / or associated data may be performed in view of the output data 324. For example, the instructions may be updated to include presenting, along with the content item, a UI element specifying that the content item includes one or more associated products, the instructions may be updated to include presenting, along with the content item, a UI element containing additional information about the associated products, etc. The instructions may depend on additional factors of the content item, such as user preferences, presentation context (e.g., search page, home page, etc.), etc. In some embodiments, the current data 322 may correspond to the same type of features in the historical data used to train the machine learning model. In some embodiments, current data 322 corresponds to a subset of the types of features in the historical data used to train the selected model 308 (e.g., a machine learning model may be trained using product association and / or contextual information and confidence values ​​from multiple sources such as text- and image-based sources, and be provided with a subset of that data as current data 322).

[0124] In some embodiments, the performance of a machine learning model (e.g., selected model 308) can be adjusted, improved, and / or updated over time. For example, additional training data can be provided to the model to improve the model's ability to correctly classify product associations with content items. In some embodiments, a portion of current data 322 can be provided (e.g., via Figure 1 The selected model 308 may be updated and / or improved periodically, continuously, etc. using the current data 322 and the current target output data 346.

[0125] In some embodiments, one or more of actions 310-320 may occur in various orders and / or with other actions not presented and described herein. In some embodiments, one or more of actions 310-320 may not be performed. For example, in some embodiments, one or more of data partitioning of block 310, model validation of block 314, model selection of block 316, or model testing of block 318 may not be performed.

[0126] System 300A has been described with respect to a fusion model. The fusion model accepts one or more indications of products detected in association with a content item (e.g., from text associated with the content item, from an image associated with the content item, etc.) and a confidence value (e.g., confidence that the product is indeed referenced in the content item), and generates as an output an overall likelihood that the product is referenced in the content item based on multiple inputs. Systems similar to system 300A may be utilized to perform other machine learning-based tasks, for example, a text or image parsing model, the output of which is provided as input to the fusion model, may operate in a similar manner as described in conjunction with system 300A with appropriate substitutions for input data, output data, etc.

[0127] Figure 3B 300B is a block diagram of an example system 300B for generating an association between a content item and one or more products according to some embodiments. The system 300B may include multiple modules, such as image recognition 330, image validation 340, text recognition 350, fusion 360, etc. In some embodiments, multiple modules of the system 300B may operate together to identify a product from a content item. For example, a product detected in an image of a content item and metadata of the content item (e.g., a title, description, commentary, etc.) may be provided to a fusion model to determine the likelihood that one or more products are associated with the content item based on multiple input channels. An action may be taken in response to the likelihood that a product appears in a content item as determined by the fusion model (e.g., updating the metadata of the content item to include an indication of the product).

[0128] The image recognition module 330 can be used to identify one or more products from the images associated with the model, for example, an image from the content item can be compared with images in a database of products (e.g., thousands of products) to identify products present in the video. The products present can include a specific theme of the content item (e.g., a product reviewed in the content item), a product included in the content item (e.g., a product that appears by chance, a product that appears without being particularly highlighted, etc.), etc. The text recognition module 350 can identify one or more products associated with the content item from metadata / text data associated with the content item (e.g., from text including the title of the content item, a description of the content item, a commentary associated with the content item, etc.). The image confirmation module 340 can use one or more images to confirm the products identified in the content item. For example, the image confirmation module 340 can work similarly to the image recognition module 330, but can be used to confirm the presence of one or more products identified by a separate module (e.g., by comparing images of potential products with a more limited range of product images provided by another module). The fusion module 360 ​​may receive candidate products and associated confidence values ​​included in a content item and determine, based on various inputs, the likelihood that one or more products appear in the content item.

[0129] Image recognition 330 may be utilized to determine products associated with a content item having a visual component, such as a video. Image recognition 330 may include frame selection 332. Frame selection 332 may be utilized to select one or more frames of the video for searching images of products. Frame selection 332 may occur via random sampling, periodic sampling, smart sampling methods, etc. For example, a content item (e.g., a video) may be provided to a machine learning model, and the machine learning model may be trained to predict frames of the video that are likely to include one or more products.

[0130] The one or more frames may be provided to an object detection model 334. Object detection 334 may extract predicted objects from the one or more frames. For example, object detection 334 may isolate a potential product from a person, an animal, a background, etc. of the image data of the content item. Object detection 334 may be or include a machine learning model.

[0131] The image of the detected object may be supplied to embedding 336. Embedding 336 may include converting one or more images to a lower dimension. Embedding 336 may include providing one or more images to a dimensionality reduction model. The dimensionality reduction model may be a machine learning model. The dimensionality reduction model may be configured to reduce the dimensionality of similar images in a similar manner. For example, embedding 336 may receive an image as input and generate a value vector as output. Embedding 336 may be configured, trained, etc. so that similar images (e.g., images of the same or similar products) are similarly represented in a reduced dimensional vector space (e.g., by Cartesian distance, by cosine distance, by another distance metric, etc.). Embedding 336 may generate data with reduced dimensionality.

[0132] The reduced dimensional image data may be provided to product identification 338. Product identification 338 may identify one or more products associated with the reduced dimensional representation provided by embedding 336. Product identification 338 may compare the reduced dimensional image data (e.g., provided by embedding 336) with reduced dimensional image data of products included in product image index 339 (e.g., generated from images of products by the same machine learning model as used by embedding 336). Product image index 339 may be stored as part of a data repository. Product image index 339 may include, for example, many products (e.g., hundreds of products, thousands of products, or more). Product image index 339 may include associations between stored image data (e.g., reduced dimensional image data) and product identifiers, product indicators, and the like. Product image index 339 may be segmented, for example, the stored dimensionality-reduced data may be classified into one or more categories, classifications, and the like. For example, product identification 338 may compare data received from embedding 336 with products of a particular category, type, classification, and the like. In some embodiments, categories, types, classifications, etc. may be provided by one or more users, one or more content creators, may be automatically detected (e.g., by one or more machine learning models), etc. A content item or one or more products associated with a content item may be associated with a category (e.g., a general category such as electronic devices, a more limited category such as screen devices, a product classification such as a tablet, a brand or trademark, a model, etc.). Product recognition 338 may generate one or more indications of products detected in an image of a content item (e.g., a list of products that may match the products represented in the product image index 339) and one or more indications of confidence values ​​(e.g., the confidence that each product in the list of products has been accurately detected). The output of image recognition 330 may be used to update the metadata of the content item (e.g., updated to include an association with one or more products, updated to include one or more product identifiers or indicators, etc.). The output of image recognition 330 may be provided to image validation 340, for example, to validate the presence of an image identified by image recognition 330 in a content item. The output of image recognition 330 may be provided to fusion 360, for example, to generate an overall and / or multi-input determination of products included in the content item via fusion model 366. The output of image recognition 330 may be provided to text recognition 350 (not shown), for example, to limit the space of products queried, searched, compared, etc. by text recognition module 350. In some embodiments, image recognition 330 is utilized to identify products found in one or more frames of a video content item. For example, image recognition module 330 may be configured to generate a list of all products detected in any selected frame and provide a confidence value for each product in each selected frame.The image recognition module 330 may generate image-based product data, such as one or more identifiers of a product that is identified based on the image of the content item.

[0133] The image confirmation module 340 may be configured to confirm the presence of the identified product of the content item using one or more images of the content item. For example, the image confirmation module 340 may include a model configured to confirm the presence of products identified by other models. The image confirmation module 340 may include a secondary recognition 345. The secondary recognition 345 may include components similar to the image recognition 330. In some embodiments, instead of or in addition to the secondary recognition 345 communicating with the product candidate image index 344, the image recognition module 330 may also communicate directly with the product candidate image index 344. In some embodiments, the secondary recognition 345 may perform a similar role as the image recognition 330, but may include a different model, a model trained using different training data, a model configured to select different frames or detect different objects, etc.

[0134] Image validation 340 may include synthesis model 341. Synthesis model 341 may receive indications of products identified by image recognition module 330, text recognition module 350, secondary recognition 345 (data flow not shown), etc. Synthesis model 341 may select object detection model 342 (e.g., an object detection model specifically configured for a category or classification of products) to provide data to. Synthesis model 341 may include image selection, e.g., synthesis model 341 may provide one or more images to object detection 342, may select one or more frames to provide to object detection 334, etc. For example, synthesis model 341 may provide one or more frames that may include products to object detection 342 based on data received from image recognition 330 and text recognition 350. Object detection model 342 may perform similar functions to object detection model 334, e.g., modified by the functions of synthesis model 341. Embedding 343 may perform similar functions to embedding 336, e.g., to reduce the dimensionality of an image of a detected product. In some embodiments, product candidate image index 344 may include reduced dimensional image data (e.g., value vectors) detected by other modules (e.g., image recognition 330, text recognition 350, etc.) Secondary recognition 345 may compare the reduced dimensional image data (e.g., embedded image data) with candidate data of product candidate image index 344 (e.g., products recognized by modules other than image verification module 340) to verify the presence of the product in the content item.

[0135] Text recognition 350 can be configured to identify one or more products from text data (e.g., metadata) associated with a content item. Text recognition 350 can generate product data based on metadata, for example, one or more product identifiers based on metadata of the content item. Text recognition 350 can generate text-based product data, for example, one or more product identifiers based on text data associated with a content item. Text recognition 350 can identify products from one or more of a content item title, a content item description, a narration associated with a content item (e.g., a machine-generated narration), a comment associated with a content item, and / or other text data or metadata of a content item. Text data (e.g., metadata) associated with a content item can be provided to a text parsing model 352. The text parsing model 352 can be a machine learning model. The text parsing model 352 can be configured to detect or predict products from text data associated with a content item. The text parsing model 352 can be configured to detect one or more products having a product identifier stored in a product identifier 354. The text parsing model 352 may provide outputs (e.g., a list of detected candidate products, associated confidence values, etc.) to the image validation module 340. The text parsing model 352 may provide outputs to the synthesis model 341. The text parsing model 352 may provide outputs that affect products validated by the image, e.g., the outputs of the text parsing model 352 may cause products detected by the text recognition module 350 to be added to the product candidate image index 344. The image validation module 340 may query an index (e.g., the product candidate image index 344) that includes products detected by other modules, such as the image recognition module 330, the text recognition module 350, etc.

[0136] The fusion module 360 ​​may receive output data (e.g., detected products, associated confidence values) from one or more sources (e.g., image recognition module 330, image validation module 340, text recognition module 350, etc.). The fusion module 360 ​​may further receive data from other sources, such as context term extraction 362 or additional feature extraction 363. The context term extraction 362 may, for example, provide context to potential products of a content item, such as categories or themes associated with some products. The context term extraction 362 may be performed by one or more machine learning models. The context term extraction 362 may detect context information from text associated with a content item, metadata associated with a content item, etc. The additional feature extraction 363 may provide additional details that may be used to determine whether one or more products appear in a content item. Additional features may include video embeddings. Additional features may include other metadata for the content item, such as the date the content item was uploaded to the content providing platform (e.g., compared to the release date of a product), the category of the content item (e.g., a shopping or product review video may be more likely to include a product than another type of video), etc.

[0137] Data from multiple sources may be provided to the fusion model 366. The fusion model 366 may be configured to receive data including, for example, one or more products with confidence values, and determine the one or more products using the confidence values ​​indicating the likelihood of the products appearing in the content item. In some embodiments, the content item may be a video. In some embodiments, the content item may be a live feed, such as a live video feed (e.g., a product review stream, an unboxing stream, etc.). The content item may be a short video.

[0138] FIG. 4A to FIG. 4E Depicted is an example UI presented on a user device including a UI element indicating associated products in accordance with some embodiments. FIG. 4A to FIG. 4E The UI may be provided as part of an application of the apparatus 400A-400E, such as a web browser application, a mobile application associated with / provided by a content platform, etc. FIG. 4A to FIG. 4EUser interaction with various elements of a UI may result in changes to the presented UI elements. For example, interaction with a UI element indicating that a content item has an associated product may result in the display of a second UI element presenting additional information (e.g., about the associated product) (e.g., replacing the first UI element, expanding the first UI element, etc.). The UI element may include an element that, when interacted with, causes the UI to display less information about the associated product (e.g., folding a panel describing one or more associated products). Interacting with a UI element associated with one or more products may result in different effects, such as transitioning to a UI environment for presenting a content item. The new UI environment may include one or more UI elements associated with one or more products of the content item. Interacting with a UI element associated with a product may result in the display of a UI element that facilitates a transaction (e.g., purchase) of the product. FIG. 4A to FIG. 4E Various interactions between the UIs and UI elements presented in are possible (e.g., interacting with elements of a first UI layout may cause a transition to a second UI layout), and any transitions between sample UIs, similar UIs, inclusion of similar UI elements, etc. are within the scope of the present disclosure. FIG. 4A to FIG. 4E is described in conjunction with a video content item, and other types of content items (e.g., image content, text content, audio content, etc.) may be presented in a similar UI. FIG. 4A to FIG. 4E Any optional features, elements, etc. presented in relation to one or more of the figures may be included in a system similar to another of those figures.

[0139] Figure 4A Depicted is an apparatus 400A presenting an example UI 402 including a UI element 404 indicating one or more associated products in accordance with some embodiments. As part of an operation for presenting one or more content items for user selection, apparatus 400A may present a UI including elements of UI 402 and / or similar UI elements. UI element 404 may be displayed in a manner similar to UI 402. Figure 4A is depicted in a folded state, eg, a folded default state.

[0140] UI 402 includes a first content item selector 406 (e.g., a video thumbnail) and a second content item selector 408. In some embodiments, more or fewer content items may be selectable, UI 402 may be scrolled to view additional content items, etc. UI element 404 is associated with the content item indicated by content item selector 406. UI element 404 (and FIG. 4A to FIG. 4DIn some embodiments, product information (e.g., product / content item associations) may be provided by a content creator. Product information may be provided by one or more users. Product information may be provided by an administrator. Product information may be retrieved from a content item, for example, via one or more machine learning models, via a system such as system 300B, or the like.

[0141] In some embodiments, a user can interact with UI element 404 to be presented with a replaced UI element, an updated UI element, etc. For example, a user can interact with expansion element 410 to display more information about a product associated with a content item. In some embodiments, expansion UI element 404 can open a panel that includes additional information about one or more products associated with the content item. Expanding UI element 404, interacting with UI element 404, etc. can adjust the presentation of UI 402 to include, for example, FIG. 4B to FIG. 4D Elements depicted in .

[0142] UI element 404 may include, for example, an expansion element 410, indications of associated products (e.g., how many products are associated with the content item), visual indications 412 of products (e.g., a visual indication that a deal or purchase is available, a visual indication that a link to a merchant of the product is available, etc.), etc. User interaction with one or more components of UI element 404 may cause device 400A to modify the presentation of UI element 404, e.g., a user selection of expansion element 410 may cause the presentation of UI element 404 to be modified to an expanded state.

[0143] In response to device 400A providing content to a platform (eg, Figure 1400A) sends a request for a content item, and a UI including UI element 404 may be presented. UI element 404 may be presented in a home feed (e.g., a list of suggested content items for a user or user account). UI element 404 may be presented in a suggested feed (e.g., a list of suggested content items based on one or more recently presented content items). UI element 404 may be presented as part of a playlist (e.g., a list of content items to be presented, populated by a user, populated by a creator, etc.). UI element 404 may be presented in a search feed (e.g., in response to a user-generated search query sent by device 400A to a content providing platform). For example, UI elements associated with one or more products may be presented based on the inclusion of products, product categories, product trademarks, etc. included in the search query. UI element 404 may be presented in a feed focused on a product (e.g., a shopping content feed). UI element 404 may be presented in response to a user selection of a content item—e.g., it may be displayed to watch an associated video after a user selection, presented together with an associated content item, etc. UI element 404 can be presented based on the detected user's interest in the content item. For example, the user can stay on the thumbnail of the video (for example, the user can stop the cursor on the thumbnail, the user can pause the scrolling of the presented thumbnail, etc.). When the stay meets one or more conditions (for example, duration conditions, thumbnail position conditions, etc.), UI element 404 can be presented. UI element 404 can be presented in response to additional data such as user account history, user settings, user preferences, etc. UI element 404 can be presented together with a list of content items to be presented, UI element 404 can be presented when presenting content items (for example, when playing a video associated with the product of UI element 404), etc.

[0144] Figure 4B An apparatus 400B is depicted presenting an example UI 420 including a UI element 422 indicating associated products in accordance with some embodiments. The UI element 422 includes information about one or more products associated with the content item. The UI element 422 includes information about one or more products associated with the content item. Figure 4B422 is presented in an expanded state. UI element 422 may include multiple components. For example, a first component may include information about a first product (e.g., pictures, product names, prices, timestamps, etc.), a second component may include information about a second product, etc. In some embodiments, UI element 422 may be scrollable, for example, to access information about additional products. UI element 422 may include multiple tabs (e.g., UI element 422 may be associated with product information and one or more other types of information). For example, UI element 422 may include product tab 424 and chapter tab 426. In some embodiments, product tab 424 may be opened by default (e.g., the content of product tab 424 may be presented by default). In some embodiments, chapter tab 426 may be opened by default (e.g., the content of chapter tab 426 may be presented by default). In some embodiments, another tab may be opened by default. For example, section tab 426 may be opened by default, except for content items with associated products, and the tab that is opened by default may be selected based on a user history (e.g., interactions with elements of section tabs, product tabs, etc.), based on a search query (e.g., a search including a product name, a related term or phrase such as "product review", etc.), etc. UI elements 422 may include additional elements, e.g., elements that a user may utilize to control the presentation of UI 420, UI elements 422, etc. For example, UI elements 422 may include a "close" element for presenting UI 420 without UI element 422, may include a "close" element for displaying less information (e.g., for collapsing a panel to resemble a Figure 4A The UI element 404 in the flowchart, a “collapse” element for modifying the presentation of the UI element 422 to a collapsed state, etc., one or more elements for displaying more information (for example, one or more listed products, listed product icons, etc. can be selected to present additional information about the products to facilitate the purchase of the products, etc.), etc.

[0145] The product tab 424 of the UI element 422 may include one or more pictures of the product, information about the product (e.g., the name of the product, a description of the product, etc.), one or more prices of the product, timestamps of content items related to the product, etc. In some embodiments, the product information may be provided by a content creator, one or more users, a system administrator, etc. In some embodiments, the product information may be retrieved by one or more models, such as machine learning models. For example, one or more machine learning models may be used to determine the presence of a product in a content item, the association of a product in a content item, the timestamp or location at which a product appears in a content item, etc. Figure 3BThe system 300B of the system 300B in the embodiment of the present invention determines one or more products associated with a content item (e.g., a video). The portion of the content item related to the content item (e.g., a timestamp) can be determined, for example, via the timing of the commentary associated with the product, the timing of the display of the image or frame of the video including the product, etc. In some embodiments, selecting a product can cause the portion of the content item associated with the product to be presented, can cause the content item to be presented starting at the time indicated by the timestamp associated with the product, etc.

[0146] In some embodiments, the portion of UI element 422 (e.g., visual component) associated with a particular product may be displayed by default, displayed differently (e.g., highlighted), etc. For example, upon receiving a search query from a user that includes the name of a product, UI element 422 may be displayed that includes a display associated with the searched product.

[0147] In some embodiments, one or more associations between content items and products may be stored, for example, as metadata associated with the content items (the metadata associated with the content items may further include content item titles, descriptions, presentation history, narration associated with the content items, etc.). In response to the device (e.g., device 400B) executing instructions for displaying a list of content items for presentation, presenting content items to a user, presenting a UI element (e.g., UI element 422) including information about one or more products, etc., the device may retrieve information about the products based on the metadata associating the products with the content items. Information about the products (e.g., images, associated products such as color variants, availability, prices, etc.) may be retrieved from a data repository. The data repository may include information about the products and may be updated, for example, when information such as the price of the product changes, the UI may retrieve updated information based on the content item / product association and display the updated information.

[0148] UI element 422 may be presented as part of a home feed (e.g., a list of suggested content items for a user or user account), a suggested feed (e.g., a list of suggested content items based on one or more recently presented content items), a playlist, a list of search results, a shopping content page, etc. UI element 422 may be presented after a content item is selected for presentation, after a content item is presented, etc. UI element 422 may be displayed upon dwell (e.g., pausing scrolling on a content thumbnail, selecting a content item, etc.) Figure 4AIn some embodiments, UI element 422 may be removed or replaced upon user action, such as upon scrolling, UI element 422 may be collapsed to resemble UI element 404 (e.g., to facilitate selection of a content item from a list of content items, to simplify scrolling a list of content items, etc.). UI element 422 may be displayed when presenting a list of content items, when presenting a single content item (e.g., when playing a video associated with a product of UI element 422), etc.

[0149] Figure 4C Depicted is an apparatus 400C presenting an example UI 430 according to some embodiments, the example UI including a UI element 432 presenting a content item and a UI element 434 presenting information about an associated product. The UI element 434 may include, for example, Figure 4B UI element 422 may include more detailed information about the product associated with the content item. UI element 434 may be product-focused, for example, it may be used to display product information to the user. UI element 434 may include one or more components, for example, a component associated with a first product, a component associated with a second product, etc. In some embodiments, UI element 434 may include a list of products associated with the content item. UI element 434 may be navigable, scrollable, etc. UI element 434 may include one or more control elements, for example, a back button for returning to a previous view, a close button for closing UI element 434 and viewing a different set of UI elements (e.g., unrelated to the product), etc. In some embodiments, UI element 434 may be removed from UI 430 in response to another user action, such as a user scrolling through an associated content item. In some embodiments, a user selection of a product presented via UI 434 may prompt the presentation of a UI element that facilitates the purchase of the product. In some embodiments, UI element 434 may be displayed in response to determining that the user is interested in one or more products associated with a content item (e.g., based on user history, based on one or more terms in a user search query, based on a user selection of a content item to be browsed and / or presented with shopping, based on a user selection of a product or a UI element associated with a product, etc.).

[0150] UI element 434 may include one or more pictures and / or additional information about one or more products associated with a content item (e.g., a content item presented via UI element 432). Pictures and / or information may be provided by a content creator, provided by one or more users, retrieved from a database (e.g., based on product / content item associated metadata), etc. In some embodiments, UI element 432 may scroll automatically. For example, UI element 432 may scroll as the content item is presented, such as making the product associated with the part of the content item currently being presented visible. UI element 432 may be presented in response to a user selection of the content item to be presented. UI element 432 may be presented in response to other factors such as user history. UI element 432 may be presented when a user stays on an associated content item, an associated UI element, etc.

[0151] In some embodiments, UI element 434 can (e.g., via Figure 4B UI element 422) presents information about a single product, such as the product selected by the user. UI element 434 can be product-focused, can focus on a single product, can display product variations (e.g., such as combining Figure 4D 444 ), or may include other elements, components and / or information described with respect to other UI elements described herein.

[0152] Figure 4D An apparatus 400D is depicted presenting an example UI 440 according to some embodiments, the example UI including a UI element 442 presenting a content item and a UI element 444 facilitating a transaction associated with a product. The UI element 444 may provide one or more fields associated with a user conducting a transaction (e.g., purchasing a product). For example, the UI element 444 may include an alternative product panel 446, which may include information, pictures, prices, etc. about alternative products (e.g., products related to one or more products associated with the content item, such as color variations, size variations, variations in product bundling, related products such as similar products of another brand, etc.).

[0153] UI element 444 may include transaction element 448. In some embodiments, transaction element 448 may facilitate a transaction (e.g., a purchase) within the application providing UI 440. In some embodiments, transaction element 448 may facilitate a transaction via another application, another website, etc. For example, interacting with transaction element 448 may direct the user to a merchant website, may direct device 400D to open an application associated with purchasing a product, etc.

[0154] UI element 444 may be navigable, scrollable, etc. UI element 444 may include one or more control elements, such as a back button for returning to a previous view, a close button for closing UI element 444 and displaying a different set of UI elements via UI 440, etc. In some embodiments, UI element 444 may be displayed as part of a list of content items for the user to select, as part of UI 440 presenting content items, etc. UI element 444 may be displayed in response to determining that the user is interested in a transaction (e.g., purchase) associated with one or more products included in the content items—e.g., from a user such as a Figure 4C The UI element 434 in the example embodiment may be displayed by selecting a product from the UI element 434 in the example embodiment, selecting a content item, hovering over a content item, including a product or product-related term in a search query, navigating the user to a list of shopping-focused content items, etc. The UI element 444 may receive information about the one or more products from a database, for example, based on metadata associating the content item with the one or more products.

[0155] FIG. 4A to FIG. 4D The UI elements depicted in can be integrated into various configurations. For example, a UI element such as UI element 404 can be presented, and after a user interaction with UI element 404, an element such as UI element 422 can be presented, after a user interaction with UI element 422, UI element 434 can be presented, after a user interaction with UI element 434, UI element 444 can be presented, and the like. One or more UI elements may include a navigation element for instructing the device to display different UI elements, for example, interacting with expansion element 410 may cause a UI element such as UI element 422 to be presented, and interacting with different elements of UI element 404 may cause UI elements such as UI element 434, UI element 444, and the like to be presented.

[0156] Other connections between UI elements are also possible, for example, interacting with a UI element such as UI element 404 may cause a UI element such as UI element 422, such as UI element 434, such as UI element 444, etc. to be displayed. Interaction with a UI element such as UI element 422 or a portion thereof may cause a UI element such as UI element 404, such as UI element 434, such as UI element 444, etc. to be presented. User interaction with a UI element such as UI element 434 or a portion thereof may cause a UI element such as UI element 404, UI element 422, such as UI element 444, etc. to be presented. User interaction with a UI element such as UI element 444 or a portion thereof may cause a UI element such as UI element 404, such as UI element 434, such as UI element 444, etc. to be presented. The default UI element presented may depend on the environment in which the UI element is presented (e.g., a list of content items presented in response to a search, a home feed, a shopping feed, a viewing feed, etc.; an environment including content items being presented; etc.). For example, the selection of the form of a UI element may be based on multiple factors. In some embodiments, the inclusion of a product name, category, etc. in a search query may alter default UI elements, e.g., may cause the UI to default to present a UI element that includes information about the product, a UI element that includes purchase options for the product, etc. The determination of the form of the UI element to be displayed may be based on user history, user account history, user actions (e.g., opening of a home feed, presentation of a watch feed, transmission of a search query, selection of a shopping feed, etc.). Transitions between forms of UI elements associated with a product may be determined by additional data similar to the data used to determine the form of the UI element presented.

[0157] Figure 4E An example device 400E is depicted with UI elements superimposed on content presentation elements 452 in accordance with some embodiments. Device 400E includes UI 450. UI 450 may be provided by an application, such as an application associated with a content providing platform. Presentation element 452 may present a content item (e.g., a video). UI 450 may present a list of additional content items, additional information associated with the presented content item (e.g., a title, description, comments, live chat, etc.), additional UI elements associated with a product (e.g., UI elements or variations such as UI elements 404, 422, 434, 444), etc.

[0158] Presentation element 452 may be overlaid with one or more UI elements. UI element 454 may indicate a product included in the content item (e.g., shown in the video). UI element 454 may be executed in conjunction with FIG. 4A to FIG. 4DThe UI element 454 may be displayed and / or removed in response to the presence of related products in the content item, for example, it may be displayed when the content item is in the video. The UI element 454 may indicate multiple products, identify products (for example, one or more products may be named), display information about products, etc. Multiple UI elements such as UI element 454 may be displayed on a thumbnail of a video, in the entire presentation of the video, at the same time during the video, etc.

[0159] Overlaid UI elements such as UI element 454 may be presented in combination with other UI elements associated with products. For example, UI element 456 may open a panel including information about multiple products associated with a content item, UI element 454 may cause a UI element including information about a product displayed in a picture to be displayed, and so on.

[0160] The superimposed UI element 454 can be displayed above the visual representation of the content item (e.g., in front of the visual representation of the content item, with a visual priority that is superior to the visual representation of the content item, etc.). For example, UI element 454 can be superimposed on a video thumbnail. UI element 454 can be displayed above the content item. For example, UI element 454 can be superimposed on the video being played. The presentation of UI element 454 can be performed in response to user actions. For example, after determining that the user is interested in one or more products (e.g., via a search query, via interaction with a product-related UI element, via user history, etc.), UI element 454 can be superimposed on another UI element and displayed. In some embodiments, the content item can be a live video. In some embodiments, the content item can be a short video.

[0161] In some embodiments, UI element 454 may perform operations similar to the description of the performance of UI element 404, for example, the user may be notified that one or more products are associated with the content item. UI element 454 may respond to interactions from the user similar to UI element 404, for example, a panel including product information may be opened or expanded, the presentation of UI element 404 may be modified to display more or different information, the UI element may be expanded to include more information, the presentation of the content item may be initiated, etc. UI element 454 may respond to interactions from the user similar to UI element 434, for example, a panel that facilitates the transaction may be opened or expanded.

[0162] FIG. 5A to FIG. 5F 5 is a flow chart of methods 500A-500F related to content items having associated products according to some embodiments. Methods 500A-500F may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing devices, etc.), software (such as instructions running on a processing device, a general purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. In some embodiments, methods 500A-500F may be performed in part by Figure 1 The method 500A may be performed in part by the product identification system 175 (e.g., Figure 1 The server machine 170 and the data set generator 172 in Figure 2 According to an embodiment of the present disclosure, the product identification system 175 may use method 500A to generate a data set for at least one of training, validating, or testing a machine learning model. Methods 500B-500D may be performed by the product identification system 175 (e.g., Figure 3B The method 500E may be performed by the client device 110. The client device 110 may utilize the method 500E to display one or more UI elements associated with a product, for example, thereby facilitating user identification of a product included in a content item. The method 500F may be performed by the content platform system 102, for example, by processing logic of the content providing platform 120, to facilitate the client device 110 to present one or more UI elements associated with a product. In some embodiments, a non-transitory machine-readable storage medium stores instructions that, when executed by a processing device (e.g., of the product identification system 175, the server machine 180, etc.), cause the processing device to perform one or more of the methods 500A-500F.

[0163] For ease of explanation, methods 500A-500F are depicted and described as a series of operations. However, operations according to the present disclosure may occur in various orders and / or concurrently, and with other operations not presented and described herein. In addition, not all of the operations shown may be performed to implement methods 500A-500F according to the disclosed subject matter. In addition, those skilled in the art will understand and appreciate that methods 500A-500F may alternatively be represented as a series of interrelated states via state diagrams or events.

[0164] Figure 5A is a flow chart of a method 500A for generating a data set for a machine learning model according to some embodiments. Figure 5A In some embodiments, at block 401 , the processing logic implementing method 500A initializes the training set T to an empty set.

[0165] At block 502, processing logic generates a first data input (e.g., a first training input, a first validation input), which may include one or more of product data, image data, metadata, text data, confidence data, etc. In some embodiments, the first data input may include a first set of features for a data type, and the second data input may include a second set of features for a data type (e.g., such as a first set of features for a data type). Figure 3A In some embodiments, the input data may include historical data.

[0166] In some embodiments, at box 503, the processing logic optionally generates a first target output for one or more of the data inputs (e.g., a first data input). In some embodiments, the input includes one or more predicted products detected in the content item and an associated confidence interval, and the target output may include a label of the product included in the content item. In some embodiments, the input includes one or more sets of data associated with the content item (e.g., image data such as a frame of a video or a portion of a frame of a video, metadata such as a title text or a commentary text, etc.), and the target output is a list of products included in the content item. In some embodiments, the first target output is predictive data. In some embodiments, the input data may be in the form of commentary text data, and the target output may be a list of possible corrections to the commentary to include a product name / reference for a machine learning model configured to correct the commentary by including product information. In some embodiments, no target output is generated (e.g., an unsupervised machine learning model is able to group the input data or find correlations in the input data, rather than requiring a target output to be provided).

[0167] At block 504, processing logic optionally generates mapping data indicating an input / output mapping. The input / output mapping (or mapping data) may refer to a data input (e.g., one or more of the data inputs described herein), a target output for the data input, and an association between the data input and the target output. In some embodiments, such as in association with a machine learning model in which a target output is not provided, block 504 may not be performed.

[0168] In some embodiments, at block 505 , processing logic adds the mapping data generated at block 504 to data set T.

[0169] At block 506, processing logic branches based on whether the data set T is sufficient for at least one of training, validating, and / or testing a machine learning model, such as Figure 1 If so, execution proceeds to block 507, otherwise, execution continues back to block 502. It should be noted that in some embodiments, the adequacy of the dataset T may be determined based solely on the number of inputs in the dataset that, in some embodiments, map to outputs, while in some other embodiments, the adequacy of the dataset T may be determined based on one or more other criteria (e.g., a measure of diversity of data examples, accuracy, etc.) in addition to or in lieu of the number of inputs.

[0170] At block 507, processing logic provides a data set T (e.g., to Figure 1In some embodiments, the data set T is a training set and is provided to the training engine 182 of the server machine 180 to perform training. In some embodiments, the data set T is a validation set and is provided to the validation engine 184 of the server machine 180 to perform validation. In some embodiments, the data set T is a test set and is provided to the testing engine 186 of the server machine 180 to perform testing. For example, in the case of a neural network, the input values ​​of a given input / output mapping (e.g., the numerical values ​​associated with the data input 210) are input to the neural network, and the output values ​​of the input / output mapping (e.g., the numerical values ​​associated with the target output 220) are stored in the output nodes of the neural network. The connection weights in the neural network are then adjusted according to a learning algorithm (e.g., back propagation, etc.), and the process is repeated for other input / output mappings in the data set T. After box 507, the model (e.g., model 190) can be at least one of: trained using the training engine 182 of the server machine 180, validated using the validation engine 184 of the server machine 180, or tested using the testing engine 186 of the server machine 180. The trained model may be implemented by product recognition system 175 to generate output data, e.g., used by product information platform 161 to provide product data to users, provided to a fusion model, utilized to update metadata for a content item to include one or more product associations, and the like.

[0171] Figure 5B 5 is a flow chart of a method 500B for updating metadata for a content item according to some embodiments. At block 510, processing logic (e.g., a processing device, a computer processor, etc.) receives first data. The first data includes a first identifier (e.g., an indicator, a pointer to further data, a code identifying the product, etc.) of a first product determined in association with the content item based on the metadata of the content item. The content item may be or include visual content, audio content, textual content, video content, etc. The metadata of the content item may have been provided to one or more trained machine learning models (e.g., Figure 3B The first data may further include a first confidence value associated with the first product and the content item. The first confidence value may indicate the likelihood that the first product is associated with the first content item, such as the likelihood that the first product appears or is referenced in the metadata of the content item. The first data may further include an identifier of a second product, and a second confidence value associated with the second product. The first data may include a list of products and associated confidence values.

[0172] At block 512, processing logic receives second data including a second identifier of the first product. The second identifier has been determined in association with the content item based on image data of the content item (e.g., one or more frames of a video, portions of one or more images, etc.). The second data also includes a second confidence value associated with the first product and the content item. The confidence value and the identifier may be generated by one or more machine learning models (e.g., Figure 3B The machine learning model may include a system configured to reduce the dimensionality of the image data. One or more candidate product images may be subjected to dimensionality reduction (e.g., converting the image to a value vector via a trained machine learning model). The machine learning model may perform an operation including comparing the reduced dimensional image data from the content item with the reduced dimensional product images of the data repository, for example, to determine the likelihood that the image includes a product. The second data may include a list of products (e.g., product identifiers) and a list of confidence values, the products including at least the first product. The confidence value may indicate the likelihood that the associated product appears in the content item, is referenced by the content item, etc.

[0173] In some embodiments, the one or more images are analyzed for potential products included in one or more images of the content item. For example, the existence of the potential product can be confirmed by providing images of the content item for further product image detection analysis (e.g., different frames, additional frames, etc. of the video content item), by text or metadata confirmation, etc. For example, further analysis for confirmation of the candidate product can be performed after the candidate product is found, for example, a search for other evidence of the identified product can be performed. In some embodiments, text data and / or metadata associated with the content item can be analyzed for potential / candidate products. The existence of the potential product can be confirmed, for example, by image-based confirmation, text confirmation, etc.

[0174] In some embodiments, the processing logic may be further provided with one or more timestamps, such as a timestamp of a frame of a video having a detected candidate product, a timestamp of a commentary associated with the video or audio content in which the detected candidate product appears, etc. The processing logic may utilize the timestamps to perform further analysis to adjust metadata for the content item, generate UI elements associated with the presentation of the content item, etc.

[0175] At block 514, processing logic provides the first data and the second data to a trained machine learning model. The trained machine learning model may be a fusion model. The trained machine learning model may be provided with one or more lists of products having associated confidence values.

[0176] At block 516, processing logic receives a third confidence value associated with the first product from the trained machine learning model.In some embodiments, processing logic may receive a list of confidence values ​​associated with a list of products including the first product.

[0177] At block 518, processing logic adjusts metadata associated with the content item in view of the third confidence value. In some embodiments, adjusting the metadata may include adding one or more connections between the content item and the product to the metadata. For example, adjusting the metadata may include adding an indication that a particular product is associated with the content item, is displayed in the content item, is included in the content item, is advertised by the content item, etc. Adjusting the metadata may include adjusting the transcript of the content item to, for example, include one or more references to the product that was incorrectly transcribed during transcript generation.

[0178] Figure 5C 5 is a flow chart of a method 500C for training a machine learning model associated with a content item product pairing according to some embodiments. In some embodiments, the machine learning model trained using the method 500C may be a fusion model. Similar methods may be utilized to train different models connected to a media item product pairing, such as an image recognition model, an image verification model, a text recognition model, a commentary update model, and the like.

[0179] At block 520, processing logic receives product image data associated with a plurality of content items. The product image data may include data associating products with content items, the association being derived from one or more images (e.g., frames of a video). The product image data includes an indication of one or more products (e.g., potential products, candidate products) detected (e.g., determined) in the image and one or more product image confidence values.

[0180] At block 522, processing logic receives product text data associated with a plurality of content items. The product text data may include data associating products with the content items. The association may be derived from text associated with the content items (e.g., metadata associated with the content items). The product text data includes indications of one or more products (associated with the content items) detected in the text and one or more product text confidence values.

[0181] The data received (or obtained in some embodiments) by the processing logic at blocks 520 and 522 may be used as training input for training the fusion model. Training a machine learning model for performing different functions may include the processing logic receiving different data as training input.

[0182] At block 524, processing logic receives data indicating products included in the plurality of content items. For example, each of the plurality of content items used to train the model (e.g., data associated with the content item may be used to train the model) may include a list of associated products, e.g., tagged by one or more users, tagged by a content creator, etc. The data received by processing logic at block 524 may be used as a target output for training the fusion model. Training machine learning models for performing different functions may include processing logic receiving different data as target outputs.

[0183] At box 526, processing logic provides the product image data and product text data as training inputs to the machine learning model. Processing logic can provide different types of data to train different machine learning models. In some embodiments, a machine learning model for frame selection can be trained by providing frames of a video as training inputs to the model. In some embodiments, a machine learning model for object detection can be trained by providing images (possibly including products) as training inputs to the machine learning model. A machine learning model for embedding can be trained by providing one or more images of an object (e.g., a product) as training inputs to the model. In some embodiments, a text parsing model can be trained by providing text associated with a content item (e.g., metadata) as training input. In some embodiments, a machine-generated commentary can be provided as training input to a model configured to correct commentary.

[0184] At box 528, the processing logic provides data indicating products included in the multiple content items (e.g., a list of products included in each of the multiple content items) to the machine learning model as a target output. The processing logic may provide different types of data to train different machine learning models. In some embodiments, a machine learning model for frame selection may be trained by providing data indicating which frames of one or more videos include products as a target output. In some embodiments, a machine learning model for object detection may be trained by providing a label of an object in an image provided to the model as a target output. In some embodiments, a text parsing model may be trained by providing a content item referenced by the text of the content item as a target output. In some embodiments, a corrected commentary (e.g., including one or more products) may be provided as a target output to a model configured to correct the commentary. In some embodiments, no target output is provided to train a machine learning model (e.g., an unsupervised machine learning model).

[0185] Figure 5D5 is a flow chart of a method 500D for adjusting metadata associated with a content item according to some embodiments. At block 530, processing logic obtains first metadata associated with the content item. The metadata may include text data. The metadata may include a content item title, description, commentary, comments, real-time chat, etc. At block 531, processing logic provides the first metadata to a first model. In some embodiments, the model is a trained machine learning model. In some embodiments, the model is a product detection model, for example, the model is configured to receive metadata and (e.g., in view of the metadata) generate an indication of a product associated with the content item.

[0186] At block 532, processing logic obtains a first product identifier based on the first metadata and a first confidence value associated with the first product identifier as an output of the first model. The product identifier may be an ID number, an indicator, a product name, or any data that (uniquely) distinguishes a product. The first product identifier may identify the first product. In some embodiments, processing logic may obtain a list of products (e.g., candidate products, potential products) and associated confidence values.

[0187] At box 533, processing logic obtains image data of the content item. In some embodiments, the image data may include one or more frames of a video or be extracted from one or more frames of a video. In some embodiments, the image data may be obtained from an object detection model. In some embodiments, the image data may include one or more products associated with the content item.

[0188] At box 534, processing logic provides the image data to a second model. In some embodiments, the second model is a machine learning model. In some embodiments, the second model is a model configured to identify a product from an image. In some embodiments, the second model is a model configured to confirm the presence of a product identified from an image. In some embodiments, the second model may reduce the dimensionality of the provided image data. In some embodiments, the second model may compare the reduced dimensional image data to second reduced dimensional image data (e.g., retrieved from a data repository, output by a machine learning model, etc.).

[0189] At block 535, processing logic obtains a second product identifier based on the image data and a second confidence value associated with the second product identifier as an output of the second model. In some embodiments, the second product identifier indicates a second product. In some embodiments, the second product is the same as the first product. In some embodiments, processing logic may obtain a list of products and associated confidence values.

[0190] At block 536, processing logic provides the data including the first product identifier, the first confidence value, the second product identifier, and the second confidence value as input to a third model.The third model may be a fusion model.

[0191] At box 537, processing logic obtains a third product identifier and a third confidence value as output of the third model. In some embodiments, the third confidence value may indicate the likelihood that the product indicated by the third product identifier is associated with (e.g., present in) the content item. In some embodiments, the third model may output a list of products and associated confidence values. In some embodiments, the third product identifier identifies the third product. In some embodiments, the third product is the same as the second product. In some embodiments, the third product is the same as the first product. In some embodiments, the first, second, and third products are all the same product.

[0192] At box 538, processing logic adjusts second metadata associated with the content item in view of the third product identifier and the third confidence value. Adjusting the metadata may include supplementing the metadata with one or more product associations, e.g., indications of the associated products. Adjusting the metadata may include updating the transcript to, for example, include products that may have been incorrectly transcribed (e.g., incorrectly transcribed by a machine-generated transcript annotation model). In some embodiments, processing logic may further receive one or more timestamps associated with the content item and the one or more products (e.g., the time at which the product was detected in an image in the video). Updating the metadata may include adding to the metadata an indication of the time at which the product was found in the content item.

[0193] Figure 5E 5 is a flowchart of a method 500E for presenting a UI element associated with one or more products according to some embodiments. At box 540, processing logic (e.g., of a user device, a client device, etc.) presents a UI. The UI includes one or more graphical representations of one or more content items (e.g., videos). A graphical representation of a content item (e.g., a video thumbnail) can be selected to initiate the presentation of an associated content item. One or more graphical representations of a content item can be displayed together with a UI element associated with one or more products. Each graphical representation of a corresponding content item can be displayed together with a UI element associated with one or more products. The UI element can be presented / displayed in a folded state, such as a folded default state. The UI element includes information identifying a plurality of products covered by the corresponding content item. The UI element can identify that one or more products are associated with the content item. The UI element can identify (e.g., via a name, picture, etc.) one or more products associated with the content item (e.g., covered in a video).

[0194] The UI may present a selectable graphical representation of a content item. The represented content item may be part of a home feed, provided in response to a search, may be part of a watch list, may be part of a play list, may be part of a shopping feed, etc. In some embodiments, the UI element may be superimposed on top of and / or in front of one or more other elements of the UI. For example, a UI element (e.g., in a folded state) may be superimposed on a graphical representation of a content item, may be superimposed on a content item (e.g., while the content item is being presented), etc.

[0195] At box 542, in response to user interaction with the UI element in the folded state, the processing logic continues to promote the presentation of the graphical representation of the corresponding video, while modifying the presentation of the UI element from the folded state to the extended state. Interaction with the UI element may include selecting the UI element. Interaction with the UI element may include staying on the UI element (e.g., placing a cursor on the UI element, scrolling to the UI element and pausing scrolling, etc.). The UI element in the extended state may include multiple visual components. Each visual component may be associated with a product in a plurality of products. The visual component may include pictures, descriptions, prices, timestamps, etc. associated with various products.

[0196] In some embodiments, the UI element (e.g., in an expanded state) may include multiple tabs. For example, the UI element may include tabs for products, tabs for chapters or sections of content items, etc. For content items with associated products, the UI element may display / open the tab for the product by default. The UI element may display the tab for the product by default in response to user actions and / or history.

[0197] At box 544, in response to a user selection of one of the multiple visual components in the UI element in the expanded state, the processing logic initiates presentation of a corresponding content item covering the product associated with the selected visual component. The processing logic may initiate playback of a video covering the product associated with the selected visual component. The processing logic may initiate presentation of a portion of the content item associated with the product of the selected visual component (e.g., initiate playback of a portion of the video associated with the product of the selected visual component).

[0198] In some embodiments, interacting with a UI element may cause the UI element to be modified to a state focused on a product. The state focused on a product may present additional information, detailed information, etc. about one or more products. Interaction with a UI element in a collapsed state may cause the UI element to be modified to a state focused on a product. Interaction with a UI element in an expanded state (e.g., interaction with a visual component of a UI element associated with a product) may cause the UI element to be modified to a state focused on a product.

[0199] In some embodiments, interacting with a UI element may cause the presentation of the UI element to be modified to a transaction state. Interaction with a UI element in a collapsed state may cause the presentation of the UI element to be modified to a transaction state. Interaction with a UI element in an expanded state (e.g., selection of a component associated with a product) may cause the presentation of the UI element to be modified to a transaction state. Interaction with a UI element in a state focused on a product may cause the presentation of the UI element to be modified to a transaction state. The following determination may be performed based on user history, user preferences, content items, content item feeds (e.g., search results, viewing feeds, etc.), etc.: whether selection or interaction with a UI element, UI element component, etc. causes a transition to a transaction state.

[0200] In some embodiments, UI elements may be superimposed on the presented content items. For example, a UI element identifying one or more products may be superimposed on the video while a video is playing, while a video is showing one or more products, etc. Selection of a superimposed UI element may cause additional UI elements to be displayed, may cause the superimposed UI element to be modified to a different state, may cause a separate UI element to be modified to a different state, etc. The superimposed UI element may be in a folded state, an extended state, a product-focused state, a transaction state, etc. The presence and / or location of the superimposed UI element may be determined by, for example, one or more trained machine learning models configured to detect one or more models of a product.

[0201] Fig. 5F 5 is a flowchart of a method 500F for instructing a device to present one or more UI elements associated with a product according to some embodiments. At block 550, processing logic provides a UI including one or more graphical representations of one or more content items to the device. The graphical representation is provided for display / presentation by the UI of the device. Each graphical representation of the corresponding content item is selectable to initiate the presentation of the corresponding content item. The content item may include a video. The content item may include a live video. The graphical representation may be provided in response to a request of the device. The graphical representation may include a home feed, a viewing feed, a playlist, a search result list, a shopping feed, etc. The instructions sent to the device (e.g., including instructions associated with any step of method 500F), the UI sent to the device, the UI elements sent to the device, etc. may be determined / selected based on obtaining the user's history. The user's history may include historical interactions and / or selections of content items including content items with associated products. The user's history may include historical interactions and / or selections of UI elements or components of UI elements associated with the product. The user's history may include one or more searches of the user, for example, searches including product names. Instructions may be provided to the device in response to the processing logic receiving the user's history.

[0202] One or more of the graphical representations of the content items are displayed with the UI element in a collapsed state. In some embodiments, each graphical representation is displayed with the UI element in a collapsed state. In some embodiments, a subset of graphical representations are displayed with the UI element in a collapsed state. The UI element in a collapsed state is presented / displayed with a first graphical representation of a first content item. The UI element includes information identifying a plurality of products covered by the first content item. The UI element may identify how many products are associated with the content item, may identify one or more products by name, may identify a category or classification of products covered by the content item, and the like. In some embodiments, the plurality of products are obtained as output from one or more trained machine learning models. The trained machine learning models may be combined with Figure 3B The models described are similar.

[0203] At block 554, in response to receiving an indication of a user interaction with a UI element in a collapsed state, processing logic causes the device to modify the presentation of the UI element. The presentation of the UI element can be modified from a collapsed state to an expanded state. The UI element in an expanded state can include multiple visual components, each visual component being associated with one of the multiple products. The visual components can include pictures, names, descriptions, prices, timestamps, etc.

[0204] At block 556, in response to receiving an indication of a user selection of one of the multiple visual components of the UI element in the expanded state, processing logic facilitates presentation of a first content item. Processing logic may provide instructions to facilitate presentation of a portion of the first content item associated with a first product, such as a product associated with one of the multiple visual components. Processing logic may provide instructions for displaying a portion of a video associated with the product (e.g., for playing the video starting from a selected point in the video based on a timestamp associated with the product).

[0205] In some embodiments, the processing logic may further provide instructions to the device for modifying the presentation of the UI element to a product-focused state. For example, upon selecting a visual component of the UI element in an expanded state, the UI element may be modified to a product-focused state. The product-focused state may include additional details about one or more products covered by, included in, associated with, or the like, the content item.

[0206] In some embodiments, the processing logic may further provide instructions to the device for modifying the presentation of the UI element to a transaction state. The transaction state may facilitate a user to initiate a transaction associated with a product (e.g., purchase a product). The transaction state may be presented in response to a user action, user history, user selection of one or more UI elements, etc. The UI element in the transaction state may include one or more components that facilitate a transaction associated with one or more products.

[0207] Figure 6 6 is a block diagram illustrating a computer system 600 according to some embodiments. In some embodiments, the computer system 600 can be connected to other computer systems (e.g., via a network such as a local area network (LAN), an intranet, an extranet, or the Internet). The computer system 600 can operate as a server or client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. The computer system 600 can be provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a web appliance, a server, a network router, a switch or a bridge, or any device capable of executing a set of instructions (sequential instructions or other instructions) specifying the actions to be taken by the device. Further, the term "computer" should include any collection of computers that execute a set (or multiple sets) of instructions alone or in combination to perform any one or more of the methods described herein.

[0208] In a further aspect, the computer system 600 may include a processing device 602, a volatile memory 604 (e.g., a random access memory (RAM)), a non-volatile memory 606 (e.g., a read-only memory (ROM) or an electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.

[0209] The processing device 602 may be provided by one or more processors, such as a general-purpose processor (e.g., such as a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor that implements other types of instruction sets, or a microprocessor that implements a combination of multiple types of instruction sets) or a special-purpose processor (e.g., such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).

[0210] The computer system 600 may further include a network interface device 622 (e.g., coupled to the network 674). The computer system 600 may also include a video display unit 610 (e.g., LCD), an alphanumeric input device 612 (e.g., keyboard), a cursor control device 614 (e.g., mouse), and a signal generating device 620.

[0211] In some embodiments, the data storage device 618 may include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable medium) on which instructions 626 encoding any one or more of the methods or functions described herein may be stored, including instructions for Figure 1 Components (eg, content providing platform 120, other platforms of content platform system 102, communication application 115, model 190, etc.) of the content platform system 102 are encoded and used to implement instructions for the methods described herein.

[0212] The instructions 626 may also reside, completely or partially, within the volatile memory 604 and / or the processing device 602 during execution of the instructions by the computer system 600, and thus, the volatile memory 604 and the processing device 602 may also constitute machine-readable storage media.

[0213] Although the computer-readable storage medium 624 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" shall include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" shall also include any tangible medium that can store or encode a set of instructions for execution by a computer, which instructions can cause the computer to perform any one or more of the methods described herein. The term "computer-readable storage medium" shall include, but is not limited to, solid-state memories, optical media, and magnetic media.

[0214] The methods, components, and features described herein may be implemented by discrete hardware components, or may be integrated into the functionality of other hardware components such as ASICS, FPGAs, DSPs, or similar devices. In addition, the methods, components, and features may be implemented by firmware modules or functional circuitry within a hardware device. Further, the methods, components, and features may be implemented in any combination of hardware devices and computer program components or in a computer program.

[0215] Unless otherwise specifically stated, terms such as "receive," "execute," "provide," "obtain," "cause," "access," "determine," "add," "use," "train," "reduce," "generate," "correct," and the like refer to actions and processes performed or implemented by a computer system that manipulate and transform data represented as physical (electronic) quantities within computer system registers and memories into other data similarly represented as physical quantities within computer system memories or registers or other such information storage, transmission, or display devices. In addition, the terms "first," "second," "third," "fourth," and the like as used herein are intended as marks for distinguishing between different elements and may not have ordinal meanings according to their numerical designations.

[0216] The examples described herein also relate to a device for performing the methods described herein. This device may be specially constructed to perform the methods described herein, or the device may include a general-purpose computer system selectively programmed by a computer program stored in the computer system. Such a computer program may be stored in a computer-readable tangible storage medium.

[0217] The methods and illustrative examples described herein are not inherently related to any particular computer or other device. Various general purpose systems may be used according to the teachings described herein, or more specialized devices may be conveniently constructed to perform the methods described herein and / or each of their individual functions, routines, subroutines, or operations. Examples of the structures of a variety of these systems are set forth in the above description.

[0218] The above description is intended to be illustrative rather than restrictive. Although the present disclosure has been described with reference to specific illustrative examples and embodiments, it will be appreciated that the present disclosure is not limited to the described examples and embodiments. The scope of the present disclosure should be determined with reference to the attached claims and the full scope of equivalents to which the claims are authorized.

[0219] References throughout this specification to "one implementation" or "an implementation" mean that a particular feature, structure, or characteristic described in connection with that implementation is included in at least one implementation. Thus, depending on the circumstances, the appearance of the phrase "in one implementation" or "in an implementation" in various places throughout this specification may, but do not necessarily, refer to the same implementation. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more implementations.

[0220] To the extent that the terms "includes," "including," "has," "contains," variations thereof, and other similar words are used in the detailed description or claims, these terms are intended to be inclusive in a manner similar to the term "comprising" as an open transitional word and do not exclude any additional or other elements.

[0221] As used in this application, the terms "component", "module", "system", etc. are generally intended to refer to a computer-related entity, that is, hardware (e.g., circuitry), software, a combination of hardware and software, or an entity associated with an operating machine having one or more specific functionalities. For example, a component can be, but is not limited to, a process, a processor, an object, an executable program, an execution thread, a program, and / or a computer running on a processor (e.g., a digital signal processor). For example, both an application running on a controller and the controller can be components. One or more components can reside within a process and / or execution thread, and a component can be confined to one computer and / or distributed between two or more computers. Further, a "device" can appear in the following forms: specially designed hardware; general-purpose hardware specialized by executing software on it that enables the hardware to perform specific functions (e.g., generating points of interest and / or descriptors); software on a computer-readable medium; or a combination thereof.

[0222] The aforementioned systems, circuits, modules, etc. have been described with respect to the interaction between several components and / or blocks. It is understood that such systems, circuits, components, blocks, etc. may include those components or specified subcomponents, some of the specified components or subcomponents, and / or additional components, and according to various permutations and combinations of the aforementioned. Subcomponents may also be implemented as components that are communicatively coupled to other components rather than being included in a parent component (layered). In addition, it should be noted that one or more components may be combined into a single component that provides aggregation functionality or divided into several separate subcomponents, and any one or more intermediate layers such as a management layer may be provided to communicatively couple to such subcomponents in order to provide integrated functionality. Any component described herein may also interact with one or more other components that are not specifically described herein but are known to those skilled in the art.

[0223] In addition, the word "example" or "exemplary" is used herein to mean serving as an example, instance or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be interpreted as being preferred or advantageous over other aspects or designs. Instead, the use of the word "example" or "exemplary" is intended to present the concept in a specific way. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X adopts A or B" is intended to mean any one of the natural inclusive arrangements. That is, if X adopts A; X adopts B; or X adopts both A and B, then "X adopts A or B" is satisfied under any of the aforementioned examples. In addition, unless otherwise specified or clear from the context that it is for a singular form, the articles "one" and "a kind of" used in this application and the appended claims should generally be interpreted as meaning "one or more".

Claims

1. A method comprising: obtaining, by a processing device, first data comprising (i) a first identifier of a first product determined in association with a content item based on first metadata of the content item, and (ii) a first confidence value associated with the first product and the content item; obtaining, by the processing device, second data comprising (i) a second identifier of the first product determined in association with the content item based on first image data of the content item, and (ii) a second confidence value associated with the first product and the content item; providing, by the processing device, the first data and the second data to a trained machine learning model; obtaining, from the trained machine learning model, a third confidence value associated with the first product; as well as Second metadata associated with the content item is adjusted in view of the third confidence value.

2. The method of claim 1, further comprising: providing the first metadata of the content item as input to a second model; as well as The first data is obtained as an output of the second model.

3. The method of claim 2, wherein: The first metadata includes at least one of the following: the title of the content item; a description of the content item; or A transcript associated with the content item.

4. The method of claim 1, further comprising: providing the first image data of the content item as input to a second model; obtaining first dimensionally reduced data as an output of the second model; as well as Second dimensionally reduced data associated with the first product is obtained from a data repository.

5. The method of claim 4, wherein: The second dimensionally reduced data is obtained from the data repository in response to obtaining the first data, and wherein the second data is generated based on at least the first dimensionally reduced data and the second dimensionally reduced data.

6. The method of claim 4, further comprising: providing the second image data to a third model, and A third identifier of the first product is obtained from the third model, wherein the second dimensionally reduced data is obtained from the data repository in response to obtaining the third identifier of the first product, and wherein the second data is generated based on at least the first dimensionally reduced data and the second dimensionally reduced data.

7. The method of claim 1, wherein: The content item is a video, and wherein the first data further comprises an indication of a timestamp of one or more frames of the video associated with the product, and wherein adjusting the second metadata comprises including an indication of the first product and the indication of the timestamp in the second metadata.

8. The method of claim 1, wherein: Adjusting the second metadata includes adjusting a caption associated with the product.

9. The method of claim 1, further comprising training a machine learning model to generate the trained machine learning model, wherein: Training the machine learning model includes: receiving image-based product data associated with a plurality of content items, wherein the image-based product data includes indications of one or more products detected in an image and one or more product image confidence values; receiving metadata-based product data associated with the plurality of content items, wherein the metadata-based product data includes indications of one or more products detected in the text and one or more confidence values; receiving data indicative of products included in the plurality of content items; providing the image-based product data and the metadata-based product data as training inputs to the machine learning model; and The data indicative of products included in the plurality of content items is provided as a target output to the machine learning model.

10. The method of claim 1, further comprising: receiving third data including a third identifier of a first product category associated with the content item; as well as The third data is provided to the trained machine learning model, wherein the third confidence value is generated based on the first data, the second data, and the third data.

11. A method comprising: obtaining, by a processing device, first metadata associated with a content item; providing the first metadata to a first model; obtaining, as an output of the first model, a first product identifier based on the first metadata and a first confidence value associated with the first product identifier; obtaining image data of the content item; providing the image data to a second model; obtaining, as an output of the second model, a second product identifier based on the image data and a second confidence value associated with the second product identifier; The following data are provided as input to the third model: the first product identifier, the first confidence value, the second product identifier, and the second confidence value; obtaining a third product identifier and a third confidence value as an output of the third model; as well as Second metadata associated with the content item is adjusted in view of the third product identifier and the third confidence value.

12. The method of claim 11, wherein: Generating the second confidence value includes: reducing a dimension of the image data to generate first dimensionally reduced data; obtaining second dimensionally reduced data from a data repository, wherein the second dimensionally reduced data is associated with the product indicated by the second product identifier; and One or more operations are performed to generate the second confidence value, wherein the second confidence value is based on one or more differences between the first dimensionality reduced data and the second dimensionality reduced data.

13. The method of claim 11, wherein: Each of the first product identifier, the second product identifier, and the third product identifier identifies a first product.

14. The method of claim 11, further comprising obtaining a time stamp associated with the image data and the content item, wherein: Adjusting second metadata associated with the content item includes adjusting the second metadata to include an indication that a product identified by the second product identifier is associated with the timestamp and the content item.

15. The method of claim 11, wherein: The second metadata includes machine-generated narration, and wherein the first product identifier is associated with a product, wherein language associated with the product was incorrectly transcribed when generating the machine-generated narration, and wherein updating the second metadata associated with the content item includes replacing a portion of the machine-generated narration associated with the product with a text identifier of the product.

16. The method of claim 11, further comprising: providing a fourth product identifier and a fourth confidence value to the third model; as well as A fifth product identifier is obtained as an output of the third model, wherein the third product identifier is associated with a first product and the fifth product identifier is associated with a second product.

17. A non-transitory machine-readable storage medium storing instructions that, when executed, cause a processing device to perform operations comprising: obtaining first data comprising (i) a first identifier of a first product determined in association with the content item based on first metadata of the content item, and (ii) a first confidence value associated with the first product and the content item; obtaining second data comprising (i) a second identifier of the first product determined in association with the content item based on first image data of the content item, and (ii) a second confidence value associated with the first product and the content item; providing the first data and the second data to a trained machine learning model; obtaining, from the trained machine learning model, a third confidence value associated with the first product; as well as Second metadata associated with the content item is adjusted in view of the third confidence value.

18. The non-transitory machine-readable storage medium of claim 17, wherein: The operations further include: providing the first image data of the content item as input to a second model; obtaining first dimensionally reduced data as an output of the second model; and Second dimensionally reduced data associated with the first product is obtained from a data repository in response to obtaining the first data, wherein the second data is based on at least the first dimensionally reduced data and the second dimensionally reduced data.

19. The non-transitory machine-readable storage medium of claim 17, wherein: The content item is a video, and wherein the first data further comprises an indication of a timestamp of one or more frames of the video associated with the product, and wherein adjusting the second metadata comprises including an indication of the first product and the indication of the timestamp in the second metadata.

20. The non-transitory machine-readable storage medium of claim 17, wherein: The operations further include: receiving third data including a third identifier of a first product category associated with the content item; and The third data is provided to the trained machine learning model, wherein the third confidence value is generated based on the first data, the second data, and the third data.

Citation Information

Patent Citations

  • Identify objects within image from user of online system matching products identified to online system by user

    CN112819025A

  • Video search engine using joint categorization of video clips and queries based on multiple modalities

    US20070255755A1

  • Selecting a product for inclusion in a content item for a user of an online system based on products previously accessed by the user and by other online system users

    US20180336621A1

  • Machine-Based Object Recognition of Video Content

    US20200134320A1

  • Automatic content recognition and information in live streaming suitable for video games

    US20220222470A1