Context-based media curation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SNAP INC
- Filing Date
- 2020-03-26
- Publication Date
- 2026-08-04
Smart Images

Figure CN122507900A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202080039248.2, entitled “Context-Based Media Curating” (filed on March 26, 2020).
[0002] Priority requirements
[0003] This application claims priority to U.S. Patent Application No. 16 / 370,373, filed March 29, 2019, the entire contents of which are incorporated herein by reference. Technical Field
[0004] Embodiments of this disclosure generally relate to mobile computing technologies, and more specifically, but not in a limited way, to systems for curating and presenting collections of media content based on user context. Background Technology
[0005] Augmented reality (AR) is a real-time, direct or indirect view of a physical, real-world environment, whose elements are enhanced by computer-generated sensory input. Attached Figure Description
[0006] To facilitate identification of any particular element or action being discussed, one or more of the most significant figures in the reference numerals refer to the figure number in which the element is first introduced.
[0007] Figure 1 This is a block diagram illustrating an example messaging system for exchanging data (e.g., messages and associated content) over a network according to some embodiments, wherein the messaging system includes a media curation system.
[0008] Figure 2 This is a block diagram illustrating further details of a messaging system according to an example embodiment.
[0009] Figure 3 This is a block diagram illustrating various modules of a media curation system according to certain example embodiments.
[0010] Figure 4 This is a flowchart depicting a method for curating a collection of media content based on input received at a client device, according to certain example embodiments.
[0011] Figure 5 This is a flowchart depicting a method for generating a custom context filter according to certain example embodiments.
[0012] Figure 6 This is a flowchart depicting a method for generating a custom context filter according to certain example embodiments.
[0013] Figure 7It is an interface flowchart depicting the interface presented by a media curation system according to certain example embodiments.
[0014] Figure 8 It is an interface flowchart depicting the interface presented by a media curation system according to certain example embodiments.
[0015] Figure 9 It is a diagram depicting a collection of input-based media content curated according to certain example embodiments.
[0016] Figure 10 This is a flowchart depicting a method for curating a collection of media content based on input, including images and input context, according to certain example embodiments.
[0017] Figure 11 This is a block diagram illustrating a representative software architecture that can be used in conjunction with the various hardware architectures described herein and to implement various embodiments.
[0018] Figure 12 This is a block diagram illustrating components of a machine according to some example embodiments, the components of which are capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more methods discussed herein. Detailed Implementation
[0019] As described above, AR systems provide users with real-time direct or indirect maps of the physical, real-world environment displayed within a graphical user interface (GUI), where elements in the map are augmented by computer-generated sensory input. For example, an AR interface can present media content at a location within the display of a view of the real-world environment, making the media content appear to interact with elements in the real-world environment. Similarly, media overlays or “lenses” comprise a set of media items that can be overlaid or filtered onto media content presented on a client device, and then modified or transformed in some way. For example, AR content from a lens can be used to make complex additions or transformations to media content presented on a client device, such as adding rabbit ears to a person's head in a video or image, adding floating hearts and stars to a video or image, changing the proportions of features of people or objects in a video or image, or many other such transformations. Transformations include real-time modifications to images or videos as the client device captures them and displays them on a screen, as well as modifications to stored content (such as video clips in a gallery or repository accessible to the client device).
[0020] The example embodiments described herein relate to a context-based media curation system that determines a user context based on one or more inputs received at a client device and curates a collection of media content based on the user context, wherein the collection of media content may include auditory content, video content, images, and AR content, including footage. According to some embodiments, the media curation system is configured to perform operations including: receiving input at a client device, wherein the input includes an input context and an image including multiple image features; identifying a category based on the multiple image features of the image; generating a query based on the category and the input context; querying a media repository based on the query, wherein the media repository includes a set of media items; and curating the collection of media content based on the set of media items in the media repository. The media curation system results in the presentation of the collection of media content at the client device, wherein the presentation of the collection of media content can be navigated by a user of the client device.
[0021] Media content can include animated graphics exchange format (GIF) images of various shapes, sizes, and themes, as well as audio content (such as songs or sound effects) and media lenses, where media lenses include AR content rendered as media filters to enhance image data displayed on client devices. In some embodiments, the media curation system can communicate with a media repository that includes a collection of categorized and tagged media content, wherein the media content within the collection is tagged or labeled based on attributes of the media content or based on association with user-specified tags or labels. For example, media content can be tagged with labels that identify object categories of the media content, such as “food,” “basketball,” “morning,” “April,” such that a reference to an object category corresponds to a set of media content in the media content collection.
[0022] In some embodiments, the input includes an image or video image comprising a depiction of one or more objects in a real-world environment, and metadata such as location and time data that identifies one or more contextual inputs associated with the input. In response to receiving input including an image, the media curation system identifies the object category of one or more objects depicted within the image. For example, the media curation system may detect one or more Quick Response (QR) codes within the image, wherein the QR codes identify objects or object categories associated with objects depicted within the image, or in further embodiments, one or more image and text recognition techniques may be employed to identify objects depicted in the image. For example, the media curation system may employ a machine learning model or neural network trained on training data based on labeled data including multiple image features to identify object categories.
[0023] Based on object or object category identification, the media curation system obtains a set of tags or labels associated with the object or object category, and generates a query based on this set of tags or labels and the input context. The media curation system then queries a media repository to identify a collection of media content based on this query.
[0024] In response to curating a collection of media content from a media repository, the media curation system prompts the presentation of the collection of media content on a client device. In some embodiments, the presentation of the collection of media content includes the vertical or horizontal arrangement of the media content within the collection, allowing a user of the client device to navigate within the collection of media content.
[0025] In some embodiments, the media curation system generates customized media content to be presented in the presentation of a set of media content based on a query and input context. In response to receiving input including an image, the media curation system selects a media template based on multiple image features of the image, wherein the media template defines the display configuration of a set of media items within the set of media content in the image. The media template may define the location for presenting a portion of a set of media content within the image on the client device. In some embodiments, the presentation of a set of media content within the image on the client device may be based on the location (or multiple locations) of multiple image features. For example, in the context of a shot, the AR content of the shot may be presented by detecting objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.), tracking these objects as they leave, enter, and move within the field of view of an image or video frame, and applying modifications based on the AR content of the shot while tracking the objects.
[0026] In some embodiments, in response to receiving a lens selection, the element to be lens-transformed is identified, then detected and tracked. Image features corresponding to the object are modified according to the AR content of the selected lens, thereby transforming frames of the video stream. Frame transformation of the video stream can be performed using different methods for different types of transformations.
[0027] Figure 1 This is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and related content) over a network. Messaging system 100 includes one or more client devices 102, each hosting multiple applications including messaging client applications 104. Each messaging client application 104 is communicatively coupled to other instances of messaging client applications 104 and a messaging server system 108 via a network 106 (e.g., the Internet).
[0028] Therefore, each messaging client application 104 can communicate and exchange data with another messaging client application 104 and the messaging server system 108 via network 106. The data exchanged between messaging client applications 104 and between messaging client applications 104 and messaging server system 108 includes functions (e.g., commands to call functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0029] Messaging server system 108 provides server-side functionality to a specific messaging client application 104 via network 106. Although some functions of messaging system 100 are described herein as being performed by messaging client application 104 or by messaging server system 108, it should be understood that the location of certain functions within messaging client application 104 or messaging server system 108 is a design choice. For example, it is technically preferable to first deploy certain technologies and functions within messaging server system 108 and then migrate those technologies and functions to messaging client application 104, in which client device 102 has sufficient processing power.
[0030] The messaging server system 108 supports various services and operations provided to the messaging client application 104. These operations include sending data to and receiving data from the messaging client application 104, and processing data generated by the messaging client application 104. In some embodiments, as an example, this data includes message content, client device information, geolocation information, media annotations and overlays, message content persistence conditions, social network information, and live event information. In other embodiments, other data is used. Data exchange within the messaging system 100 is invoked and controlled via functions available through the GUI of the messaging client application 104.
[0031] Now, specifically to the messaging server system 108, the application programming interface (API) server 110 is coupled to the application server 112 and provides a programming interface to the application server 112. The application server 112 is communicatively coupled to the database server 118, which facilitates access to the database 120, in which data associated with the messages processed by the application server 112 is stored.
[0032] Specifically, Application Programming Interface (API) server 110 handles the receiving and sending of message data (e.g., commands and message payloads) between client device 102 and application server 112. Specifically, API server 110 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by messaging client application 104 to call functions of application server 112. API server 110 exposes various functions supported by application server 112, including: account registration; login functionality; sending messages from one messaging client application 104 to another messaging client application 104 via application server 112; sending media files (e.g., images or videos) from messaging client application 104 to messaging server application 114 for possible access by another messaging client application 104; setting up collections of media data (e.g., stories); retrieving the friend list of the user of client device 102; retrieving such collections; retrieving messages and content; adding and deleting friends in a social graph; the position of friends in the social graph; and opening application events (e.g., related to messaging client application 104).
[0033] Application server 112 hosts multiple applications and subsystems, including messaging server application 114, image processing system 116, social networking system 122, and media curation system 124. According to some example embodiments, media curation system 124 is configured to receive input including an input context and an image containing multiple image features; identify one or more objects depicted in the image based on the multiple image features; select one or more categories based on the one or more objects, where each of the one or more categories corresponds to a set of tags; generate a query based on the set of tags; query a media repository based on the query; and curate a collection of media content including a set of media items from the media repository. Further details of media curation system 124 can be found below. Figure 3 Found it.
[0034] Messaging server application 114 implements several messaging techniques and functions, particularly involving the aggregation and other processing of content (e.g., text and multimedia content) received from multiple instances of messaging client application 104. As will be described in further details, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). Messaging server application 114 then makes these collections available to messaging client application 104. Considering the hardware requirements of such processing, additional processor- and memory-intensive processing of the data can also be performed on the server side by messaging server application 114.
[0035] Application server 112 also includes an image processing system 116, which is dedicated to performing various image processing operations, typically relating to images or videos received within the payload of messages at messaging application server 114.
[0036] Social networking system 122 supports various social networking functions and services, and makes these functions and services available to messaging server application 114. To this end, social networking system 122 maintains and accesses entity graph 304 within database 120. Examples of functions and services supported by social networking system 122 include identifying other users of messaging system 100 who have relationships with a particular user or who are "following" that particular user, as well as other entities and interests of a particular user.
[0037] Application server 112 is communicatively coupled to database server 118, which facilitates access to database 120, which stores data associated with messages processed by messaging server application 114.
[0038] Figure 2 This is a block diagram illustrating further details of a messaging system 100 according to an example embodiment. Specifically, the messaging system 100 is shown to include a messaging client application 104 and an application server 112, which in turn embody several specific subsystems, namely, a short-timer system 202, a collection management system 204, and an annotation system 206.
[0039] The short-timer system 202 is responsible for performing temporary access to content permitted by the messaging client application 104 and the messaging server application 114. To this end, the short-timer system 202 incorporates multiple timers that selectively display and enable access to messages and associated content via the messaging client application 104 based on duration and display parameters associated with messages, message sets (e.g., SNAPCHAT stories), or graphical elements. Further details regarding the operation of the short-timer system 202 are provided below.
[0040] The collection management system 204 is responsible for managing collections of media (e.g., text, image, video, and audio data). In some examples, collections of content (e.g., messages including images, videos, text, and audio) may be organized into "event galleries" or "event stories." Such collections may be available for a specified time period (e.g., the duration of an event related to the content). For example, content related to a concert may be available as a "story" for the duration of the concert. The collection management system 204 may also be responsible for publishing icons that notify the user interface of the messaging client application 104 of the existence of a specific collection.
[0041] The collection management system 204 also includes a curatorial interface 208, which allows collection managers to manage and curate specific collections of content. For example, the curatorial interface 208 enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some embodiments, compensation (e.g., currency, non-monetary credits or points associated with a messaging system or third-party reward system, travel miles, access to artwork or special lenses, etc.) may be paid to users for including user-generated content in the collection. In this case, the curatorial interface 208 operates to automatically pay such users for using its content.
[0042] Annotation system 206 provides various functionalities that enable users to annotate or otherwise modify or edit media content associated with messages. For example, annotation system 206 provides functionality related to media overlays generated and published for messages processed by messaging system 100. Annotation system 206 operatively provides media overlays (e.g., SNAPCHAT filters, lenses) to messaging client application 104 based on the geographic location of client device 102. In another example, annotation system 206 operatively provides media overlays to messaging client application 104 based on other information (e.g., the social network information of the user of client device 102). Media overlays may include audio and visual content as well as visual effects. Examples of audio and visual content include images, text, logos, animations and sound effects, and animated facial models, such as those generated by media curation system 124. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photographs) at client device 110. For example, media overlays may include text that can be overlaid on a photograph taken by client device 102. In another example, media overlays may include identifiers for location overlays (e.g., Venice Beach), names of on-site events, or names of merchant overlays (e.g., Beach Cafe). In yet another example, annotation system 206 uses the geographic location of client device 102 to identify media overlays that include the name of a merchant at that geographic location. Media overlays may include additional tags associated with the merchant. Media overlays may be stored in database 120 and accessible via database server 118.
[0043] In one example embodiment, annotation system 206 provides a user-based publishing platform that allows users to select geographic locations on a map and upload content associated with those locations. Users can also specify under what circumstances a particular media overlay should be made available to other users. Annotation system 206 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geographic location.
[0044] In another example embodiment, annotation system 206 provides a merchant-based publishing platform that enables merchants to select specific media overlays associated with geographic locations via a bidding process. For example, annotation system 206 associates the media overlay of the highest bidder with the corresponding geographic location within a predetermined time period.
[0045] Figure 3 This is a block diagram illustrating the components of a media curation system 124, which, according to certain example embodiments, is configured to perform operations to receive input including an input context and an image containing multiple image features, identify one or more objects depicted in the image based on the multiple image features, select one or more categories based on one or more objects, wherein each of the one or more categories corresponds to a set of tags, generate a query based on the set of tags, query a media repository based on the query, and curate a collection of media content including a set of media items from the media repository.
[0046] In a further embodiment, the components of the media curation system 124 may be configured to perform operations, according to some example embodiments, to acquire an image including a depiction of an object from a client device 102, to identify one or more objects or object categories within the image based on the depiction of the object, to select one or more markers or tags based on the object or object category, to obtain a set of media content based on the markers or tags, and to cause the presentation of the set of media content to be displayed within the image on the client device.
[0047] The media curation system 124 is shown to include a presentation module 302, a media module 304, a communication module 306, and an identification module 308, all of which are configured to communicate with each other (e.g., via a bus, shared memory, or a switch). Any or more of these modules may be implemented using one or more processors 310 (e.g., by configuring one or more such processors to perform the functions described for that module), and thus may include one or more processors 310.
[0048] Any one or more of the described modules can be implemented using individual hardware (e.g., one or more processors 310 of a machine) or a combination of hardware and software. For example, any module described in the media curation system 124 may physically comprise an arrangement of one or more processors 310 (e.g., a subset of one or more processors of a machine or a subset thereof) configured to perform the operations described herein for that module. As another example, any module of the media curation system 124 may include software, hardware, or both, that configures an arrangement of one or more processors 310 (e.g., among one or more processors of a machine) to perform the operations described herein for that module. Thus, different modules of the media curation system 124 may include and configure different arrangements of such processors 310 or a single arrangement of such processors 310 at different points in time. Furthermore, any two or more modules of the media curation system 124 may be combined into a single module, and the functionality described herein for a single module may be subdivided among multiple modules. Moreover, modules described herein as being implemented within a single machine, database, or device may be distributed across multiple machines, databases, or devices, according to various example embodiments.
[0049] Figure 4 This is a flowchart depicting a method 400 for curating a collection of media content based on input received at a client device, according to certain example embodiments. The operation of method 400 can be described by referring to the above. Figure 3 The module described is executed. For example... Figure 4 As shown, method 400 includes one or more operations 402, 404, 406, 408 and 410.
[0050] In operation 402, the presentation module 302 receives input at the client device 102, wherein the input includes an input context and image data including multiple image features. For example, the input context may include metadata, which includes location data identifying the location of the client device 102 and time data indicating the time or date.
[0051] In some embodiments, the context data may also include user profile information associated with the user profile of the client device 102. For example, the user profile data may include identifiers of one or more user similarities (i.e., "likes") associated with the user, and identifiers of one or more network connections (i.e., friends) of the user.
[0052] In some embodiments, receiving input at client device 102 may include the operation of acquiring an image at client device 102, wherein the image includes a depiction of an object at a certain location within the image. For example, presentation module 302 may activate the camera of client device 102 and cause the camera of client device 102 to acquire an image.
[0053] In operation 404, the identification module 308 identifies one or more objects depicted within the image based on multiple image features of the image data. In some embodiments, to identify one or more objects, the identification module 308 may utilize computer vision to perform one or more image or pattern recognition techniques. In a further embodiment, the identification module 308 may identify one or more QR codes within the image and identify objects based on the QR codes.
[0054] Based on one or more objects identified in an image, the identification module 308 determines one or more object categories corresponding to the one or more objects, wherein the one or more object categories are associated with a set of labels.
[0055] In operation 406, media module 304 accesses a media repository, such as database 120, to identify a set of media content based on at least one or more object categories and input context. In some embodiments, in response to identifying one or more objects depicted within an image based on multiple image features of image data, media module 304 generates a query based on a tag associated with each of the one or more objects. For example, media module 304 may identify a first object and a second object, where the first object corresponds to a first media tag (or a first set of media tags), and the second object corresponds to a second media tag (or a second set of media tags). Media module 304 may generate a query by combining the first media tag, the second media tag, and the input context, and query the media repository (i.e., database 120) based on this query.
[0056] In operation 408, media module 304 curates a set of media content based on media content accessed within a media repository, tags associated with categories corresponding to one or more objects identified by identification module 308, and the input context of client device 102. For example, in some embodiments, media module 304 accesses media content associated with one or more objects within a media repository (e.g., database 120). Identification module 308 may select one or more tags or markers based on the identifiers of objects within an image, and cause media module 304 to query the media repository based on the selected tags or markers. Therefore, media module 304 can access a set of media content associated with identified objects within an image by referencing tagged media content within the media repository using the selected tags or markers.
[0057] At operation 408, presentation module 302 causes the presentation of a collection of media content to be displayed at client device 102. In some embodiments, the presentation of the collection of media content includes a navigable list of media items displayed at client device 102. For example, the presentation of the collection of media content may be horizontal or vertical at client device 102, e.g. Figure 8 The collection of media content is 825.
[0058] In some embodiments, the presentation of a collection of generated media content may include ranking the collection of media content based on corresponding tags, queries, and input context.
[0059] In some embodiments, the presentation of media content may include the AR display of one or more media items at locations within an image based on multiple features of image data. For example, such as Figure 7 As shown, the presentation of media content may include the presentation of media content 735, as depicted in interface 715, wherein the presentation of a set of media content 735 includes multiple media items, which include GIFs and images related to objects detected based on multiple image features of image data.
[0060] In some embodiments, in order to generate a presentation of media content, such as a presentation of a set of media content 735, the presentation module 302 obtains a media template, wherein the media template defines the presentation format and layout to be applied to the set of media content. For example, the media template may define the position and orientation of the media content to be presented within an image at the client device 102. For example, the media template may be based on one or more image attributes (including multiple image features) and the input context.
[0061] In some embodiments, in order to generate a presentation of media content to be displayed at the client device 102, the media module 304 provides the client device 102 with an identifier for each media content in a set of media content, so that the client device 102 can identify the relevant media content in a local storage repository (e.g., database 120).
[0062] Figure 5 This is a flowchart depicting a method 500 for generating a custom context filter according to certain example embodiments. The operation of method 500 can be described by the above references. Figure 3 The module described is used for execution. For example... Figure 5 As shown, method 500 includes one or more operations 502, 504 and 506.
[0063] In some embodiments, context filters may include media overlays or “lenses,” wherein the media overlay includes AR content to be rendered over images and videos at client device 102. As an illustrative example, a user of media curation system 124 may activate the camera of client device 102, resulting in the display of an image depicting a real-world environment. The image may include pictures and videos (real-time or pre-recorded). In response to the display of the image, media curation system 124 may generate custom context filters (i.e., media overlays, lenses) based on the input context of client device 102 and one or more objects detected within the image.
[0064] At operation 502, media module 304 obtains a media template based on input received at the client device, wherein the input includes an image containing a set of image features and an input context. For example, in response to receiving input, media module 304 accesses a template repository to obtain a media template based on at least a portion of multiple image features and the input context.
[0065] In operation 504, media module 304 generates a context filter (i.e., a shot) based on at least a portion of a set of media content curated according to multiple image features and input context, and a media template. For example, in some embodiments, media module 304 may select a portion of media content from the set of media content based on multiple image features, input context, and media template. Media module 304 populates the template with a portion of media content from the set of media content. For example, the media template may define the display instructions for media items within the template based on attributes of the media items themselves.
[0066] At operation 506, presentation module 302 causes the context filter to be displayed on client device 102, such as... Figure 7 As seen in interface 715. In some embodiments, presentation module 302 can present context filters within a context filter set. The presentation of the media content set allows the user to provide input to select a context filter. For example, as... Figure 8 As shown, graphic icons representing context filters can be displayed between media content sets 825.
[0067] Figure 6 This is a flowchart depicting a method 500 for generating a context filter to be displayed at client device 102 according to certain example embodiments. The operation of method 600 can be described by the above references. Figure 3 The module described is used for execution. For example... Figure 6 As shown, method 600 includes one or more operations 602, 604 and 606.
[0068] In operation 602, as in operation 402, the rendering module 302 receives input including an image that depicts objects. For example, the image may include multiple image features corresponding to one or more objects depicted at locations in the image. For example, the rendering module 302 may activate the camera of the client device 102 and cause the camera of the client device 102 to acquire an image. In some embodiments, the image acquired by the rendering module 302 may include image metadata containing location data, time data, and device data of the client device 102.
[0069] In operation 604, the identification module 308 determines the context of the client device 102 in response to received input. For example, the context may include the location of the client device, the time of day when images were acquired, and the device type of the client device 102.
[0070] In some embodiments, the identification module 308 can parse the image's metadata to determine relevant contextual information from the metadata's location data, time data, and device data.
[0071] In operation 608, media module 304 accesses media content within a media repository based on the context of client device 102 and one or more objects depicted in an image from client device 102. For example, media content may correspond to one or more tags based on object categories and to location or time information within the media repository, such that references to a specific time of day, season, day of the week, month, or location can identify a set of related media content.
[0072] Figure 7 This is an interface flowchart 700 depicting an interface presented by a media curation system 124 according to certain example embodiments. The operations depicted by the interface in flowchart 700 can be derived from the above description... Figure 3 The module described is used for execution.
[0073] Interface 705 depicts an image captured by client device 102. For example... Figure 7 As seen, interface 705 includes a depiction of object 720 located within interface 705.
[0074] In some embodiments, client device 102 may activate media curation system 124 in response to receiving user input selecting user option 725 and cause media curation system 124 to capture images depicted within interface 705, such as those displayed within interface 705.
[0075] In response to receiving input selecting user option 725, media curation system 124 may cause graphic icon 730 to be displayed in interface 710 to indicate that media curation system 124 has been activated.
[0076] Interface 715 includes the presentation of a media overlay (i.e., a shot), comprising a set of media content 735 from a collection of media content curated based on one or more objects (e.g., object 720) depicted within interface 705 and displayed within an image captured by client device 102. As seen in interface 715, the set of media content 735 of the media overlay may be displayed at a location within the image based on the position of object 720, as defined by a media template. As seen in interface 715, the presentation of the set of media content 735 may include multiple media items, including images and GIFs associated with object 720.
[0077] For example, such as Figure 7 As shown, object 720 is a bag of potato chips. The media curation system 124 identifies the object category of object 720 (e.g., food, snacks, etc.) and obtains a set of media content 735, wherein the set of media content 735 includes media content labeled or marked with the object category of object 720.
[0078] As discussed in methods 400, 500, and 600, the location of each media content in a set of media content 735 within an image depicted in interface 705 can be determined based on a media template. A user of client device 102 can thus generate a message comprising a set of media content 735 to be distributed to one or more recipients identified by the user of client device 102. In some embodiments, the message may include a short message.
[0079] Figure 8 This is an interface flowchart 800 depicting an interface presented by a media curation system 124 according to certain example embodiments. The operations depicted in the interface of flowchart 800 can be described above regarding... Figure 3 The module described is used for execution.
[0080] Interface 805 depicts input received at client device 102. According to some embodiments, the input may include image or video data that can be streamed from a camera associated with client device 102. For example, a user of client device 102 may activate the camera of client device 102, causing client device 102 to display image data.
[0081] In response to receiving input at client device 102, as seen in interface 810, media curation system 124 presents indicator 820 in response to detecting a set of features corresponding to one or more objects depicted within an image. In some embodiments, indicator 820 may vary based on object category or type. For example, as shown in interface 810, indicator 820 includes a display of a set of musical notes indicating one or more objects corresponding to auditory content detectable within the image presented in interfaces 805 and 810.
[0082] In response to the collection of curatorial media content, such as Figure 4 In operation 408 of the method 400 described herein, the media curation system 124 causes the presentation of a media content set 825 to be displayed at a location within an image displayed in interfaces 805, 810, and 815. In some embodiments, the presentation of the media content set 825 may further include the display of a result indicator 830, which includes an indication of the number of media items in the media content set 825.
[0083] Figure 9 Figure 900 depicts a collection of media content curated based on input 905 according to certain example embodiments. For example... Figure 9 As shown, the collection of media content can include media content 910, 915, and 920, where each media content corresponds to tags 925, 930, and 935.
[0084] like Figure 4 The method described in section 400, operation 402 and Figure 6 As described in operation 602 of the method 600 depicted in the figure, input 905 may include image 940 (or a frame of video) and context input 945, wherein context input 945 may include one or more of user profile data, time data, location data and device data.
[0085] In response to receiving input 905, the media curation system 124 curates a collection of media content based on multiple image features of image 940 and input content 945. For example, as seen in Figure 900, image 940 includes a depiction of an object (a cup of coffee), and the input context includes time data (i.e., 9:30 AM).
[0086] The media curation system 124 generates a query based on one or more tags associated with the object's category and the input context 945 of input 905. As shown in Figure 900, tags may include tags 925, 930, and 935. Figure 10 As described in operation 1008 of method 1000, media curation system 124 generates a query that includes a set of query items, wherein the query items are based on one or more labels associated with the object category of the object depicted in the input (i.e., the image) and the input context (i.e., time of day, location, device type, user profile information).
[0087] In some embodiments, media content 910, 915, and 920 may be associated with corresponding tags 925, 930, and 935 within a media repository (e.g., database 120). In some embodiments, users of the media curation system 124 may create and assign tags to media content within the media repository.
[0088] Figure 10 This is a flowchart depicting a method 1000 for curating a collection of media content based on input including images and input context, according to certain example embodiments. The operation of method 1000 can be described by the above references. Figure 3 The module described is used for execution. For example... Figure 10 As shown, method 1000 includes one or more operations 1002, 1004, 1006, 1008 and 1010.
[0089] In operations 1002 and 1004, the identification module 308 identifies a first object and a second object depicted in the image based on a first subset and a second subset of multiple image features of the image. For example, in response to receiving input including the image and the input context, the identification module 308 analyzes multiple image features of the image to identify one or more objects depicted within the image.
[0090] In response to detecting a first object and a second object based on multiple image features, the media module 304 selects a first category corresponding to the first object and a second category corresponding to the second object. The first category and the second category may each include a corresponding set of labels (i.e., tags).
[0091] In operation 1008, media module 304 generates a query that includes a set of query items based on the first category and the second category. For example, the query may include a set of tags associated with the first category and the second category, as well as one or more tags associated with the input context. Based on this query, in operation 1010, media module 304 queries a media repository (e.g., database 120) to identify a collection of media content.
[0092] In some embodiments, media items in the collection of media content may include images, videos, audio content, and shots and media overlays, wherein shots and media overlays include content containing AR content to be presented at the client device 102.
[0093] Software Architecture
[0094] Figure 11 This is a block diagram illustrating an example software architecture 1106 that can be used in conjunction with various hardware architectures described herein. Figure 11 This is a non-limiting example of a software architecture, and it should be understood that many other architectures can be implemented to facilitate the functionality described herein. Software architecture 1106 can be implemented in, for example... Figure 12 The machine 1200 executes on hardware, which, among other things, includes a processor 1204, a memory 1214, and I / O components 1218. A representative hardware layer 1152 is shown and can represent, for example... Figure 11The machine 1100. A representative hardware layer 1152 includes a processing unit 1154 having associated executable instructions 1104. The executable instructions 1104 represent executable instructions of the software architecture 1106, including implementations of the methods, components, etc., described herein. Hardware layer 1152 also includes memory and / or storage modules, memory / storage devices 1156, which also have the executable instructions 1104. Hardware layer 1152 may also include other hardware 1158.
[0095] exist Figure 11 In the example architecture, software architecture 1106 can be conceptualized as a stack of layers, where each layer provides specific functionality. For example, software architecture 1106 may include layers such as operating system 1102, library 1120, application 1116, and presentation layer 1114. Operationally, application 1116 and / or other components within these layers can invoke application programming interface (API) API calls 1108 via the software stack and receive responses in response to API calls 1108. The layers shown are representative in nature, and not all software architectures have all layers. For example, some mobile or dedicated operating systems may not provide a framework / middleware 1118, while other operating systems may provide such a layer. Other software architectures may include other or different layers.
[0096] Operating system 1102 manages hardware resources and provides general services. Operating system 1102 may include, for example, a kernel 1122, services 1124, and drivers 1126. Kernel 1122 can serve as an abstraction layer between hardware and other software layers. For example, kernel 1122 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. Services 1124 may provide other public services to other software layers. Drivers 1126 are responsible for controlling or interfacing with underlying hardware. For example, depending on the hardware configuration, drivers 1126 may include display drivers, camera drivers, Bluetooth® drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi® drivers, audio drivers, power management drivers, etc.
[0097] Library 1120 provides common infrastructure used by application 1116 and / or other components and / or layers. Library 1120 provides functionality that allows other software components to perform tasks more easily than by directly interfaceing with the underlying operating system 1102 functions (e.g., kernel 1122, services 1124, and / or drivers 1126). Library 1120 may include system libraries 1144 (e.g., the C standard library), which provide functions such as memory allocation, string manipulation, mathematical functions, etc. Furthermore, library 1120 may include API libraries 1146, such as media libraries (e.g., libraries supporting the rendering and manipulation of various media formats such as MPREG4, H.264, MP3, AAC, AMR, JPG, PNG), graphics libraries (e.g., OpenGL frameworks for rendering 2D and 3D content on a display), database libraries (e.g., SQLite providing various relational database functions), network libraries (e.g., WebKit providing web browsing functionality), etc. Library 1120 may also include a wide variety of other libraries 1148 to provide many other APIs to application 1116 and other software components / modules.
[0098] The framework / middleware 1118 (sometimes also called middleware) provides a higher level of common infrastructure that can be used by application 1116 and / or other software components / modules. For example, the framework / middleware 1118 can provide various graphical user interface (GUI) functions, advanced resource management, advanced location services, etc. The framework / middleware 1118 can provide a wide range of other APIs that can be used by application 1116 and / or other software components / modules, some of which may be specific to a particular operating system 1102 or platform.
[0099] Application 1116 includes built-in application 1138 and / or third-party application 1140. Examples of representative built-in applications 1138 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, and / or game applications. Third-party applications 1140 may include applications developed by entities other than platform-specific vendors using the Android™ or iOS™ Software Development Kit (SDK), and may be mobile software running on a mobile operating system such as iOS™, Android™, Windows® Phone, or other mobile operating systems. Third-party applications 1140 may invoke API calls 1108 provided by the mobile operating system (e.g., operating system 1102) to facilitate the functionality described herein.
[0100] Application 1116 can use built-in operating system functions (e.g., kernel 1122, service 1124, and / or driver 1126), libraries 1120, and frameworks / middleware 1118 to create a user interface to interact with the system's user. Alternatively, or additionally, in some systems, interaction with the user can be achieved through a presentation layer (e.g., presentation layer 1114). In these systems, the application / component "logic" can be separated from the various aspects of the application / component that interact with the user.
[0101] Figure 12 This is a block diagram illustrating components of a machine 1200, according to some example embodiments, capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and executing any one or more of the methods discussed herein. Specifically, Figure 12 A schematic representation of a machine 1200 in an example form having a computer system is shown, in which instructions 1210 (e.g., software, programs, applications, applets, or other executable code) are executable to cause the machine 1200 to perform any or more methods discussed herein. Thus, instructions 1210 can be used to implement the modules or components described herein. Instructions 1210 transform a generic, unprogrammed machine 1200 into a specific machine 1200 programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, machine 1200 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, machine 1200 may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1200 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, networked appliances, network routers, network switches, bridges, or any machine capable of executing instructions 1210 sequentially or otherwise, specifying the actions that machine 1200 will take. Furthermore, although only a single machine 1200 is shown, the term "machine" should also be considered to include a collection of machines that individually or collectively execute instructions 1210 to perform any or more of the methods discussed herein.
[0102] Machine 1200 may include processor 1204, memory / storage device 1206, and I / O components 1218, which may be configured to communicate with each other, for example, via bus 1202. Memory / storage device 1206 may include memory 1214 (such as main memory or other memory) and storage cell 1216, both of which may be accessed by processor 1204, such as via bus 1202. Storage cell 1216 and memory 1214 store instructions 1210 embodying any one or more of the methods or functions described herein. Instructions 1210 may also reside wholly or partially within memory 1214, storage cell 1216, at least one processor of processor 1204 (e.g., within the processor's cache memory), or any suitable combination thereof during execution by machine 1200. Thus, the memory of memory 1214, storage cell 1216, and processor 1204 are examples of machine-readable media.
[0103] I / O component 1218 may include a wide variety of components for receiving input, providing output, generating output, sending information, exchanging information, acquiring measurements, etc. The specific I / O component 1218 included in a particular machine 1200 will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that I / O component 1218 may be included in... Figure 12 Several other components are not shown. For the purpose of simplifying the discussion below, the I / O components 1218 are grouped according to function, and this grouping is by no means limiting. In various example embodiments, the I / O components 1218 may include output components 1226 and input components 1228. Output components 1226 may include visual components (e.g., displays, such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), auditory components (e.g., speakers), haptic components (e.g., vibration motors, reactive mechanisms), other signal generators, etc. Input components 1228 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens providing the location and force of a touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.
[0104] In a further example embodiment, I / O component 1218 may include, among a variety of other components, a biometric component 1230, a motion component 1234, an environmental component 1236, or a position component 1238. For example, biometric component 1230 may include components for detecting expressions (e.g., hand gestures, facial expressions, voice expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), and identifying a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1234 may include an accelerometer component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. Environmental component 1236 may include, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of harmful gases or measuring pollutants in the atmosphere for safety purposes), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Position component 1238 may include a position sensor component (e.g., a Global Positioning System (GPS) receiver component), an altitude sensor component (e.g., an altimeter or barometer for detecting from which altitude the air pressure can be obtained), an orientation sensor component (e.g., a magnetometer), etc.
[0105] Various technologies can be used to implement communication. I / O component 1218 may include communication component 1240, which is operable to couple machine 1200 to network 1232 or device 1220 via coupling 1222 and coupling 1224, respectively. For example, communication component 1240 may include a network interface component or other suitable device interfaced with network 1232. In a further example, communication component 1240 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components that provide communication via other forms. Device 1220 may be another machine or any of a wide variety of peripheral devices, such as peripheral devices coupled via Universal Serial Bus (USB).
[0106] Furthermore, the communication component 1240 may detect identifiers or include components operable to detect identifiers. For example, the communication component 1240 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes (e.g., Quick Response (QR) codes, Aztec codes, data matrices, digital graphics, maximum codes, PDF417, super codes, UCCRSS-2D barcodes), and other optical codes), or an acoustic detection component (e.g., a microphone for identifying the tagged audio signal). Additionally, various types of information can be obtained via the communication component 1240, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via detecting NFC beacon signals indicating a specific location, etc.
[0107] Glossary
[0108] In this context, "carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions to be executed by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions can be sent or received over a network via a transmission medium using a network interface device and employing any of a number of well-known transmission protocols.
[0109] In this context, "client device" refers to any machine that is connected to a communication network interface to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to: mobile phones, desktop computers, portable computers, portable digital assistants (PDAs), smartphones, tablets, ultrabooks, netbooks, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.
[0110] In this context, "communication network" refers to one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless local area network (WLAN), wide area network (WAN), wireless wide area network (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a POTS network, a cellular telephone network, a wireless network, a Wi-Fi® network, another type of network, or a combination of two or more such networks. For example, a network or a part of a network may include a wireless or cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, coupling can enable any of a variety of data transmission technologies, such as single-carrier radio transmission technology (1xRTT), evolved data optimization (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standard, other standards defined by various standards-setting organizations, other remote protocols, or other data transmission technologies.
[0111] In this context, a "short-lived message" refers to a message that is accessible for a limited time. Short-lived messages can be text, images, videos, etc. The access time for a short-lived message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is short-lived.
[0112] In this context, "machine-readable medium" means a component, device, or other tangible medium capable of temporarily or permanently storing instructions and data, and may include, but is not limited to: random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage devices (e.g., erasable programmable read-only memory (EEPROM)) and / or any suitable combination thereof. The term "machine-readable medium" should be understood to include a single medium or multiple media capable of storing instructions (e.g., a centralized or distributed database, or associated caches and servers). The term "machine-readable medium" should also be understood to include any medium or combination of media capable of storing machine-executable instructions (e.g., code) such that, when executed by one or more processors of the machine, these instructions cause the machine to perform any or more methods described herein. Therefore, "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network comprising multiple storage devices or apparatuses. The term "machine-readable medium" does not include the signal itself.
[0113] In this context, a “component” refers to a device, physical entity, or logic having boundaries defined by functional or subroutine calls, branch points, application programming interfaces (APIs), or other partitions or modular technologies for specific processing or control functions. Components can be combined through their interfaces with other components to perform machine processes. A component can be an encapsulated functional hardware unit designed to be used with other components, or part of a program that typically performs a specific function related to that function. Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example embodiments, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or an application portion) to operate to perform certain operations described herein. Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic permanently configured to perform certain operations. Hardware components can be dedicated processors, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). Hardware components may also include programmable logic or circuitry temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine) specifically tailored to perform the configured function and is no longer a general-purpose processor. It should be understood that the decision to implement hardware components mechanically in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software) may be made for cost and time considerations. Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include tangible entities that are physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein. Consider embodiments where hardware components are temporarily configured (e.g., programmed), each hardware component does not need to be configured or instantiated at any instance of time. For example, where the hardware components include a general-purpose processor configured by software to be a dedicated processor, the general-purpose processor can be configured as different dedicated processors (e.g., including different hardware components) at different times. Therefore, the software accordingly configures one or more specific processors, for example, to constitute a specific hardware component at one time instance and different hardware components at different time instances. Hardware components can provide information to or receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled.In the presence of multiple hardware components, communication can be achieved through signal transmission between two or more hardware components (e.g., on appropriate circuitry and buses). In embodiments where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, by storing and retrieving information in a memory structure accessible to the multiple hardware components. For example, one hardware component may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. Another hardware component may then access the memory device at a later time to retrieve and process the stored output. Hardware components may also initiate communication with input or output devices and may operate on resources (e.g., collections of information). The various operations of the example methods described herein can be performed at least in part by one or more processors that are temporarily (e.g., via software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented at least in part by processors, where one or more specific processors are examples of hardware. For example, at least some operations of the method may be performed by one or more processors or processor-implemented components. Furthermore, one or more processors may also be operable to support the performance of the relevant operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations may be performed by a group of computers (as an example of a machine including processors), and these operations may be accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application programming interfaces (APIs)). The performance of some operations may be distributed among processors, residing not only in a single computer but also deployed across multiple computers. In some example embodiments, the processor or processor-implemented component may reside in a single geographic location (e.g., in a home environment, office environment, or server farm). In other example embodiments, the processor or processor-implemented component may be distributed across multiple geographic locations.
[0114] In this context, a "processor" refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values according to control signals (e.g., "commands," "opcodes," "machine codes," etc.) and generates corresponding output signals for operating a machine. A processor can be, for example, a Central Processing Unit (CPU), a Reduced Instruction Set Computing (RISC) processor, a Complex Instruction Set Computing (CISC) processor, a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Radio Frequency Integrated Circuit (RFIC), or any combination thereof. A processor can further be a multi-core processor having two or more independent processors (sometimes called "cores") capable of executing instructions simultaneously.
[0115] In this context, a "timestamp" refers to a sequence of characters or encoded information that identifies when a specific event occurred, such as giving a date and time, sometimes accurate to a fraction of a second.
[0116] In this context, "lifting force (LIFT)" is a measure relative to a randomly selected target model, as it is a measure of the performance of a target model in prediction or classification scenarios with an enhanced response (relative to the overall population).
[0117] In this context, "phoneme alignment" refers to a phoneme, which is a phonological unit that distinguishes one word from another. A phoneme may consist of a series of closing, plosive, and inspiratory events; or, liaison may transition from a back vowel to a front vowel. Therefore, a speech signal can be described not only by the phonemes it contains but also by the position of the phonemes. Thus, phoneme alignment can be described as the "time alignment" of phonemes in a waveform to determine the proper order and position of each phoneme in the speech signal.
[0118] In this context, "audio-to-visual conversion" refers to the conversion of audible speech signals into visible speech, where visible speech may include mouth shapes that represent audible speech signals.
[0119] In this context, "Time-Delay Neural Network (TDNN)" refers to an artificial neural network architecture whose primary purpose is to process sequential data. An example is converting continuous audio into a stream of categorized phoneme labels for speech recognition.
[0120] In this context, "Bidirectional Long Short-Term Memory (BLSTM)" refers to a recurrent neural network (RNN) architecture that can remember values over arbitrary time intervals. The stored values are not modified as learning progresses. RNNs allow forward and backward connections between neurons. Given time delays of unknown size and duration between events, BLSTMs are well-suited for classifying, processing, and predicting time series data.
Claims
1. A method comprising: This results in an image containing multiple image features being displayed on the client device; Based on the multiple image features of the image, identify the object depicted by the image; Based on the identified object, a category is selected, which corresponds to one or more media tags; Access a set of media items in the media repository based on the one or more media tags corresponding to the category; as well as This results in the presentation of the set of media items from the media repository at the client device.
2. The method according to claim 1, wherein, The presentation that causes the display of the set of media items to further include: In response to accessing the set of media items within the media repository, a notification is displayed; Receive input selecting the notification; and In response to the input that selects the notification, the presentation of the set of media items is caused.
3. The method according to claim 2, wherein, The notification includes a display of an indication of the number of media items in the set of media items.
4. The method according to claim 1, wherein, The presentation of the group of media items includes displaying the group of media items in a horizontal array.
5. The method according to claim 1, wherein, The presentation of the group of media items includes displaying the group of media items in a vertical array.
6. The method according to claim 1, wherein, Identifying the object depicted by the image based on the multiple image features further includes: Detect the multiple image features; and In response to the detection of the plurality of image features, an indicator is presented at a location within the image.
7. The method according to claim 6, wherein, The indicator includes graphical characteristics based on the category of the object.
8. A system comprising: Memory; as well as At least one hardware processor coupled to the memory and including instructions to cause the system to perform operations, said operations including: This results in an image containing multiple image features being displayed on the client device; Based on the multiple image features of the image, identify the object depicted by the image; Based on the identified object, a category is selected, which corresponds to one or more media tags; Access a set of media items within the media repository based on the one or more media tags corresponding to the category; and This results in the presentation of the set of media items from the media repository at the client device.
9. The system according to claim 8, wherein, The presentation that causes the display of the set of media items to further include: In response to accessing the set of media items within the media repository, a notification is displayed; Receive input selecting the notification; and In response to the input that selects the notification, the presentation of the set of media items is caused.
10. The system according to claim 9, wherein, The notification includes a display of an indication of the number of media items in the set of media items.
11. The system according to claim 8, wherein, The presentation of the group of media items includes displaying the group of media items in a horizontal array.
12. The system according to claim 8, wherein, The presentation of the group of media items includes displaying the group of media items in a vertical array.
13. The system according to claim 8, wherein, Identifying the object depicted by the image based on the multiple image features further includes: Detect the multiple image features; and In response to the detection of the plurality of image features, an indicator is presented at a location within the image.
14. The system according to claim 13, wherein, The indicator includes graphical characteristics based on the category of the object.
15. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of the machine, cause the machine to perform operations, the operations including: This results in an image containing multiple image features being displayed on the client device; Based on the multiple image features of the image, identify the object depicted by the image; Based on the identified object, a category is selected, which corresponds to one or more media tags; Access a set of media items in the media repository based on the one or more media tags corresponding to the category; as well as This results in the presentation of the set of media items from the media repository at the client device.
16. The non-transitory machine-readable storage medium according to claim 15, wherein, The presentation that causes the display of the set of media items to further include: In response to accessing the set of media items within the media repository, a notification is displayed; Receive input selecting the notification; and In response to the input that selects the notification, the presentation of the set of media items is caused.
17. The non-transitory machine-readable storage medium according to claim 16, wherein, The notification includes a display of an indication of the number of media items in the set of media items.
18. The non-transitory machine-readable storage medium according to claim 15, wherein, The presentation of the group of media items includes displaying the group of media items in a horizontal array.
19. The non-transitory machine-readable storage medium according to claim 15, wherein, The presentation of the group of media items includes displaying the group of media items in a vertical array.
20. The non-transitory machine-readable storage medium according to claim 15, wherein, Identifying the object depicted by the image based on the multiple image features further includes: Detect the multiple image features; and In response to the detection of the plurality of image features, an indicator is presented at a location within the image.