Determining visual theme in collection of media item

A machine learning-based method for clustering media items by visual similarity addresses inefficiencies in organizing large image collections, enhancing user experience through automated theme identification and reduced power consumption.

JP2025124812AActive Publication Date: 2025-08-26GOOGLE LLC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025092492
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-25
Filing Date
2025-06-03
Publication Date
2025-08-26
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

Existing methods for organizing large collections of images and videos in user accounts lack efficiency and accuracy in identifying visual themes without manual intervention, leading to disorganized and inefficient user experiences.

Method used

A computer-implemented method using a trained machine learning model to generate vector representations of media items, cluster them based on visual similarity, and display a user interface with selected clusters, incorporating temporal and location diversity, user feedback, and machine learning model updates.

Benefits of technology

The method provides efficient and accurate clustering of visually similar media items, reducing power consumption and improving user experience by leveraging machine learning for automated theme identification and dynamic model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124812000001_ABST
    Figure 2025124812000001_ABST
Patent Text Reader

Abstract

To provide a method and system for determining, based on pixels of images or videos from a collection of media items, clusters of media items and a subset of the clusters.SOLUTION: A method includes: Step 608 of determining, based on pixels of images or videos from a collection of media items, clusters of media items such that the media items in each cluster have a visual similarity; Step 610 of selecting a subset of the clusters of media from corresponding clusters of media items based on the media items in each cluster having a visual similarity within a range of threshold similarity values; and Step 612 of displaying a user interface that includes the subset of the clusters of media.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 17 / 509767, filed October 25, 2021, entitled "Determining a Visual Theme from Pixels in a Media Item Collection," which claims priority to both U.S. Provisional Patent Application No. 63 / 187390, filed March 11, 2021, entitled "Determining a Visual Theme from Pixels in a Media Item Collection," and U.S. Provisional Patent Application No. 63 / 189658, filed May 17, 2021, entitled "Determining a Visual Theme from Pixels in a Media Item Collection," each of which is incorporated herein in its entirety. [Background technology]

[0002] background Users of devices such as smartphones or other digital cameras take numerous photographs and videos and store them in image libraries. Users utilize such libraries to view their photographs and videos to reminisce about various events, such as birthdays, weddings, vacations, and trips. Users may have large image libraries containing thousands of images taken over an extended period of time.

[0003] The discussion of the background art set forth herein is intended to provide a general context for the present disclosure. To the extent that it is set forth in this background art section, the work of the presently named inventors, as well as statements that do not qualify as prior art at the time of filing, are not expressly or implicitly admitted as prior art to the present disclosure. Summary of the Invention [Means for solving the problem]

[0004] overview The computer-implemented method includes using a trained machine learning model to generate vector representations of media items from a media item collection associated with a user account; and determining media item clusters based on the vector representations of the media items such that the media items in each cluster have visual similarity, wherein a vector distance between the vector representations of pairs of media items indicates the visual similarity of the media items, and the clusters are selected such that the vector distance between each pair of media items in the cluster is outside a visual similarity threshold range, and the method further includes displaying a user interface including a subset of the media item clusters.

[0005] In some embodiments, each media item has an associated timestamp, and media items acquired within a predetermined time period are associated with an episode, and selecting the subset of media item clusters is performed based on the corresponding associated timestamps such that corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of corresponding media items from a particular episode. In some embodiments, the method further includes excluding media items associated with categories on a prohibited category list from the media item collection before selecting the subset of media item clusters. In some embodiments, the method further includes excluding media items corresponding to categories on a prohibited category list before determining the media item clusters. In some embodiments, each media item is associated with a location, and in response to the subset of media item clusters including more than a predetermined number of media items, Selecting the subset of media item clusters is based on location, such that the subset of media item clusters meets a location diversity criterion. In some embodiments, the media item clusters are further determined based on corresponding media items associated with labels having semantic similarity. In some embodiments, the method further includes scoring each media item in the subset of media item clusters based on analyzing a likelihood that a user associated with the user account will reference the media items and perform a positive behavior, and selecting media items from the subset of media item clusters based on corresponding scores that meet a threshold score. In some embodiments, the method further includes receiving feedback from a user regarding one or more media items in the subset of media item clusters and modifying corresponding scores of one or more media items in the subset of media item clusters based on the feedback. In some embodiments, the feedback includes an explicit behavior indicated by deleting one or more media items from the subset of media item clusters from a user interface or an implicit behavior indicated by one or more of viewing corresponding media items in the subset of media item clusters or sharing corresponding media items in the subset of media item clusters. In some embodiments, the method includes receiving aggregate feedback of the aggregated subset of media item clusters from users and providing the aggregate feedback to a trained machine learning model, where parameters of the trained machine learning model are updated, and the method further includes modifying the media item clusters based on updating the parameters of the trained machine learning model. In some embodiments, the method further includes selecting a particular media item from each cluster in the subset of media item clusters as a cover photo for each cluster in the subset of media item clusters based on the particular media item including a maximum number of objects corresponding to a visual similarity.In some embodiments, the method further includes adding a title to each cluster in the subset of media item clusters based on the visual similarity type and the common expression. In some embodiments, the user interface is displayed at predetermined intervals. In some embodiments, the method further includes providing a notification to a user associated with the user account that the subset of media item clusters is available, the notification including a title corresponding to each of the clusters in the subset of media item clusters. In some embodiments, the method further includes determining computations to be performed on individual devices to optimize the computations, and implementing the trained machine learning model on the multiple devices based on the computations performed on the individual devices.

[0006] In some embodiments, a method includes receiving media items from a media item collection associated with a user account as input to a trained machine learning model, and generating output image embeddings of media item clusters using the trained machine learning model, wherein the media items in each cluster have visual similarity, and dividing the vector space generates the media item clusters such that media items with visual similarity are closer to each other in the vector space than dissimilar media items, and the method further includes selecting a subset of the media item clusters based on corresponding media items in each cluster having visual similarity within a visual similarity threshold, and displaying a user interface including the subset of the media item clusters.

[0007] In some embodiments, functional images are removed from the media item collection before the media item collection is provided to the trained machine learning model. In some embodiments, the trained machine learning model is configured to generate a set of media items based on user feedback including reactions to the set of media items or titles of the set of media items. It is trained with user feedback including modifications.

[0008]

[0010] Embodiments may further include a system comprising one or more processors and a memory storing instructions executed by the one or more processors. The instructions include determining media item clusters based on pixels of images or videos from a media item collection, such that media items in each cluster have a visual similarity, the media item collection being associated with a user account; selecting a subset of the media item clusters based on corresponding media items in each cluster having a visual similarity within a visual similarity threshold; and displaying a user interface including the subset of the media item clusters. In some embodiments, each media item has an associated timestamp, and media items acquired within a predetermined time period are associated with an episode, and selecting the subset of the media item clusters is performed based on the corresponding associated timestamps such that corresponding media items in the subset of the media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of corresponding media items from a particular episode.

[0009]

[0010] Embodiments may further include a non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform the following operations: determining media item clusters based on pixels of images or videos from a media item collection, such that media items in each cluster have a visual similarity, the media item collection being associated with a user account; selecting a subset of the media item clusters based on corresponding media items in each cluster having a visual similarity within a visual similarity threshold; and displaying a user interface including the subset of the media item clusters. In some embodiments, each media item has an associated timestamp, and media items acquired within a predetermined time period are associated with an episode, and selecting the subset of media item clusters is performed based on the corresponding associated timestamps such that corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of corresponding media items from a particular episode. [Effects of the Invention]

[0010] This specification describes a method for identifying clusters of similar images (or other media items) using a machine learning model, advantageously without the need to manually identify the images or manually provide categories for the images (or other media items). In this manner, an improved method for classifying images or other media items can be provided. The method can, for example, provide classifications into events that more reliably reflect underlying trends in the data than predefined classifications or categories. Furthermore, the machine learning model can advantageously reduce power consumption and increase efficiency by using a static training set and updating the event machine learning model in response to the update size being less than a threshold size. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary network environment in accordance with certain embodiments described herein. [Figure 2] FIG. 1 is a block diagram illustrating an exemplary computing device in accordance with some embodiments described herein. [Figure 3A] According to some embodiments, an exemplary set of various media items each matching a particular visual theme is shown in FIG. 3A, which shows a first set of media items matching a first visual theme including an object with a curved shape, a second visual theme including three images that are the same still life, and a third visual theme including a cat in a stuffed shark in different poses. [Figure 3B] FIG. 3B illustrates a fourth visual theme in which the same object (a backpack) is seen in each image taken at different times and in different locations, according to some embodiments. [Figure 4] 1 is an example of a visual theme of nature images of different mountain ranges with both temporal and spatial diversity, according to some embodiments. [Figure 5] 1 is an example of a user interface including clusters with visual themes according to some embodiments described herein. [Figure 6] 1 is a flowchart illustrating an example method for displaying a subset of a media item cluster according to some embodiments described herein. [Figure 7] 1 is a flowchart illustrating an example method for generating embeddings of media item clusters and selecting a subset of media item clusters using a machine learning model, according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0012] Detailed Description Network Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes media server 101, user device 115a, user device 115n, and network 105. Users 125a, 125n may be associated with user devices 115a, 115n, respectively. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1 or may not include media server 101. In FIG. 1 and other figures, a letter following a reference number, e.g., "115a," indicates a reference to the element with that specific reference number. A reference number in text without a letter following it, e.g., "115," indicates a general reference to an embodiment of the element with that reference number.

[0013] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively connected to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits and receives data to and from one or more user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0014] The media application 103a may include code and routines operable to, with user permission, determine media item clusters based on pixels of images or videos from a media item collection associated with a user account, such that media items in each cluster have visual similarity, where the media item collection is associated with a user account. For example, one cluster may include objects with similar shapes and colors, another cluster may include parks with similar environmental attributes, and another cluster may include images of pets in different situations. The media application 103a selects a subset of media item clusters based on corresponding media items in each cluster that have visual similarity within a visual similarity threshold. The media application 103a displays a user interface including the subset of media item clusters.

[0015] In some embodiments, the media application 103a may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.

[0016] The database 199 can store media collections associated with user accounts, training sets for machine learning models, and user actions related to media (views, shares, annotations, etc.). The database 199 can store media items that are indexed and associated with the identity of the user 125 of the user device 115. The database 199 can also store social network data associated with the user 125, user preferences for the user 125, etc.

[0017] The user device 115 may be a computing device that includes a memory and a hardware processor. For example, the user device 115 may include a desktop computer, a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.

[0018] In the illustrated implementation, user device 115a is connected to network 105 via signal line 108, and user device 115n is connected to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a, 115n are utilized by users 125a, 125n, respectively. The user devices 115a, 115n in FIG. 1 are used for illustrative purposes. While FIG. 1 shows two user devices 115a and 115n, the present disclosure applies to system architectures including one or more user devices 115.

[0019] In some embodiments, a user account includes a media item collection. For example, a user captures images and videos from their camera (e.g., a smartphone or other camera), uploads images from a digital single-lens reflex (DSLR) camera, and adds media captured and shared by other users to their media item collection. The media application 103 determines media item clusters based on pixels of images or videos from the media item collection, such that media items within each cluster have visual similarity. For example, FIG. 3A shows a first visual theme 300 of visually similar images (brown objects having curved shapes). Specifically, the first object is a drink with ice in a glass, the second object is a latte with a heart in a coffee cup, and the third object is a bowl made of wood in a different shade of brown. Other examples can include mountain ranges, natural arches, ocean waves with people in them, horizontally extending parallel lines (e.g., train tracks, roads, etc.), changes over time (plant growth, sun movement, paint in progress), etc.

[0020] A media item cluster can contain images from the same episode, e.g., multiple images of the same piece taken by a user at different angles. For example, Figure 3A shows three images: Figure 3 shows a second example 325 containing three images of the same still life photographed in different ways so that the leaves become progressively more distinct.

[0021] The media application 103 selects a subset of media item clusters based on corresponding media items in each cluster that have visual similarities within a visual similarity threshold. The visual similarity threshold may range from very similar media items to media items that are more similar than items that are only distantly related. For example, the theme of the first example 300 in FIG. 3A is brown circular objects. This may be in the middle of the similarity threshold range. Conversely, the third example 350 in FIG. 3A is a media item cluster with a theme of cats in stuffed sharks photographed at different times of the day. This is a more visually similar theme. The fourth example 375 in FIG. 3B includes a theme of orange backpacks used by people on different trips. Yet another example of very similar media that is closer to the threshold similarity value is when the media items are pink flowers of slightly different shapes.

[0022] If the media items are not sufficiently visually similar, it may be difficult to identify a theme among the media items, and as a result, the media items may appear more like a random collection of media items than what the user wants to see. In some embodiments, the media application 103 limits the number of media items to keep the visual theme more consistent, so that the media item collection does not appear, for example, as a group of all cat pictures available from the user's library.

[0023] The media application 103 can display a user interface that includes a subset of the media item clusters. In some embodiments, the media application 103 displays the user interface that includes the subset of the media item clusters at predetermined intervals. For example, the media application 103 may display the user interface that includes the subset of the clusters daily, weekly, monthly, etc. The media application 103 may change the frequency with which it displays the subset of the clusters based on feedback. For example, if the user views the subset of the clusters every time they become available, the media application 103 may maintain the frequency of display, but if the user views the subset of the clusters less frequently, the media application 103 may reduce the frequency of display.

[0024] The media application 103 can provide a notification to a user associated with a user account that a subset of media item clusters is available, along with a title corresponding to the subset of clusters. For example, the media application 103 can provide the user with daily notifications, weekly notifications, monthly notifications, etc. In some embodiments, the user interface includes options for limiting the frequency of notifications and / or the display of the subset of media item clusters.

[0025] Exemplary Computing Device 200 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is user device 115 used to run media application 103. In another example, computing device 200 is media server 101. In yet another example, media application 103 is located partially on user device 115 and partially on media server 101.

[0026] One or more methods described herein can be implemented as a standalone program running on any type of computing device, a program running on a web browser, or a mobile application (app) running on a mobile computing device (e.g., a mobile phone, a smartphone, a tablet computer, a wearable device (e.g., a watch, an armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, a head-mounted display), or a laptop computer). In a primary example, all computations are performed in the mobile application on the mobile computing device. However, a client / server architecture can be used. For example, the mobile computing device sends user input data to a server device and receives and outputs (e.g., displays) final output data from the server. In another example, computations may be shared between the mobile computing device and one or more server devices.

[0027] In some embodiments, computing device 200 includes processor 235, memory 237, I / O interface 239, display 241, camera 243, and storage device 245. Processor 235 may be connected to bus 218 via signal line 222. Memory 237 may be connected to bus 218 via signal line 224. I / O interface 239 may be connected to bus 218 via signal line 226. Display 241 may be connected to bus 218 via signal line 228. Camera 243 may be connected to bus 218 via signal line 230. Storage device 245 may be connected to bus 218 via signal line 232.

[0028] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling basic operations of the computing device 200. A "processor" includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., having a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a dedicated circuit for achieving a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a system with a processor optimized for performing matrix calculations (e.g., matrix multiplication), or other systems. In some implementations, the processor 235 may include one or more coprocessors for performing neural network processing. In some implementations, the processor 235 may be a processor that generates a probabilistic output by processing data. For example, the output generated by the processor 235 may be inaccurate or accurate within a range of expected output values. Processing need not be limited to a particular geographic location, nor need it be limited in time. For example, a processor may perform functions in real time, offline, or batch mode. Portions of processing may be performed by different (or the same) processing systems at different times and in different locations. A computer may be any processor in communication with a memory.

[0029] Memory 237 is typically provided within computing device 200 for use by processor 235 and may be any suitable processor-readable storage medium, e.g., random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, for storing instructions executed by the processor or set of processors. Memory 237 may be located separately from processor 235 and / or may be integrated therewith. Memory 237 stores software executed on computing device 200 by processor 235. The application can store applications (including media applications 103).

[0030] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, a camera application, an image library application, an image management application, an image gallery application, a media display application, a communication application, a web hosting engine or application, a mapping application, a media sharing application, etc. One or more methods disclosed herein may be implemented in a number of environments and platforms, for example, as a standalone computer program capable of running on any type of computing device, as a web application having web pages, or as a mobile application (“app”) running on a mobile computing device.

[0031] Application data 266 may also be data generated by other applications 264 or hardware of computing device 200. For example, application data 266 may include images captured by camera 243, user behavior identified by other applications 264 (e.g., social networking applications), etc.

[0032] The I / O interface 239 can provide functionality that allows the computing device 200 to interface with other systems and devices. The interfacing devices may be included as part of the computing device 200 or may be separate but in communication with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or database 199), and input / output devices can communicate through the I / O interface 239. In some embodiments, the I / O interface 239 can connect to interfacing devices, such as input devices (keyboards, pointing devices, touchscreens, microphones, cameras, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.). For example, when a user provides touch input, the I / O interface 239 transmits data to the media application 103.

[0033] Some exemplary interface connection devices that can be connected to I / O interface 239 may include display 241, which can be used to display the content described herein, e.g., images, videos, and / or user interfaces of output applications, and to receive touch (or gesture) input from a user. For example, display 241 can be used to display a user interface that includes a subset of media item clusters. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), light emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in an eyeglass form factor or headset device, or a monitor screen of a computing device.

[0034] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103. do.

[0035] Storage device 245 stores data related to media application 103. For example, storage device 245 may store media item collections associated with a user account, subsets of media clusters, training sets for machine learning models, etc. In embodiments in which media application 103 is part of media server 101, storage device 245 is the same as database 199 of FIG.

[0036] Exemplary Media Application 103 2 illustrates an exemplary media application 103. The media application 103 includes a filtering module 202, a clustering module 204, a machine learning module 205, a selection module 206, and a user interface module 208. In some embodiments, the media application 103 uses either the clustering module 204 or the machine learning module 205.

[0037] The filtering module 202 excludes media items from the media item collection that correspond to categories on the prohibited category list. In some embodiments, the filtering module 202 includes a set of instructions executable by the processor 235 to exclude media items that correspond to categories on the prohibited category list. In some embodiments, the filtering module 202 may be stored in memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0038] In some embodiments, the filtering module 202 excludes media from the media item collection before the clustering module 204 performs clustering. In an alternative embodiment, the filtering module 202 excludes media from the media item collection after the clustering module 204 performs clustering. For example, the filtering module 202 excludes media items associated with visual similarity to categories from a prohibited category list. The prohibited category list can include media items that are photographed not for their photographic value but as functional images, such as receipt images, document images, parking meter images, screenshot images, etc.

[0039] In some embodiments in which the media application 103 includes a machine learning module 205, the filtering module 202 filters out functional images from the media item collection before the media item collection is provided to the machine learning model. For example, the filtering module 202 filters out receipts, instructions, documentation, and screenshots before the media item collection is provided to the machine learning model.

[0040] The clustering module 204 determines media item clusters based on pixels of images or videos from the media item collection such that the media items in each cluster have visual similarity. In some embodiments, the clustering module 204 includes a set of instructions executable by the processor 235 to generate the media item clusters. In some embodiments, the clustering module 204 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0041] In some embodiments, the clustering module 204 may The filtering module 202 accesses a media item collection associated with the user, e.g., a library associated with the user. If the filtering module 202 filters out a media item, the clustering module 204 accesses a media item collection that does not include a media item corresponding to the prohibited category list. The clustering module 204 can determine media item clusters based on pixels of images or videos from the media item collection such that media items within each cluster have visual similarity. In some embodiments, the clustering uses an N-dimensional Gaussian diversity function to determine visual similarity.

[0042] In some embodiments, the machine learning module 205 includes a machine learning model trained to generate output image embeddings for media clusters such that media items within each cluster have visual similarity. In some embodiments, the machine learning module 205 includes a set of instructions executable by the processor 235 to generate the image embeddings. In some embodiments, the machine learning module 205 may be stored in memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0043] In some embodiments, the machine learning module 205 can determine the visual similarity of clusters using vectors (embeddings) in a multidimensional feature space. Images with similar features may have similar feature vectors. For example, the vector distance between feature vectors of images with similar features may be smaller than the vector distance between dissimilar images. The feature space may be a function of various image factors, such as the depicted subject matter (objects detected from the image), image composition, color information, image orientation, image metadata, and specific objects recognized from the image (e.g., known faces, with user permission).

[0044] In some embodiments, the training may be performed using supervised learning. In some embodiments, the machine learning module 205 includes a set of instructions executable by the processor 235. In some embodiments, the machine learning module 205 may be stored in memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0045] In some embodiments, the machine learning module 205 can generate a trained model, specifically a machine learning model, using training data (obtained with permission for training). For example, the training data can include ground truth data in the form of media clusters associated with visual similarity descriptions of the clusters. In some embodiments, the visual similarity descriptions can include user feedback on whether the clusters are related and contain clear themes. In some embodiments, the visual similarity descriptions can be added automatically through image analysis. The training data can be obtained from any source, such as a data repository specified for training or data that has been granted permission to be used as training data for machine learning.

[0046] In some embodiments, the training data may include synthetic data generated for training purposes, e.g., data that is not based on activity in the situation being trained, e.g., data generated from simulations or computer-generated images / videos. In some embodiments, the machine learning module 205 uses weights obtained and unedited / transferred from another application. For example, in these embodiments, the trained model may be generated, e.g., on a different device and provided as part of the media application 103. In various embodiments, the trained model may include a model structure or morphology (e.g., defining the number and type of neural network nodes, the connections between the nodes, and the organization of the nodes into multiple layers) and associated weights. The machine learning module 205 may read the trained model data file and implement a neural network including node connections, layers, and weights based on the model structure or form specified in the trained model.

[0047] The machine learning module 205 generates a trained model, referred to herein as an event machine learning model. In some embodiments, the machine learning module 205 is configured to apply the event machine learning model to data, such as application data 266 (e.g., input media), to identify one or more features in input media items and generate feature vectors (embeddings) that represent the media items. In some embodiments, the machine learning module 205 can include software code executed by the processor 235. In some embodiments, the machine learning module 205 can specify circuitry (e.g., a programmable processor, a field programmable gate array (FPGA)) that enables the processor 235 to apply the machine learning model. In some embodiments, the machine learning module 205 can include software instructions, hardware instructions, or a combination thereof. In some embodiments, the machine learning module 205 can provide an application programming interface (API). The operating system 262 and / or other applications 264 can utilize this API to invoke the machine learning module 205 and output image embeddings of media clusters, for example, by applying the machine learning model to the application data 266. In some embodiments, media items that match visual similarity are closer to each other in vector space than dissimilar images. Therefore, media item clusters are generated by partitioning the vector space.

[0048] In some embodiments, the machine learning model is a classifier that receives the collection of media items as input. Examples of classifiers include neural networks, support vector machines, k-nearest neighbors, logistic regression, naive Bayes, decision trees, perceptrons, etc.

[0049] In some embodiments, the machine learning model may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., "hidden layers" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results obtained from the processing of each tile), or a sequence-to-sequence neural network (e.g., a network that receives sequential data as input, such as words in a sentence or frames in a video, and produces a sequence of results as output).

[0050] The model form or structure may specify the connections between various nodes and the organization of the nodes into layers. For example, nodes in an initial layer (e.g., input layer) may receive data as input data or application data 266. For example, when using a machine learning model to analyze an input image, e.g., a first image, associated with a user account, such data may include, for example, one or more pixels per node. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connections specified in the model form or structure. These layers are sometimes referred to as hidden layers. The final layer (e.g., output layer) generates an output of the machine learning model. For example, this output may be an image embedding of media clusters. In some embodiments ,The model form or structure specifies the number and / or type of nodes in each ,layer.

[0051] The features output by the machine learning module 205 may include the subject matter (e.g., a sunset versus a particular person), the colors present in the image (green hills versus a blue lake), color balance, lighting source, angle and intensity, the location of objects in the image (e.g., adhering to the rule of thirds), the relative location of objects (e.g., depth of field), the location of the shot, the focus (foreground versus background), or shadows. While the aforementioned features are human-understandable, the output features may be embeddings or other numerical values ​​that are representative of the image and not human-analyzable (e.g., individual feature values ​​may not correspond to specific features such as the colors present, the location of objects, etc.). However, the trained model is robust to images, outputting similar features for similar images and dissimilar features for significantly different images.

[0052] In some embodiments, the model is a neural network (CNN) that includes network layers, each of which extracts image features at a different level of abstraction. A CNN used to identify features within an image may be used to classify the image. The model architecture may include layer combinations and sequences of multidimensional convolutions, average pooling, max pooling, activation functions, normalization, regularization, and other layers and modules commonly used in applied deep neural networks.

[0053] In different embodiments, the machine learning model may include one or more models. The one or more models may include multiple nodes arranged in layers according to a model structure or morphology. In some embodiments, the node may be, for example, a memoryless computational node configured to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and generating a node output by adjusting the weighted sum with a bias value or an intercept value. For example, the machine learning module 205 may adjust each weight based on feedback in response to automatically updating one or more parameters of the machine learning model.

[0054] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, the computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, using individual processing units of a graphical processing unit (GPU), or using dedicated neural circuitry. In some embodiments, the nodes may include memory. The nodes may, for example, store one or more previous inputs and use the one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a long-short-term memory (LSTM) node. The LSTM node can use the memory to maintain state that allows the node to operate like a finite state machine (FSM). Models including such nodes can be used to process sequential data, for example, multiple words in a sentence or paragraph, a series of images, or videos. This may be useful when processing frames in video, speech, or other audio. For example, a heuristics-based model used in a gating model may remember one or more features previously generated for previous images.

[0055] In some embodiments, the machine learning model may include embeddings or weights for individual nodes. For example, the machine learning model may be initialized as a plurality of nodes organized into layers as specified by the model topology or structure. At initialization, each pair of nodes connected according to the model topology, e.g., each node in successive layers of a neural network, may be assigned an embedding or weight. Respective weights can be applied to the connections between pairs. For example, each weight can be randomly assigned or initialized to a default value. A machine learning model can then be trained using, for example, a training set of media clusters to generate results. In some embodiments, a subset of the entire architecture can be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.

[0056] For example, training can include applying supervised learning techniques. In supervised learning, training data can include multiple inputs (e.g., media items from a media item collection associated with a user account) and expected outputs (e.g., image embeddings of media clusters) corresponding to each input. For example, weight values ​​are automatically adjusted based on a comparison of the machine learning model's output and the expected output to increase the probability that the machine learning model will generate the expected output when provided with similar inputs.

[0057] In some embodiments, training can include applying unsupervised learning techniques, in which only input data (e.g., media items from a media item collection associated with a user account) may be provided, and a machine learning model may be trained to distinguish between the data, for example, to cluster image features into groups.

[0058] In various embodiments, the trained model includes a set of weights corresponding to the model structure. In embodiments that omit a training set, the machine learning module 205 may generate the machine learning model based on prior training, such as by the developer of the machine learning module 205 or a third party. In some embodiments, the machine learning model may include a set of fixed weights downloaded from a server that provides the weights.

[0059] In some embodiments, the machine learning module 205 may be implemented in an offline manner. Implementing the machine learning module 205 may include using a static training set that does not include updates when data in the static training set changes. This advantageously results in increased efficiency of processing performed by the computing device 200 and reduced power consumption of the processing device 200. In these embodiments, the machine learning model may be generated in a first stage and provided as part of the machine learning module 205. In some embodiments, small updates to the machine learning model may be implemented in an online manner, where updates to the training data are included as part of training the machine learning model. Small updates are updates having a size smaller than a threshold size. The size of the update is related to the number of variables in the machine learning model affected by the update. In such embodiments, an application invoking the machine learning module 205 (e.g., the operating system 262, one or more other applications 264, etc.) can utilize image embeddings of media item clusters to identify visually similar clusters. The machine learning module 205 may also generate a system log periodically, for example, hourly, monthly, or quarterly. The system log may be used to update the machine learning model, for example, to update the embeddings of the machine learning model.

[0060] In some embodiments, the machine learning module 205 may be implemented in a manner that is compatible with the particular configuration of the computing device 200 on which the machine learning module 205 executes. For example, the machine learning module 205 may determine a computation graph that utilizes available computational resources, such as the processor 235. If the machine learning module 205 is implemented as a distributed application on multiple devices, for example, if the media server 101 includes multiple media servers 101, the machine learning module 205 may The computations performed on each device may be determined to optimize the computations. In another example, if the machine learning module 205 determines that the processor 235 includes a GPU with a certain number of GPU cores (e.g., 1000), the machine learning module 205 may be implemented (e.g., as 1000 separate processes or threads).

[0061] In some embodiments, the machine learning module 205 can implement a set of trained models. For example, an event machine learning model can include multiple trained models, each applicable to the same input data. In these embodiments, the machine learning module 205 can select a particular trained model based on, for example, available computational resources, success rates using previous inferences, etc.

[0062] In some embodiments, the machine learning module 205 can execute multiple trained models. In these embodiments, the machine learning module 205 can combine outputs, for example, using a majority vote to score the outputs obtained by applying each trained model, or by selecting one or more specific outputs. In some embodiments, such a selector is part of the model itself and serves as a connection layer between the trained models. Furthermore, in these embodiments, the machine learning module 205 can apply a time threshold (e.g., 0.5 ms) for applying each trained model and only utilize individual outputs that are available within the time threshold. Outputs that are not received within the time threshold may not be utilized, e.g., discarded. For example, such an approach may be appropriate when there is a specified time limit between invoking the machine learning module 205, e.g., by the operating system 262 or one or more other applications 264. In this way, the responsiveness of the media application 103 can be improved because the maximum time it takes for the machine learning module 205 to perform a task, e.g., to identify one or more features of an input media item and generate a feature vector (embedding) representing the media item, can be limited, thereby enabling the machine learning module 205 to provide the best classification in real time.

[0063] In some embodiments, the machine learning module 205 receives feedback. For example, the machine learning module 205 can receive feedback from a single user or a set of users via the user interface module 208. If a single user provides feedback, the machine learning module 205 provides the feedback to the machine learning model, and the machine learning model 205 uses the feedback to update parameters of the machine learning model to modify the output image embeddings of the media item clusters. If a set of users provides feedback, the machine learning module 205 provides aggregate feedback to the machine learning model, and the machine learning model 205 uses the aggregate feedback to update parameters of the machine learning model to modify the output image embeddings of the media item clusters. For example, the aggregate feedback may include a subset of media clusters and user reactions to the subset of media clusters. User responses include viewing only one image and refusing to view the remaining media, viewing all of the corresponding media items in the subset, sharing the corresponding media items, providing an indication of approval or disapproval of the corresponding media items (e.g., thumbs up / down, like, +1, etc.), removing / adding individual media items from the subset of media item clusters, changing the title, etc. The machine learning module 205 can modify the media clusters based on updating the parameters of the machine learning model.

[0064] In some embodiments, the machine learning model is trained using user feedback, which includes responses to subsets of clusters and responses within the subsets. The machine learning module 205 provides feedback to the machine learning model to modify parameters to filter out media item clusters that have certain types of visual similarity (e.g., images of ocean waves that are visually similar but are not the type of media the user wants to see, and images of surfers on waves taken at different times and / or in different locations).

[0065] The selection module 206 selects a subset of the media item clusters based on the visual similarities determined by the clustering module 204. In some embodiments, the selection module 206 includes a set of instructions executable by the processor 235 to select the subset of the media item clusters. In some embodiments, the selection module 206 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0066] In some embodiments, the selection module 206 selects a subset of media item clusters whose media items have a visual similarity within a visual similarity threshold range. For example, the range may be between 0.05 and 0.3, out of 0 to 4. Other ranges and scales are possible. The subset of media item clusters within the visual similarity threshold range may be considered to have a visual theme that is recognized as related and cohesive.

[0067] In some embodiments, if a media item cluster exceeds a predetermined number (e.g., more than 15 media items), the selection module 206 can impose additional constraints when selecting a subset of media item clusters. For example, the selection module 206 can impose temporal diversity by identifying a timestamp associated with each media item, identifying based on the timestamps that multiple media items (e.g., media items associated with the same time period and the same location) are associated with the same episode, and selecting a subset of media item clusters based on their associated timestamps so that they meet a temporal diversity criterion that excludes more than a predetermined number of media items from a particular episode (i.e., selecting a subset of media item clusters based on their associated timestamps such that no more than a certain number (e.g., 3) of media items are associated with the same episode). This avoids media item clusters that are too similar and may overlap, as a user may take multiple images of an object at the same time and in the same location. This also avoids a situation where a user takes the same image and edits it to post, for example, to another photo-sharing application. The selection module 206 can use temporal diversity to select a subset of clusters that show the progression of an object over a period of time. For example, a cluster may contain different images of a child taken at different time periods to show the child growing larger, or different images of a plant from a seedling to a flowering bush.

[0068] In some embodiments, the selection module 206 imposes a spatial diversity on the subset of media clusters. For example, the selection module 206 may identify a location associated with each media item, and if the number of corresponding media items available for a cluster exceeds a predetermined number (e.g., more than 10 media items), the selection module 206 selects a subset of media item clusters based on location, such that the subset of media item clusters meets a spatial diversity criterion. Figure 4 includes an example visual theme 400 of nature images of different mountain ranges that have both temporal diversity, because the images were taken in different months and years, and spatial diversity, because the images were taken in different locations. Although the images have two types of diversity, hidden similarities emerge through the visual theme.

[0069] In some embodiments, the selection module 206 imposes a semantic theme on a subset of the media clusters. The selection module 206 can identify labels associated with images and group the subset of media item clusters based on corresponding media items that have the same or similar labels. For example, the selection module 206 can select a subset of media item clusters ranging from puppies to adult dogs using labels that identify depictions of dogs in images. In some embodiments, the media application 103 combines the semantic theme of the Golden Gate Bridge with the visual theme of other bridges that are visually similar to the golden color of the Golden Gate Bridge.

[0070] In some embodiments, the selection module 206 scores each media item in the subset of media item clusters based on analyzing the likelihood that a user associated with the user account will view the media item and perform a positive action. Positive actions may include viewing the subset, sharing the subset, ordering prints from the subset, etc. The selection module 206 may score a media item as being associated with the likelihood that a user associated with the user account will perform a positive action if the subject matter is more interesting, for example, if the subject matter includes babies, people the user knows, places the user has visited, etc. Conversely, the selection module 206 may determine that a user is less likely to perform a positive action related to a particular subject, for example, a static object such as a bunk bed. In some embodiments, the selection module 206 scores the subset of media item clusters based on personal information related to the user or aggregate information regarding the user's general reaction to the media. In some embodiments, the selection module 206 scores media items based on qualities, such as being too blurry, because the quality of the media item reduces the likelihood that a user associated with the user account will perform a positive action associated with the media item.

[0071] The selection module 206 can select media items in a subset of a media item cluster if the score corresponding to each media item meets a threshold score. In some embodiments, the threshold score is a static value that is the same for all users. In some embodiments, the threshold score is user-specific. In some embodiments, the threshold score is specified by the user.

[0072] Once the selection module 206 has determined the subset of clusters, it can direct the user interface module 208 to display a user interface that includes the subset of clusters. In some embodiments, a user can provide feedback related to the subset of clusters. For example, a user can view the subset, provide an indication of approval of the subset, share the subset, and order prints of photos from the subset.

[0073] In some embodiments, the selection module 206 receives feedback and modifies the scores corresponding to subsets of media item clusters based on the feedback. For example, the feedback may include explicit behavior indicated by deleting a subset of clusters from a user interface, or implicit behavior indicated by one or more of viewing a subset of clusters, watching a subset of clusters, or sharing a subset of clusters. In some embodiments, the selection module 206 may identify a pattern in the feedback. For example, if positive feedback occurs when objects in a cluster are of a particular type (e.g., babies, families, trees, etc.), the selection module 206 may modify the scores so that the subset of clusters includes objects of similar types. In another example, a pattern may indicate that a user prefers themes with lower visual similarity over higher visual similarity, causing the selection module 206 to modify the scores corresponding to the subsets of media item clusters. The module 206 can modify the scores to select themes with lower visual similarity more frequently.

[0074] In some embodiments, the selection module 206 can receive feedback from a set of users who use the media application 103 and aggregate the feedback. For example, the selection module 206 can create aggregate feedback of the users for a subset of the media clusters and can modify the scoring based on the aggregate feedback.

[0075] The user interface module 208 generates the user interface. In some embodiments, the user interface module 208 includes a set of instructions executable by the processor 235 to generate the user interface. In some embodiments, the user interface module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0076] The user interface module 208 displays a user interface that includes a subset of the media clusters. Figure 5 shows an example user interface 500 that includes clusters 505 with a visual theme, according to some embodiments described herein. In this example, clusters 505 are displayed at the top of the user interface, along with a group of recent highlights and a group of images from a year ago. User interface 500 also includes images taken yesterday (March 9) in San Francisco.

[0077] In some embodiments, the user interface module 208 generates a user interface for suggesting a subset of the clusters and for viewing, editing, and sharing the media. As shown in Figure 5, the user interface may, for example, include the clusters located at the top of the user interface, and when the user selects an image, the user interface includes options for editing or sharing the image.

[0078] In some embodiments, in response to a user selecting a cluster in the user interface, the user interface module 208 displays corresponding media items from the cluster at predetermined intervals. For example, the user interface module 208 may display each media item for 2 seconds, 3 seconds, etc.

[0079] In some embodiments, the user interface module 208 provides cover photos for a subset of the clusters. The cover photo may be the most recent photo, the photo with the highest score, etc. In some embodiments, the user interface module 208 selects a particular media item from each cluster in the subset of media item clusters as the cover photo for each cluster in the subset of media item clusters based on the particular media item containing the greatest number of objects corresponding to the visual similarity. For example, the clusters may have a visual theme of a group of people skiing, and the user interface module 208 may select a cover photo for the cluster that shows an image depicting the highest number of people from the cluster of people skiing. In another example, if a cluster has a visual theme of people engaged in outdoor water activities, the user interface module 208 may determine that an image of people surfing is the most representative media item for the cover, compared to other images in which people are near water rather than in it (e.g., building sandcastles), or in which people are engaged in more vigorous outdoor activities (e.g., sunbathing along the water). The user interface module 208 may also select a cover photo based on having the highest visual quality (e.g., clear, high resolution, not blurry, well exposed, etc.) within the cluster. It is also possible to do so.

[0080] In some embodiments, the user interface module 208 adds a title to each cluster in the subset of media item clusters based on the visual theme type and / or template representation. For example, the title may describe the action occurring in the image (e.g., "surf's up" for an ocean cluster, "into the blue" for a sky cluster, "on the road" for a road cluster, "stairway to heaven" for a church image), food metaphors (e.g., "the ocean" for an ocean cluster, "into the blue" for a sky cluster, "on the road" for a road cluster, "stairway to heaven" for a church image), or a combination of the image and the title of the media item cluster. (e.g., "mixed nuts," "smorgasbord," "mixed bag," "goody bag," "wine flight," "cheese pairing," "sampler," "treasure trove," "overlooked treasures," "have a drink"), photo trails (e.g., "photo detective," "photo mystery," "mystery photos," "photo sphinx"), creative combinations, correlations (e.g., "photo detective," "photo mystery," "mystery photos," "photo sphinx"), Examples of common words include: connection, feather photo, photo club, photo weaving, patterned, coincidence, cause and effect, slot machine, one of these is the same as the other, similarity), titles referring to patterns (e.g., beta pattern, pattern hunter, small pattern, pattern portal, connect the dots, image pattern, photo pattern), synonyms for patterns such as theme (e.g., photo story, photo story, photo tale, the story of two photos, photo theme, lucky theme), set (e.g., photo set, surprise set), or match (e.g., memory match), onomatopoeia (e.g., zigzag, boom, boom, smack, smack photo, photo smack), verbs (e.g., "look what we found in the couch cushions," "look what appeared," "help us sleuth," "will it blend," The template expression may refer to a cluster title that references a connection (e.g., "time flies," "some things never change") that results in the selection module having a higher confidence score that the guess is correct (e.g., magical pattern, lightheartedness). In some embodiments, the template expression may be a funny or common expression that is more colloquial and engaging than simply adding a title such as "Birthdays 1997-2001." In some embodiments, the user interface module 208, for the fourth example 375 of FIG. 3B, may include a title that includes a general title such as "Look what we found" and a subtitle that identifies the theme, such as "your orange backpack brought you far."

[0081] In some embodiments, the user interface module 208 provides a notification to a user associated with a user account that a subset of the clusters is viewable. The user interface module 208 may provide the notification periodically, such as daily, weekly, monthly, etc. In some embodiments, if the user stops viewing the notification when the notification is provided daily (weekly, monthly, etc.), the user interface module 208 may generate the notification less frequently. The user interface module 208 may additionally provide a notification that includes a title corresponding to the subset of the clusters.

[0082] Exemplary Flowchart 6 is a flowchart illustrating an example method 600 for displaying a subset of a media item cluster, according to some embodiments. The method illustrated in flowchart 600 may be performed by computing device 200 of FIG.

[0083] Method 600 may begin at block 602, where a request for access to a media item collection associated with a user account is generated. In some embodiments, the request is generated by the user interface module 208. Block 602 may be followed by block 604.

[0084] At block 604, an authorization interface element is displayed. For example, a user interface element is displayed. The interface module 208 may display a user interface that includes a permission interface element to request that the user provide permission to access the media item collection. Block 606 may be executed after block 604.

[0085] At block 606, it is determined whether permission for access to the media item collection has been granted by the user. In some embodiments, block 606 is performed by the user interface module 208. If the user has not provided permission, the method ends. If the user has provided permission, block 608 can be performed after block 606.

[0086] In block 608, media item clusters are determined based on pixels of images or videos from the media item collection, such that media items in each cluster have visual similarity. The media item collection is associated with a user account. In some embodiments, block 606 is performed by the clustering module 204. Block 610 can be performed after block 608.

[0087] In block 610, a subset of media item clusters is selected based on corresponding media items in each cluster that have a visual similarity within a visual similarity threshold. In some embodiments, block 610 is performed by the selection module 206. Block 612 can be performed after block 610.

[0088] A user interface including the subset of media clusters is displayed at block 612. In some embodiments, block 610 is performed by the user interface module 208.

[0089] 7 is a flowchart illustrating an example method 700 for generating embeddings for media item clusters and selecting a subset of the media item clusters using a machine learning model, according to some embodiments. The method illustrated in flowchart 700 may be performed by computing device 200 of FIG.

[0090] Method 700 may begin at block 702, where a request for access to a media item collection associated with a user account is generated. In some embodiments, the request is generated by the user interface module 208. Block 702 may be followed by block 704.

[0091] In block 704, a permission interface element is displayed. For example, the user interface module 208 may display a user interface that includes a permission interface element to request that the user provide permission to access the media item collection. Block 704 may be followed by block 706.

[0092] At block 706, it is determined whether permission for access to the media item collection has been granted by the user. In some embodiments, block 706 is performed by the user interface module 208. If the user has not provided permission, the method ends. If the user has provided permission, block 708 can be performed after block 706.

[0093] In block 708, the trained machine learning model receives as input media items from a media item collection associated with the user account. In some embodiments, block 708 is performed by the machine learning module 205. Block 710 can be performed after block 708.

[0094] In block 710, the trained machine learning model generates output image embeddings of media item clusters. The media items in each cluster have visual similarity, and media items with visual similarity are closer to each other than media items that are dissimilar in vector space, such that media item clusters are generated by partitioning the vector space. In some embodiments, block 710 is performed by the machine learning module 205. Block 712 can be performed after block 710.

[0095] At block 712, a subset of media item clusters is selected based on corresponding media items in each cluster having a visual similarity within a visual similarity threshold. In some embodiments, block 712 is performed by the machine learning module 205. Block 714 can be performed after block 712.

[0096] A user interface including the subset of media item clusters is displayed at block 714. In some embodiments, block 714 is performed by the user interface module 208.

[0097] In addition to the above, the systems, programs, or features described herein may provide users with controls to choose whether and when to enable collection of user information (e.g., information about the user's media items, such as photos or videos; the user's interactions with a media application displaying the media items; the user's social networks, social behavior or activities; occupation; viewing preferences for image-based creations; user preferences, such as settings for hiding people or pets; user interface preferences; or information about the user's current location) and to transmit content or information from the server. Furthermore, certain data may be processed in one or more ways to remove identifiable personal information before being stored or used. For example, the user's ID may be processed so that the user's personal information cannot be identified. Also, when obtaining location information (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's location cannot be identified. Thus, users can control what user information is collected, how the information is used, and what information is provided to them.

[0098] In the above description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments may be practiced without these specific details. In some instances, structures and devices are shown in block diagrams to avoid obscuring the description. For example, the embodiments may be described above primarily with reference to user interfaces and specific hardware. However, the embodiments may apply to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0099] References herein to "some embodiments" or "some instances" mean that a particular feature, structure, or characteristic described in connection with the embodiments or instances may be included in at least one implementation of the description. Appearances of the phrase "in some embodiments" in various places in the specification do not necessarily all refer to the same embodiments.

[0100] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is herein, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It is convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0101] It should be understood that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise stated, or as will be apparent from the discussion, throughout the description, discussions utilizing terms including "processing," "operating," "calculating," "determining," or "displaying," etc., refer to operations and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical quantities in the computer system memory, registers, or other information storage, transmission, or display devices.

[0102] Embodiments herein also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on any type of non-transitory computer-readable storage medium, including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each connected to a computer system bus.

[0103] This specification may include some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0104] Furthermore, the descriptions may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, machine, or device.

[0105] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, mass storage devices, and cache memory that provides temporary storage of at least some program code to reduce the number of times the code must be retrieved from mass storage devices during execution.

Claims

1. 1. A computer-implemented method comprising: generating vector representations of media items from a media item collection associated with a user account using the trained machine learning model; determining media item clusters based on the vector representations of the media items such that the media items in each cluster have a visual similarity, wherein the vector distance between vector representations of pairs of media items indicates the visual similarity of the media items, and the clusters are selected such that the vector distance between each pair of media items in the cluster is outside a visual similarity threshold range, the method further comprising: selecting a subset of the media item clusters based on corresponding media items in each cluster having the visual similarity within a visual similarity threshold; and displaying a user interface including the subset of the media item clusters.

2. Each media item has an associated timestamp; The media items acquired within a predetermined time period are associated with episodes; 2. The method of claim 1, wherein selecting the subset of media item clusters is based on corresponding associated timestamps such that corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of the corresponding media items from a particular episode.

3. The method of claim 1 , further comprising excluding from the media item collection media items associated with categories that are on a prohibited category list before selecting the subset of media item clusters.

4. The method of claim 1 , further comprising filtering out media items that correspond to categories that are on a prohibited category list before determining the media item clusters.

5. Each media item is associated with a location, 2. The method of claim 1, wherein selecting the subset of media item clusters in response to the subset of media item clusters containing more than a predetermined number of media items is based on location such that the subset of media item clusters meets a location diversity criterion.

6. The method of claim 1 , wherein the media item clusters are further determined based on the corresponding media items associated with labels having semantic similarity.

7. scoring each media item in the subset of media item clusters based on analyzing the likelihood that a user associated with the user account will view the media item and perform a positive action; The method of claim 1 , further comprising: selecting the media items from the subset of media item clusters based on corresponding scores that meet a threshold score.

8. receiving feedback from a user regarding one or more media items in the subset of media item clusters; and modifying corresponding scores of the one or more media items in the subset of media item clusters based on the feedback. Item 7. The method according to item 7.

9. 9. The method of claim 8, wherein the feedback includes an explicit action indicated by removing one or more media items from the subset of the media item cluster from the user interface, or an implicit action indicated by one or more of viewing the corresponding media items in the subset of the media item cluster or sharing the corresponding media items in the subset of the media item cluster.

10. receiving aggregate feedback from a user of the aggregated subset of the media item clusters; providing the aggregate feedback to the trained machine learning model, whereby parameters of the trained machine learning model are updated; the method further comprising: The method of claim 1 , further comprising modifying the media item clusters based on updating the parameters of the trained machine learning model.

11. The method of claim 1 , further comprising selecting a particular media item from each cluster in the subset of media item clusters as a cover photo for each cluster in the subset of media item clusters based on the particular media item containing the largest number of objects corresponding to the visual similarity.

12. The method of claim 1 , further comprising adding a title to each cluster in the subset of media item clusters based on visual similarity type and common expression.

13. The method of claim 1 , wherein the subset of the media item clusters is displayed in the user interface at predetermined intervals.

14. providing a notification to a user associated with the user account that the subset of the media item clusters is available; The method of claim 1 , wherein the notification includes a title corresponding to each of the clusters in the subset of the media item clusters.

15. determining the computations to be performed on each individual device to optimize the computations; 10. The method of claim 1, further comprising: implementing the trained machine learning model on multiple devices based on the calculations performed on the individual devices.

16. 1. A computer-implemented method comprising: receiving media items from a media item collection associated with a user account as input to a trained machine learning model; generating output image embeddings of media item clusters using the trained machine learning model, wherein the media items in each cluster have a visual similarity, and partitioning the vector space such that media items with the visual similarity are closer to each other in vector space than dissimilar media items generates the media item clusters; selecting a subset of the media item clusters based on corresponding media items in each cluster having a visual similarity within a visual similarity threshold; and displaying a user interface including the subset of the media item clusters.

17. The method of claim 16 , wherein functional images are removed from the media item collection before the media item collection is provided to the trained machine learning model.

18. 17. The method of claim 16, wherein the trained machine learning model is trained using user feedback including reactions to a set of media items or user feedback including changes to a title of the set of media items.

19. 1. A system comprising: a processor; a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform the following operations: The operation is determining media item clusters based on pixels of images or videos from a media item collection such that the media items within each cluster have a visual similarity, the media item collection being associated with a user account; selecting a subset of the media item clusters based on corresponding media items in each cluster having the visual similarity within a visual similarity threshold; and displaying a user interface including the subset of the media item clusters.

20. Each media item has an associated timestamp; The media items acquired within a predetermined time period are associated with episodes; 20. The system of claim 19, wherein selecting the subset of media item clusters is based on corresponding associated timestamps such that corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of the corresponding media items from a particular episode.

Citation Information

Patent Citations

  • Digital asset search user interface

    CN112088370A

  • Retrieval method of web page and clustering method of web page

    JP2007080061A

  • Array-based media item discovery

    JP2009530741A

  • Method and apparatus for navigating image data set, and program

    JP2011154687A

  • Image selection suggestions

    JP2021504803A