Determining the visual theme within the media item collection

A machine learning-based method for clustering visually similar media items addresses the inefficiencies in organizing large photo and video collections by using vector representations and user feedback, enhancing the organization and presentation of image libraries.

JP7843895B2Active Publication Date: 2026-04-10GOOGLE LLC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2025-06-03
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for organizing large collections of photos and videos lack efficiency in identifying and presenting visually similar media items without manual intervention, leading to disorganized and unstructured image libraries.

Method used

A computer-implemented method using a trained machine learning model to generate vector representations of media items, cluster them based on visual similarity, and display a user interface with selected clusters, incorporating temporal, spatial, and semantic diversity criteria, along with user feedback for refinement.

Benefits of technology

This approach enables efficient and automated clustering of visually similar media items, improving the organization and presentation of image libraries by reflecting underlying trends and user preferences, reducing manual effort and enhancing user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843895000001
    Figure 0007843895000001
  • Figure 0007843895000002
    Figure 0007843895000002
  • Figure 0007843895000003
    Figure 0007843895000003
Patent Text Reader

Abstract

To provide a method and system for determining, based on pixels of images or videos from a collection of media items, clusters of media items and a subset of the clusters.SOLUTION: A method includes: Step 608 of determining, based on pixels of images or videos from a collection of media items, clusters of media items such that the media items in each cluster have a visual similarity; Step 610 of selecting a subset of the clusters of media from corresponding clusters of media items based on the media items in each cluster having a visual similarity within a range of threshold similarity values; and Step 612 of displaying a user interface that includes the subset of the clusters of media.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 187390, filed on March 11, 2021, titled "Determination of Visual Themes from Pixels in a Media Item Collection" and U.S. Provisional Patent Application No. 63 / 189658, filed on May 17, 2021, titled "Determination of Visual Themes from Pixels in a Media Item Collection", and claims priority to U.S. Patent Application No. 17 / 509767, filed on October 25, 2021, titled "Determination of Visual Themes in a Media Item Collection", and the entire contents of each application are incorporated herein by reference.

Background Art

[0002] Background Users of devices such as smartphones or other digital cameras take numerous photos and videos and store them in an image library. By viewing their photos and videos using such a library, users recall various events such as birthdays, weddings, vacations, and trips. A user can have a large image library containing thousands of images taken over a long period of time.

[0003] The description of the background art described herein is intended to schematically show the context of the present disclosure. Within the scope described in this background art section, the research of the inventors, as currently named, is not expressly or implicitly recognized as prior art to the present disclosure, similar to the descriptions that are not considered prior art at the time of filing.

Summary of the Invention

Means for Solving the Problems

[0004] Summary The computer-implemented method includes using a trained machine learning model to generate vector representations of media items from a collection of media items associated with a user account, determining media item clusters based on the vector representations of media items such that the media items in each cluster have visual similarity, where the vector distance between the vector representations of media item pairs indicates the visual similarity of the media items, and the clusters are selected such that the vector distance between each pair of media items in the cluster is outside the range of a visual similarity threshold, and the method further includes displaying a user interface containing a subset of the media item clusters.

[0005] In some embodiments, each media item has an associated timestamp, and media items acquired within a predetermined period are associated with an episode, and the selection of a subset of media item clusters is based on the corresponding associated timestamps such that the corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of corresponding media items from a particular episode. In some embodiments, the method further includes excluding media items associated with categories on a prohibited category list from the media item collection before selecting a subset of media item clusters. In some embodiments, the method further includes excluding media items corresponding to categories on a prohibited category list before determining the media item clusters. In some embodiments, each media item is associated with a location, and in response to the subset of media item clusters containing more than a predetermined number of media items, The selection of a subset of media item clusters is based on locations such that the subset of media item clusters satisfies a spatial diversity criterion. In some embodiments, media item clusters are further determined based on corresponding media items associated with labels having semantic similarity. In some embodiments, the method further includes scoring each media item in the subset of media item clusters based on an analysis of the likelihood that a user associated with a user account will refer to the media item and take a positive action, and selecting media items from the subset of media item clusters based on corresponding scores that satisfy a threshold score. In some embodiments, the method further includes receiving feedback from the user regarding one or more media items in the subset of media item clusters, and modifying the corresponding scores of one or more media items in the subset of media item clusters based on the feedback. In some embodiments, the feedback includes explicit actions indicated by removing one or more media items from the subset of media item clusters from the user interface, or implicit actions indicated by one or more of viewing the corresponding media items in the subset of media item clusters or sharing the corresponding media items in the subset of media item clusters. In some embodiments, the method includes receiving aggregated feedback from a user for an aggregated subset of media item clusters, providing the aggregated feedback to a trained machine learning model, updating the parameters of the trained machine learning model, and further including modifying the media item clusters based on the updated parameters of the trained machine learning model. In some embodiments, the method further includes selecting a specific media item from each cluster in the subset of media item clusters as the cover photo for each cluster in the subset of media item clusters, based on a specific media item containing the maximum number of objects corresponding to visual similarity.In some embodiments, the method further includes adding a title to each cluster within a subset of media item clusters based on the type and common representation of visual similarity. In some embodiments, the user interface is displayed at predetermined intervals. In some embodiments, the method further includes providing a notification to a user associated with a user account that a subset of media item clusters is available, the notification including a title corresponding to each of the clusters within the subset of media item clusters. In some embodiments, the method further includes determining computations to be performed on individual devices to optimize the computation, and implementing a trained machine learning model on multiple devices based on the computations performed on individual devices.

[0006] In some embodiments, the method includes receiving media items from a media item collection associated with a user account as input to a trained machine learning model, and using the trained machine learning model to generate output image embeddings of media item clusters, wherein the media items in each cluster have visual similarity, and the vector space is divided such that media items with visual similarity are closer to each other than media items that are not similar in the vector space, thereby generating media item clusters, and the method further includes selecting a subset of media item clusters based on the corresponding media items in each cluster having visual similarity within a visual similarity threshold, and displaying a user interface containing the subset of media item clusters.

[0007] In some embodiments, functional images are removed from the media item collection before the media item collection is provided to the trained machine learning model. In some embodiments, the trained machine learning model receives user feedback, including reactions to the media item set, or the titles of the media item set. It is trained using user feedback, including changes.

[0008] Embodiments may further include a system comprising one or more processors and memory for storing instructions executed by the one or more processors. An instruction includes determining media item clusters such that media items in each cluster have visual similarity based on pixels of images or videos from a media item collection, the media item collection being associated with a user account, selecting a subset of media item clusters based on corresponding media items in each cluster having visual similarity within a visual similarity threshold, and displaying a user interface containing the subset of media item clusters. In some embodiments, each media item has an associated timestamp, media items acquired within a predetermined period are associated with an episode, and the selection of a subset of media item clusters is done based on the corresponding associated timestamp such that the corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of corresponding media items from a particular episode.

[0009] The embodiments may further include non-temporary computer-readable media that, when executed by one or more computers, stores instructions causing one or more computers to perform the following actions: The actions include determining media item clusters such that media items in each cluster have visual similarity based on pixels of images or videos from a media item collection, the media item collection being associated with a user account, selecting a subset of media item clusters based on corresponding media items in each cluster having visual similarity within a visual similarity threshold, and displaying a user interface containing the subset of media item clusters. In some embodiments, each media item has an associated timestamp, media items acquired within a predetermined period are associated with an episode, and the selection of a subset of media item clusters is done based on the corresponding associated timestamp such that corresponding media items in the subset of media item clusters satisfy a temporal diversity criterion that excludes more than a predetermined number of corresponding media items from a particular episode. [Effects of the Invention]

[0010] This specification describes a method for identifying clusters of similar images (or other media items) using a machine learning model, advantageously without the need to manually identify images or manually provide categories for images (or other media items). In this way, an improved method for classifying images or other media items can be provided. This method can, for example, provide classification to events that more reliably reflect underlying trends in the data than in the case of predefined classifications or categories. Furthermore, the machine learning model can be made more efficient and power-efficient by updating the event machine learning model in response to the update size being less than a threshold size, using a static training set. [Brief explanation of the drawing]

[0011] [Figure 1] This block diagram shows an exemplary network environment according to some embodiments described herein. [Figure 2] A block diagram illustrating an exemplary computing device according to some embodiments described herein. [Figure 3A] According to several embodiments, each is an exemplary set of various media items that conform to a particular visual theme, and Figure 3A shows the first set of media items conforming to a first visual theme that includes an object having a curved shape, a second visual theme that includes three images that are the same still life, and a third visual theme that includes a cat in a shark stuffing in different poses. [Figure 3B] According to several embodiments, each is an exemplary set of various media items that conform to a particular visual theme, and Figure 3B shows a fourth visual theme in which the same object (a backpack) is seen in each image taken at different times and in different locations, according to several embodiments described herein. [Figure 4] According to several embodiments, this is an example of a visual theme for natural images of different mountain ranges having both temporal and spatial diversity. [Figure 5] An example of a user interface including a cluster having a visual theme, according to some embodiments described herein. [Figure 6] This flowchart illustrates an exemplary method for displaying a subset of media item clusters according to some embodiments described herein. [Figure 7] This flowchart illustrates an exemplary method for generating media item cluster embeddings using a machine learning model and selecting a subset of media item clusters, according to some embodiments described herein. [Modes for carrying out the invention]

[0012] Detailed explanation Network environment 100 Figure 1 shows a block diagram of an exemplary environment 100. In some embodiments, the environment 100 includes a media server 101, a user device 115a, a user device 115n, and a network 105. Users 125a and 125n may be associated with user devices 115a and 115n, respectively. In some embodiments, the environment 100 may include other servers or devices not shown in Figure 1, or may not include the media server 101. In Figure 1 and other drawings, letters following a reference number, such as "115a," indicate a reference to the element having that particular reference number. Reference numbers in the text without letters following them, such as "115," indicate a general reference to embodiments of the element having that reference number.

[0013] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicably connected to the network 105 via a signal line 102. The signal line 102 may be a wired connection such as Ethernet®, coaxial cable, or fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0014] The media application 103a may, with the user's permission, include code and routines that can be operated to determine media item clusters such that the media items in each cluster have visual similarity based on the pixels of images or videos from the media item collection, and the media item collection is associated with the user account. For example, one cluster may contain objects having similar shapes and colors, another cluster may contain parks having similar environmental attributes, and yet another cluster may contain images of pets in different situations. The media application 103a selects a subset of media item clusters based on the corresponding media items in each cluster having visual similarity within a visual similarity threshold. The media application 103a displays a user interface containing the subset of media item clusters.

[0015] In some embodiments, the media application 103a may be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.

[0016] Database 199 can store media collections associated with user accounts, training sets for machine learning models, and user behavior related to media (such as browsing, sharing, and annotating). Database 199 can also store indexed media items associated with the ID of user 125 on user device 115. In addition, database 199 can store social network data related to user 125, user preferences of user 125, and so on.

[0017] The user device 115 may be a computing device including a memory and a hardware processor. For example, the user device 115 may include a desktop computer, a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.

[0018] In the illustrated implementation, the user device 115a is connected to the network 105 via the signal line 108, and the user device 115n is connected to the network 105 via the signal line 110. The media application 103 may be stored on the user device 115a as the media application 103b or on the user device 115n as the media application 103c. The signal lines 108 and 110 may be a wired connection such as Ethernet (registered trademark), coaxial cable, fiber optic cable, or a wireless connection such as Wi-Fi (registered trademark), Bluetooth (registered trademark), or other wireless technologies. The user devices 115a, 115n are each utilized by the users 125a, 125n. The user devices 115a, 115n in FIG. 1 are used as an example. FIG. 1 shows two user devices 115a and 115n, but the present disclosure is applicable to a system architecture including one or more user devices 115.

[0019] In some embodiments, a user account includes a media item collection. For example, a user can acquire images and videos from their camera (e.g., a smartphone or other camera), upload images from a digital single-lens reflex (DSLR) camera, and add media taken and shared by other users to the media item collection. The media application 103 determines media item clusters such that media items within each cluster have a visual similarity based on the pixels of the images or videos from the media item collection. For example, FIG. 3A shows a first visual theme 300 of images (brown objects having a curved shape) having a visual similarity. Specifically, the first object is a drink with ice in a glass, the second object is a latte with a heart in a coffee cup, and the third object is a bowl made of brown wood of a different shade. Other examples can include mountains, natural arches, sea waves with people, horizontally extending parallel lines (e.g., train tracks, roads, etc.), changes over time (plant growth, sun movement, painting in progress), and the like.

[0020] A media item cluster can include images from the same episode, e.g., multiple images of the same work taken by the user from different angles. For example, FIG. 3A shows a second example 325 including three images. The three images are of the same still life taken in different ways such that the leaves of the tree become increasingly distinguishable. 含む第2の例325を示している。3つの画像は、木の葉が次第により区別できるように、異なる方法で撮影された同じ静物画である。

[0021] The media application 103 selects a subset of media item clusters based on the corresponding media items within each cluster that have a visual similarity within a visual similarity threshold. The visual similarity threshold may range from extremely similar media items to media items that are more similar than items that only have a distant relationship. For example, the theme of the first example 300 in Figure 3A is a brown circular object. This may be in the middle of the similarity threshold range. Conversely, the third example 350 in Figure 3A is a media item cluster with the theme of a cat in a shark stuffing, photographed at different times of day. This is a more visually similar theme. The fourth example 375 in Figure 3B includes the theme of an orange backpack used by people on different trips. Yet another example of very similar media closer to the threshold similarity value is when the media items are pink flowers of slightly different shapes.

[0022] If media items are not sufficiently visually similar, it can be difficult to identify themes between them, and as a result, the media items may appear more like a random collection of media items than what the user wants to see. In some embodiments, the media application 103 maintains greater visual theme consistency by limiting the number of media items so that the collection of media items does not appear, for example, to be a group of all cat images available from the user library.

[0023] The media application 103 can display a user interface that includes a subset of media item clusters. In some embodiments, the media application 103 displays the user interface that includes a subset of media item clusters at predetermined intervals. For example, the media application 103 may display the user interface that includes a subset of clusters daily, weekly, monthly, etc. The media application 103 may change the frequency of displaying the subset of clusters based on feedback. For example, if the user views the subset of clusters whenever it becomes available, the media application 103 may maintain the display frequency, but if the user views the subset of clusters less frequently, the media application 103 may reduce the display frequency.

[0024] The media application 103 can provide users associated with a user account with a notification that a subset of a cluster is available, along with the title corresponding to the cluster subset. For example, the media application 103 can provide users with daily, weekly, or monthly notifications. In some embodiments, the user interface includes options for limiting the frequency of notifications and / or the display of subsets of media item clusters.

[0025] Exemplary computing device 200 Figure 2 is a block diagram of an exemplary computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a user device 115 used to run a media application 103. In another example, the computing device 200 is a media server 101. In yet another example, the media application 103 is partially located on the user device 115 and partially on the media server 101.

[0026] One or more of the methods described herein can be performed as a standalone program running on any type of computing device, a program running on a web browser, or a mobile application (app) running on a mobile computing device (e.g., a mobile phone, smartphone, tablet computer, wearable device (e.g., a watch, armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted display), or laptop computer). In the primary example, all computations are performed by a mobile application on a mobile computing device. However, a client / server architecture can be used. For example, the mobile computing device sends user input data to a server device and receives and outputs (e.g., displays) the final output data from the server. In another example, the computations may be shared between the mobile computing device and one or more server devices.

[0027] In some embodiments, the computing device 200 includes a processor 235, memory 237, I / O interface 239, display 241, camera 243, and storage device 245. The processor 235 may be connected to a bus 218 via a signal line 222. The memory 237 may be connected to a bus 218 via a signal line 224. The I / O interface 239 may be connected to a bus 218 via a signal line 226. The display 241 may be connected to a bus 218 via a signal line 228. The camera 243 may be connected to a bus 218 via a signal line 230. The storage device 245 may be connected to a bus 218 via a signal line 232.

[0028] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling the basic operation of the computing device 200. “Processor” includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. The processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configurations), multiple processing units (e.g., having a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for achieving a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a system with a processor optimized for performing matrix calculations (e.g., matrix multiplication), or other systems. In some implementations, the processor 235 may include one or more coprocessors for performing neural network processing. In some implementations, the processor 235 may be a processor that generates a probabilistic output by processing data. For example, the output generated by the processor 235 may be inaccurate or accurate within the range of an expected output. The processing does not need to be limited to a specific geographical location or time. For example, a processor can perform functions in real time, offline, or batch mode. Parts of the processing may be performed by different (or the same) processing systems at different times and in different locations. The computer may be any processor that communicates with memory.

[0029] Memory 237 is typically located within the computing device 200 for use by the processor 235 and may be any suitable processor-readable storage medium for storing instructions executed by the processor or set of processors, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory. Memory 237 may be located separately from the processor 235 and / or integrated with it. Memory 237 stores software executed on the computing device 200 by the processor 235. It can store A (including media application 103).

[0030] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, a camera application, an image library application, an image management application, an image gallery application, a media display application, a communication application, a web hosting engine or application, a mapping application, a media sharing application, and the like. One or more of the methods disclosed herein may run on several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, and as a mobile application ("App") that runs on a mobile computing device.

[0031] The application data 266 may also be data generated by other applications 264 or hardware of the computing device 200. For example, the application data 266 may include images captured by the camera 243, user behavior identified by other applications 264 (e.g., a social networking application), and so on.

[0032] The I / O interface 239 can provide functionality that enables the computing device 200 to interface with other systems and devices. Interface devices may be included as part of the computing device 200 or may be separate but capable of communicating with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or database 199), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can be connected to interface devices, such as input devices (keyboard, pointing device, touchscreen, microphone, camera, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.). For example, when a user provides touch input, the I / O interface 239 transmits data to the media application 103.

[0033] Some exemplary interface connection devices that can be connected to the I / O interface 239 may include a display 241 that can be used to display the user interface of the content described herein, e.g., images, video and / or output applications, and to receive touch (or gesture) input from the user. For example, the display 241 may be used to display a user interface that includes a subset of media item clusters. The display 241 may include any suitable display device, e.g., a liquid crystal display (LCD), a light-emitting diode (LED), or a plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen on a computer device.

[0034] Camera 243 may be any type of image capture device capable of capturing images and / or videos. In some embodiments, camera 243 captures images or videos that the I / O interface 239 transmits to the media application 103. do.

[0035] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store a collection of media items associated with a user account, a subset of media clusters, a training set for a machine learning model, and so on. In embodiments where the media application 103 is part of the media server 101, the storage device 245 is the same as the database 199 in Figure 1.

[0036] Exemplary media application 103 Figure 2 shows an exemplary media application 103. The media application 103 includes a filtering module 202, a clustering module 204, a machine learning module 205, a selection module 206, and a user interface module 208. In some embodiments, the media application 103 uses either the clustering module 204 or the machine learning module 205.

[0037] The filtering module 202 excludes media items from the media item collection that correspond to categories in the prohibited category list. In some embodiments, the filtering module 202 includes a set of instructions that can be executed by the processor 235 to exclude media items that correspond to categories in the prohibited category list. In some embodiments, the filtering module 202 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0038] In some embodiments, the filtering module 202 excludes media from the media item collection before the clustering module 204 performs clustering. In alternative embodiments, the filtering module 202 excludes media from the media item collection after the clustering module 204 has performed clustering. For example, the filtering module 202 excludes media items associated with visual similarity from the prohibited category list. The prohibited category list may include media items taken as functional images, such as images of receipts, documents, parking meters, and screenshots, rather than for their photographic value.

[0039] In some embodiments where the media application 103 includes a machine learning module 205, the filtering module 202 excludes functional images from the media item collection before it is provided to the machine learning model. For example, the filtering module 202 excludes receipts, instruction manuals, documents, and screenshots before the media item collection is provided to the machine learning model.

[0040] The clustering module 204 determines media item clusters such that the media items in each cluster have visual similarity, based on the pixels of images or videos from the media item collection. In some embodiments, the clustering module 204 includes a set of instructions that can be executed by the processor 235 to generate media item clusters. In some embodiments, the clustering module 204 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0041] In some embodiments, the clustering module 204 is a user account The clustering module 204 accesses the media item collection associated with the user, for example, the library associated with the user. If the filtering module 202 excludes media items, the clustering module 204 accesses the media item collection that does not contain media items corresponding to the prohibited category list. The clustering module 204 can determine media item clusters such that the media items in each cluster have visual similarity based on the pixels of images or videos from the media item collection. In some embodiments, clustering determines visual similarity using an N-dimensional Gaussian diversity function.

[0042] In some embodiments, the machine learning module 205 includes a machine learning model trained to generate output image embeddings of media clusters such that media items within each cluster have visual similarity. In some embodiments, the machine learning module 205 includes a set of instructions that can be executed by the processor 235 to generate the image embeddings. In some embodiments, the machine learning module 205 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0043] In some embodiments, the machine learning module 205 can determine the visual similarity of clusters using vectors (embeddings) in a multidimensional feature space. Images with similar features may have similar feature vectors. For example, the vector distance between feature vectors of similar images may be smaller than the vector distance between dissimilar images. The feature space may be a function of various factors of the image, such as the subject depicted (objects detected from the image), the composition of the image, color information, the orientation of the image, the metadata of the image, and specific objects recognized from the image (e.g., known faces, if user permission is granted).

[0044] In some embodiments, training may be performed using supervised learning. In some embodiments, the machine learning module 205 includes a set of instructions that can be executed by the processor 235. In some embodiments, the machine learning module 205 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0045] In some embodiments, the machine learning module 205 can generate a trained model, specifically a machine learning model, using training data (obtained with permission for training). For example, the training data may include ground truth data in the form of media clusters associated with visual similarity descriptions of clusters. In some embodiments, the visual similarity descriptions may include user feedback on whether the clusters are related and contain a clear theme. In some embodiments, the visual similarity descriptions may be automatically added by image analysis. The training data may be obtained from any source, for example, a data repository specifically designated for training, or data for which permission has been granted to use as training data for machine learning.

[0046] In some embodiments, the training data may include synthetic data generated for training purposes, such as data not based on activities in the situation being trained, such as data generated from simulations or computer-generated images / videos. In some embodiments, the machine learning module 205 uses weights obtained from another application and not edited / transferred. For example, in these embodiments, the trained model may be generated on a different device and provided as part of a media application 103. In various embodiments, the trained model has a model structure or form (defining, for example, the number and types of neural network nodes, the connections between nodes, and organizing the nodes into multiple layers) and associated weights. The data may be provided as a data file containing the data. The machine learning module 205 can read the data file of the trained model and implement a neural network, including node connections, layers, and weights, based on the model structure or form specified in the trained model.

[0047] The machine learning module 205 generates a trained model, referred to herein as an event machine learning model. In some embodiments, the machine learning module 205 is configured to identify one or more features within input media items and generate feature vectors (embeddings) representing the media items by applying the event machine learning model to data such as application data 266 (e.g., input media). In some embodiments, the machine learning module 205 may include software code executed by the processor 235. In some embodiments, the machine learning module 205 may specify a circuit configuration (e.g., a programmable processor, a field-programmable gate array (FPGA)) that enables the processor 235 to apply the machine learning model. In some embodiments, the machine learning module 205 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the machine learning module 205 may provide an application programming interface (API). The operating system 262 and / or other applications 264 can use this API to call the machine learning module 205 and output image embeddings of media clusters, for example, by applying the machine learning model to application data 266. In some embodiments, media items that match visual similarity are closer to each other in vector space than dissimilar images. Therefore, media item clusters are generated by dividing the vector space.

[0048] In some embodiments, the machine learning model is a classifier that receives a media item collection as input. Examples of classifiers include neural networks, support vector machines, k nearest neighbors, logistic regression, naive Bayes, decision trees, and perceptrons.

[0049] In some embodiments, a machine learning model may include one or more model forms or structures. For example, a model form or structure may include any kind of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., "hidden layers" between input and output layers, each of which is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results obtained from processing each tile), or a sequence-to-sequence neural network (e.g., a network that takes sequential data as input, such as words in a sentence or frames in a video, and produces a result sequence as output).

[0050] The model configuration or structure can specify the connections between various nodes and the organization of nodes into layers. For example, the nodes in the first layer (e.g., the input layer) can receive data as input data or application data. For example, if a machine learning model is used to analyze input images associated with a user account, e.g., the first image, such data may include, for example, one or more pixels per node. Subsequent intermediate layers can receive the outputs of the nodes in the previous layer as input, according to the connections specified in the model configuration or structure. These layers are sometimes called hidden layers. The final layer (e.g., the output layer) generates the output of the machine learning model. For example, this output may be an image embedding of a media cluster. In some embodiments, The model form or structure specifies the number and / or types of nodes in each layer.

[0051] Features output by machine learning module 205 may include the subject (e.g., sunset vs. a specific person), the colors present in the image (e.g., green hills vs. a blue lake), color balance, lighting source, angle and intensity, the location of objects in the image (e.g., adhering to the rule of thirds), the relative locations of objects (e.g., depth of field), the shooting location, focus (foreground vs. background), or shadows. While the aforementioned features are human-readable, the output features may be representative of the image and may be embeddings or other numerical values ​​that are not human-analyzable (e.g., individual feature values ​​may not correspond to specific features such as present colors or object locations). However, because the trained model is robust to images, it will output similar features for similar images and different features for significantly different images.

[0052] In some embodiments, the model configuration is a CNN including network layers, each network layer extracting image features at a different level of abstraction. The CNN used to identify features in an image may also be used to classify the image. The model architecture may include combinations and sequences of layers consisting of multidimensional convolution, mean pooling, max pooling, activation functions, normalization, regularization, and other layers and modules actually used in applied deep neural networks.

[0053] In different embodiments, a machine learning model may include one or more models. One or more models may include multiple nodes arranged in layers according to the model structure or form. In some embodiments, a node may be a memoryless computation node configured to process one unit of input and produce one unit of output. The computation performed by the node may include, for example, the steps of multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and producing a node output by adjusting the weighted sum with a bias value or intercept value. For example, the machine learning module 205 may adjust each weight based on feedback in response to automatically updating one or more parameters of the machine learning model.

[0054] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to a tuned weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multicore processor, using individual processing units of a graphical processing unit (GPU), or using dedicated neural circuits. In some embodiments, the nodes may include memory. The nodes may, for example, store one or more previous inputs and use one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a long-short-term memory (LSTM) node. The LSTM node can use memory to maintain state that allows the node to behave like a finite state machine (FSM). Models including such nodes can handle sequential data, for example, multiple words in a sentence or paragraph, a series of images, etc. This would be useful when processing frames, conversations, or other audio within a video. For example, heuristic-based models used in gating models can remember one or more previously generated features for a given image.

[0055] In some embodiments, the machine learning model may include embeddings or weights for individual nodes. For example, the machine learning model may be initialized as a group of nodes organized into layers as specified by the model form or structure. At initialization, each pair of nodes connected according to the model form, for example, each node in a continuous layer of a neural network, is initialized. Each weight can be applied to the connection between pairs. For example, each weight may be assigned randomly or initialized to a default value. Then, for example, a machine learning model can be trained using a training set of media clusters to produce results. In some embodiments, a subset of the entire architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.

[0056] For example, training may involve applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., media items from a collection of media items associated with a user account) and expected outputs corresponding to each input (e.g., image embeddings in a media cluster). For example, the weight values ​​may be automatically adjusted based on a comparison between the machine learning model's output and the expected output to increase the probability that the machine learning model will produce the expected output when given similar inputs.

[0057] In some embodiments, training may include applying unsupervised learning techniques. In unsupervised learning, only input data (e.g., media items from a collection of media items associated with a user account) may be provided, and the machine learning model may be trained to distinguish the data, for example, to cluster image features into multiple groups.

[0058] In various embodiments, the trained model includes a set of weights corresponding to the model structure. In embodiments where the training set is omitted, the machine learning module 205 may generate a machine learning model based on prior training by, for example, the developer of the machine learning module 205 or a third party. In some embodiments, the machine learning model may include a set of fixed weights downloaded from a server that provides the weights.

[0059] In some embodiments, the machine learning module 205 may be implemented offline. Implementing the machine learning module 205 may include using a static training set that does not include updates when the data in the static training set changes. This is advantageous as it improves the efficiency of the processing performed by the computing device 200 and reduces the power consumption of the processing device 200. In these embodiments, the machine learning model may be generated in a first stage and provided as part of the machine learning module 205. In some embodiments, small updates to the machine learning model may be implemented online, where updates to the training data are included as part of training the machine learning model. A small update is an update with a size smaller than a threshold size. The size of the update is related to the number of variables in the machine learning model that the update affects. In such embodiments, an application calling the machine learning module 205 (e.g., operating system 262, one or more other applications 264, etc.) can use image embeddings of media item clusters to identify visually similar clusters. The machine learning module 205 may also generate system logs periodically, for example, every hour, every month, or every three months. The system logs may be used to update the machine learning model, for example, to update the embeddings of the machine learning model.

[0060] In some embodiments, the machine learning module 205 may be implemented in a manner that conforms to a specific configuration of the computing device 200 on which the machine learning module 205 is executed. For example, the machine learning module 205 can determine the computation graph that utilizes available computing resources, such as the processor 235. If the machine learning module 205 is implemented as a distributed application on multiple devices, for example, if the media server 101 includes multiple media servers 101, the machine learning module 205 may... The calculations performed on individual devices can be determined to optimize the computations. In another example, if the machine learning module 205 determines that the processor 235 contains a GPU with a certain number (e.g., 1000) GPU cores, the machine learning module 205 can be implemented (e.g., as 1000 separate processes or threads).

[0061] In some embodiments, the machine learning module 205 can implement a set of trained models. For example, an event machine learning model may include multiple trained models, each applicable to the same input data. In these embodiments, the machine learning module 205 can select a particular trained model based, for example, on available computational resources, the success rate of previous inferences, etc.

[0062] In some embodiments, the machine learning module 205 can run multiple trained models. In these embodiments, the machine learning module 205 can synthesize outputs, for example, by using a majority vote to score the outputs obtained by applying each trained model, or by selecting one or more specific outputs. In some embodiments, such selectors are part of the model itself and function as a connecting layer between the trained models. Furthermore, in these embodiments, the machine learning module 205 can apply a time threshold (e.g., 0.5 ms) for applying each trained model and utilize only the individual outputs available within the time threshold. Outputs not received within the time threshold are not utilized and may, for example, be discarded. For example, such a technique would be appropriate when there is a specified time limit between calling the machine learning module 205, for example, by the operating system 262 or one or more other applications 264. In this way, the maximum time required for the machine learning module 205 to perform a task, for example, to identify one or more features of an input media item and generate a feature vector (embedding) representing the media item, can be limited, thereby improving the responsiveness of the media application 103, and as a result, the machine learning module 205 can provide the best classification in real time.

[0063] In some embodiments, the machine learning module 205 receives feedback. For example, the machine learning module 205 can receive feedback from one user or a group of users via the user interface module 208. If one user provides feedback, the machine learning module 205 provides the feedback to the machine learning model, which uses the feedback to update the parameters of the machine learning model to modify the output image embedding of the media item clusters. If a group of users provides feedback, the machine learning module 205 provides aggregated feedback to the machine learning model, which uses the aggregated feedback to update the parameters of the machine learning model to modify the output image embedding of the media item clusters. For example, the aggregated feedback may include a subset of media clusters and user responses to that subset of media clusters. User responses include viewing only one image and refusing to view the rest of the media, viewing all corresponding media items within a subset, sharing corresponding media items, providing approval or disapproval instructions for corresponding media items (e.g., thumbs up / down, like, +1, etc.), deleting / adding individual media items from a subset of the media item cluster, and changing titles. The machine learning module 205 can modify the media cluster based on updates to the parameters of the machine learning model.

[0064] In some embodiments, machine learning models are trained using user feedback, which is used to respond to a subset of clusters and within that subset. This includes changing the title of one of the clusters. Machine learning module 205 provides feedback to the machine learning model to modify parameters to exclude media item clusters with a specific type of visual similarity (for example, images of ocean waves that are visually similar but are not the type of media the user wants to see, and images of surfers on waves taken at different times and / or locations).

[0065] The selection module 206 selects a subset of media item clusters based on the visual similarity determined by the clustering module 204. In some embodiments, the selection module 206 includes a set of instructions that can be executed by the processor 235 to select a subset of media item clusters. In some embodiments, the selection module 206 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0066] In some embodiments, the selection module 206 selects a subset of media item clusters whose media items have a visual similarity within a visual similarity threshold range. For example, the range may be between 0.05 and 0.3, out of 0 to 4. Other ranges and scales are also possible. The subset of media item clusters within the visual similarity threshold range may be considered to have a visual theme that is perceived as relevant and cohesive.

[0067] In some embodiments, if a media item cluster exceeds a predetermined number (e.g., more than 15 media items), the selection module 206 can impose additional restrictions when selecting a subset of the media item cluster. For example, the selection module 206 can impose temporal diversity by identifying the timestamp associated with each media item, identifying based on the timestamps that multiple media items (e.g., media items associated with the same time period and location) are associated with the same episode, and selecting a subset of media item clusters based on the associated timestamps to satisfy a temporal diversity criterion that excludes more than a predetermined number of media items from a particular episode (i.e., selecting a subset of media item clusters based on the associated timestamps so that no media items of a certain number (e.g., 3) or less are associated with the same episode). This avoids media item clusters that may be too similar and duplicate, as a user may take multiple images of an object in the same place at the same time. This also avoids situations where a user takes the same image and edits it for posting to, for example, another photo-sharing application. The selection module 206 can use temporal diversity to select a subset of clusters that show the progress of an object over a period of time. For example, a cluster could include different images of a child taken at different time periods to show how the child grows, or different images of a plant from a seedling to a flowering bush.

[0068] In some embodiments, the selection module 206 imposes spatial diversity on a subset of media clusters. For example, the selection module 206 can identify a location associated with each media item, and if the number of corresponding media items available in a cluster exceeds a predetermined number (e.g., more than 10 media items), the selection module 206 selects a subset of media item clusters based on location so that the subset of media item clusters satisfies the spatial diversity criterion. Figure 4 includes an example 400 of a visual theme of nature images of different mountain ranges, which have both temporal diversity because the images were taken in different months and years, and spatial diversity because the images were taken in different locations. The images have two kinds of diversity, but a hidden similarity emerges through the visual theme.

[0069] In some embodiments, the selection module 206 assigns a semantic theme to a subset of media clusters. The selection module 206 can identify labels associated with images and group subsets of media item clusters based on corresponding media items having the same or similar labels. For example, the selection module 206 can select a subset of media item clusters ranging from puppies to adult dogs using labels that identify depictions of dogs in images. In some embodiments, the media application 103 combines the semantic theme of the Golden Gate Bridge with the visual theme of other bridges that are visually similar to the golden color of the Golden Gate Bridge.

[0070] In some embodiments, the selection module 206 scores each media item within a subset of media item clusters based on an analysis of the likelihood that a user associated with a user account will refer to the media item and take a positive action. Positive actions may include browsing the subset, sharing the subset, or ordering prints from the subset. The selection module 206 may score a media item as being more likely to be associated with a positive action by a user associated with a user account if the subject matter is more interesting, for example, if the subject matter includes a baby, an acquaintance of the user, or a place the user has visited. Conversely, the selection module 206 may determine that a user is less likely to take a positive action related to a particular subject, such as a static object like a bunk bed. In some embodiments, the selection module 206 scores a subset of media item clusters based on personal information related to the user or on aggregated information about the user's general reaction to the media. In some embodiments, the selection module 206 scores a media item based on its quality, such as being too blurry, because the quality of the media item reduces the likelihood that a user associated with a user account will take a positive action associated with the media item.

[0071] The selection module 206 can select media items within a subset of the media item cluster if the score corresponding to each media item satisfies a threshold score. In some embodiments, the threshold score is a static value that is the same for all users. In some embodiments, the threshold score is unique to the user. In some embodiments, the threshold score is specified by the user.

[0072] Once the selection module 206 has determined a subset of the cluster, it can instruct the user interface module 208 to display a user interface containing the subset of the cluster. In some embodiments, the user can provide feedback related to the subset of the cluster. For example, the user can browse the subset, provide instructions indicating approval of the subset, share the subset, and order prints of photos from the subset.

[0073] In some embodiments, the selection module 206 receives feedback and modifies the score corresponding to a subset of media item clusters based on this feedback. For example, the feedback may include explicit actions indicated by removing a subset of clusters from the user interface, or implicit actions indicated by one or more of the following: browsing a subset of clusters, viewing a subset of clusters, or sharing a subset of clusters. In some embodiments, the selection module 206 may identify patterns in the feedback. For example, if positive feedback occurs when objects in a cluster are of a particular type (e.g., babies, families, trees), the selection module 206 may modify the score so that the subset of clusters includes objects of similar types. In another example, the pattern may indicate that the user prefers themes with lower visual similarity to those with higher visual similarity, and the selection module Joule 206 can modify its score to more frequently select themes with lower visual similarity.

[0074] In some embodiments, the selection module 206 can receive and aggregate feedback from a set of users using the media application 103. For example, the selection module 206 can create aggregated user feedback for a subset of media clusters and modify the scoring based on the aggregated feedback.

[0075] The user interface module 208 generates a user interface. In some embodiments, the user interface module 208 includes a set of instructions that can be executed by the processor 235 to generate the user interface. In some embodiments, the user interface module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0076] The user interface module 208 displays a user interface that includes a subset of media clusters. Figure 5 shows an example of a user interface 500 that includes a cluster 505 having a visual theme, according to some embodiments described herein. In this example, cluster 505 is displayed at the top of the user interface along with a group of recent highlights and a group of images from a year ago. The user interface 500 also includes images taken in San Francisco yesterday (March 9).

[0077] In some embodiments, the user interface module 208 suggests a subset of clusters and generates a user interface for browsing, editing, and sharing media. As shown in Figure 5, the user interface may include, for example, clusters located at the top of the user interface, and when the user selects an image, the user interface includes options for editing or sharing the image.

[0078] In some embodiments, in response to a user selecting a cluster within the user interface, the user interface module 208 displays corresponding media items from the cluster at predetermined intervals. For example, the user interface module 208 may display each media item at intervals of 2 seconds, 3 seconds, etc.

[0079] In some embodiments, the user interface module 208 provides cover photos for a subset of clusters. The cover photos may be the most recent photos, the photos with the highest scores, and so on. In some embodiments, the user interface module 208 selects a specific media item from each cluster within a subset of media item clusters as the cover photo for each cluster within the subset of media item clusters, based on a specific media item containing the maximum number of objects corresponding to the visual similarity. For example, a cluster may have a visual theme of groups of people skiing, and the user interface module 208 may select a cover photo for the cluster that shows an image depicting the highest number of people from the cluster of people skiing. In another example, if a cluster has a visual theme of people engaged in outdoor water activities, the user interface module 208 may determine that an image of people surfing is the most representative media item for the cover, compared to other images where people are near the water rather than in it (e.g., building sandcastles) or where people are not engaged in more active outdoor activities (e.g., sunbathing along the water). The user interface module 208 may also select the cover photo based on having the highest visual quality within the cluster (e.g., sharp, high resolution, not blurry, good exposure, etc.). It is also possible to do so.

[0080] In some embodiments, the user interface module 208 adds a title to each cluster within a subset of media item clusters based on the type of visual theme and / or template representation. For example, the title could be an action that occurred in the image (e.g., "surf's up" for the ocean cluster, "into the blue" for the sky cluster, "on the road" for the road cluster, "stairway to heaven" for the church image), a food metaphor ( For example, "mixed nuts," "smorgasbord," "mixed bag," "goody bag," "wine flight," "cheese pairing," "sampler," "treasure trove," "overlooked treasures," "have a drink"), photo trails (for example, "photo detective," "photo mystery," "mystery photos," "photo sphinx"), creative combinations, correlations ( For example, connections, feather photos, photo club, photo weaving, patterned, coincidence, cause and effect, slot machine, one of these is the same as the other, similar), titles that refer to patterns (e.g., beta pattern, pattern hunter, small pattern, pattern portal, dot linking, picture pattern, photo pattern), synonyms for patterns, for example, themes (e.g., photo story, picture story, photo narrative, two-photo story, photo theme, lucky theme), sets (e.g., photo set, surprise set), or matches (e.g., memory match), onomatopoeia (e.g., zigzag, boom, bang bang, smack, smack photo, photo smack), verbs (e.g., "look what we found in the couch cushions", "look what appeared", "help us sleuth", "will it blend") The title of a cluster may refer to a connection that the selection module will have a higher confidence score of inferring correctness, such as "time flies," "some things never change," or "magic pattern," "lightheartedness." In some embodiments, the template expression may be a more colloquial, engaging, and humorous expression or a common expression than simply adding a title such as "Birthday 1997-2001." In some embodiments, the user interface module 208 may include a title with a general title such as "Look what we found" and a theme-specific subtitle such as "your orange backpack brought you far," with respect to the fourth example 375 in Figure 3B.

[0081] In some embodiments, the user interface module 208 provides users associated with a user account with a notification that a subset of clusters is viewable. The user interface module 208 can provide notifications periodically, such as daily, weekly, or monthly. In some embodiments, if users stop viewing notifications when they are provided daily (weekly, monthly, etc.), the user interface module 208 can generate notifications less frequently. The user interface module 208 can additionally provide notifications that include titles corresponding to the subset of clusters.

[0082] Exemplary flowchart Figure 6 is a flowchart illustrating an exemplary method 600 for displaying a subset of media item clusters according to several embodiments. The method shown in flowchart 600 may be performed by the computing device 200 of Figure 2.

[0083] Method 600 may begin with block 602. In block 602, a request is generated to seek access to the media item collection associated with the user account. In some embodiments, the request is generated by the user interface module 208. Block 604 may be executed after block 602.

[0084] In block 604, the authorization interface element is displayed. For example, user input The interface module 208 can display a user interface that includes permission interface elements for requesting the user to provide permission to access the media item collection. Block 606 can be executed after Block 604.

[0085] In block 606, it is determined whether permission has been granted by the user to access the media item collection. In some embodiments, block 606 is performed by the user interface module 208. If the user has not provided permission, the method terminates. If the user has provided permission, block 608 may be executed after block 606.

[0086] In block 608, media item clusters are determined such that the media items within each cluster have visual similarity, based on the pixels of images or videos from the media item collection. The media item collection is associated with a user account. In some embodiments, block 606 is performed by the clustering module 204. Block 610 may be performed after block 608.

[0087] In block 610, a subset of media item clusters is selected based on the corresponding media items within each cluster having a visual similarity within a visual similarity threshold range. In some embodiments, block 610 is performed by the selection module 206. Block 612 can be performed after block 610.

[0088] In block 612, a user interface including a subset of media clusters is displayed. In some embodiments, block 610 is executed by the user interface module 208.

[0089] Figure 7 is a flowchart illustrating an exemplary method 700 for generating embeddings for media item clusters using a machine learning model and selecting a subset of media item clusters, according to several embodiments. The method shown in flowchart 700 may be performed by the computing device 200 of Figure 2.

[0090] Method 700 may begin with block 702. In block 702, a request is generated to seek access to the media item collection associated with the user account. In some embodiments, the request is generated by the user interface module 208. Block 704 may be executed after block 702.

[0091] In block 704, the authorization interface element is displayed. For example, user interface module 208 may display a user interface that includes an authorization interface element for requesting the user to provide permission to access a media item collection. Block 706 can be executed after block 704.

[0092] In block 706, it is determined whether permission has been granted by the user to access the media item collection. In some embodiments, block 706 is performed by the user interface module 208. If the user has not provided permission, the method terminates. If the user has provided permission, block 708 may be executed after block 706.

[0093] In block 708, the trained machine learning model receives media items from the media item collection associated with the user account as input. In some embodiments, block 708 is executed by the machine learning module 205. Block 710 can be executed after block 708.

[0094] In block 710, the trained machine learning model generates output image embeddings of media item clusters. Media items within each cluster have visual similarity, and media items with visual similarity are closer to each other than media items that are dissimilar in the vector space, so as to generate media item clusters by partitioning the vector space. In some embodiments, block 710 is performed by the machine learning module 205. Block 712 can be performed after block 710.

[0095] In block 712, a subset of media item clusters is selected based on the corresponding media items within each cluster that have a visual similarity within a visual similarity threshold. In some embodiments, block 712 is performed by the machine learning module 205. Block 714 can be performed after block 712.

[0096] In block 714, a user interface containing a subset of media item clusters is displayed. In some embodiments, block 714 is executed by the user interface module 208.

[0097] In addition to the above description, the systems, programs, or functions described herein may give the user control over whether and when they enable the collection of user information (e.g., information about the user's media items such as photographs or videos, the user's interactions with media applications that display media items, the user's social networks, social behavior or activities, occupation, viewing preferences for image-based creations, settings for hiding people or pets, user interface preferences, or information about the user's current location), and whether they transmit content or information from the server. Furthermore, certain data may be processed to remove personally identifiable information in one or more ways before being stored or used. For example, the user's ID may be processed so that the user's personal information cannot be identified. Also, when obtaining location information (e.g., city, zip code, or state level) so that the user's location cannot be identified, the user's geographic location may be generalized. Thus, the user can control what user information is collected, how the information is used, and what information is provided to the user.

[0098] In the above description, for the purpose of explanation, many specific details are provided to give a complete understanding of the various embodiments described. However, it will be apparent to those skilled in the art that the various embodiments described can be carried out even without these specific details. In some cases, structures and devices are shown in block diagrams to avoid obscuring the description. For example, embodiments can be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0099] In this specification, any reference to “some embodiments” or “some instances” means that certain features, structures, or characteristics described in relation to an embodiment or instance may be included in at least one implementation of the description. The phrase “in some embodiments” in various places in this specification does not necessarily refer to the same embodiment.

[0100] Some parts of the detailed description above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic descriptions and representations are means used by those skilled in the field of data processing to communicate the nature of their work to others skilled in the field in the most effective way. In this specification, an algorithm is generally considered to be a set of consistent steps that produce a desired result. These steps require the physical manipulation of physical quantities. These quantities usually take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and other manipulated. In some cases, and primarily for reasons of common use, it is convenient to refer to these data as bits, numbers, elements, symbols, characters, terms, digits, etc.

[0101] It is important to understand that all these and similar terms are merely convenient labels associated with and applied to appropriate physical quantities. Unless otherwise specified or as evident from the discussion, any discussion using terms including “process,” “operate,” “calculate,” “determine,” or “display” throughout the description refers to the operation and processes of a computer system or similar electronic computing device that processes and transforms data represented as physical quantities in computer system memory, registers, or other information storage devices, transmission devices, or display devices.

[0102] Embodiments of this specification also relate to processors for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively started or reconfigured by a computer program stored in the computer. Such computer programs may be stored in non-temporary computer-readable storage media, including, but not limited to, optical discs, ROMs, CD-ROMs, magnetic disks, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory including USB keys with non-volatile memory, or any type of medium suitable for storing electronic instructions, each of which is connected to a computer system bus.

[0103] This specification may include several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, and microcode.

[0104] Furthermore, the description may take the form of a computer program product accessible from a computer-enabled or computer-readable medium that provides program code used by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-enabled or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program used by or in connection with an instruction execution system, machine, or apparatus.

[0105] A data processing system suitable for storing or executing program code includes at least one processor directly or indirectly connected to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, mass storage, and cache memory providing temporary storage for at least some program code to reduce the number of times the code must be retrieved from mass storage during execution.

Claims

1. A method performed by a computer, Based on the pixels of images or videos from the media item collection, determine the image embedding of media item clusters such that the media items within each cluster have visual similarity. Each media item is associated with a location and associated timestamp. Media items acquired within a specified period are associated with the episode. The aforementioned media item collection is associated with a user account. The aforementioned method, Selecting the subset of media item clusters based on corresponding associated timestamps such that the corresponding media items in each cluster having a visual similarity within a visual similarity threshold, and the corresponding media items in a subset of the media item clusters, satisfy a temporal diversity criterion that excludes more than a first predetermined number of the corresponding media items from the episode. A method comprising, in response to the number of corresponding media items in a subset of the cluster being greater than a second predetermined number of media items, excluding one or more of the corresponding media items based on locations such that the subset of the media item cluster satisfies a spatial diversity criterion.

2. Displaying a user interface that includes the subset of the media item cluster, Receiving aggregated feedback from users for aggregated subsets of media item clusters, The method includes providing the aggregated feedback to a machine learning model, wherein the parameters of the machine learning model are updated based on the aggregated feedback, and the method is The method according to claim 1, further comprising modifying the image embedding of the media item cluster using the parameters of the machine learning model having the updated parameters.

3. The method according to claim 1 or 2, further comprising excluding from the media item collection media items associated with categories in a prohibited category list before selecting the subset of the media item cluster.

4. The method according to any one of claims 1 to 3, further comprising excluding media items corresponding to categories in a prohibited category list before determining the media item cluster.

5. The method according to any one of claims 1 to 4, wherein the media item cluster is further determined based on the corresponding media items associated with labels having semantic similarity.

6. Based on an analysis of the likelihood that a user associated with the user account will refer to the media item and take a positive action, each media item within the subset of the media item cluster is scored. The method according to any one of claims 1 to 5, further comprising selecting the media items from the subset of the media item cluster based on corresponding scores that satisfy a threshold score.

7. Receiving feedback from the user regarding one or more media items within the subset of the media item cluster, The method according to claim 6, further comprising modifying the corresponding scores of one or more media items in the subset of the media item cluster based on the feedback.

8. The method according to claim 7, wherein the feedback includes an explicit action indicated by removing one or more media items from the subset of the media item cluster from the user interface, or an implicit action indicated by viewing the corresponding media item in the subset of the media item cluster, or sharing the corresponding media item in the subset of the media item cluster.

9. The method according to any one of claims 1 to 8, further comprising updating the user interface based on changing the image embedding of the media item cluster.

10. The method according to any one of claims 1 to 9, further comprising selecting the particular media item as the cover photograph for each cluster in the subset of the media item cluster, based on a particular media item containing the maximum number of objects corresponding to the visual similarity.

11. The method according to any one of claims 1 to 10, further comprising adding a title to each cluster within the subset of media item clusters based on the type of visual similarity and common representation.

12. The method according to any one of claims 1 to 11, wherein the subset of the media item cluster is displayed on the user interface at predetermined intervals.

13. The further includes providing a notification to the user associated with the user account that the subset of the media item cluster is available, The method according to any one of claims 1 to 12, wherein the notification includes a title corresponding to each of the clusters in the subset of the media item cluster.

14. The determination includes the step of generating a vector representation of each media item using a trained machine learning model, The vector distance between the vector representations of a pair of media items indicates the visual similarity of the media items. The vector representation is an image embedding generated by the trained machine learning model, The method according to any one of claims 1 to 13, wherein the cluster is selected such that the vector distance between each pair of media items within the cluster is outside the range of the visual similarity threshold.

15. A method performed by a computer, This includes receiving media items from a media item collection associated with a user account as input to a trained machine learning model, where each media item is associated with an associated timestamp, and media items acquired within a predetermined period are associated with an episode. The method includes generating output image embeddings of media item clusters using the aforementioned trained machine learning model, wherein the media items within each cluster have visual similarity, and the vector space is divided such that media items having visual similarity are closer to each other than media items that are not similar in the vector space, thereby generating the media item clusters, and each media item is associated with a location, and the method is Selecting a subset of media item clusters based on corresponding associated timestamps such that the corresponding media items within each cluster having a visual similarity within a visual similarity threshold, and the corresponding media items within a subset of media item clusters, satisfy a temporal diversity criterion that excludes more than a first predetermined number of the corresponding media items from the episode; A method comprising, in response to the number of corresponding media items in a subset of the cluster being greater than a second predetermined number of media items, excluding one or more of the corresponding media items based on locations such that the subset of the media item cluster satisfies a spatial diversity criterion.

16. The method according to claim 15, wherein functional images are removed from the media item collection before the media item collection is provided to the trained machine learning model.

17. The method according to claim 15 or 16, wherein aggregated user feedback includes reactions to a media item set or changes to the title of the media item set.

18. It is a system, Processor and The system comprises a memory connected to the processor, and the memory, when executed by the processor, stores instructions that cause the processor to perform the following operations: The aforementioned operation is, The operation includes determining image embedding of media item clusters such that the media items within each cluster have visual similarity, based on the pixels of images or videos from the media item collection, each media item is associated with a location and associated timestamp, media items acquired within a predetermined period are associated with an episode, the media item collection is associated with a user account, and the operation is performed as follows: Selecting a subset of media item clusters based on corresponding associated timestamps such that the corresponding media items in each cluster having a visual similarity within a visual similarity threshold, and the corresponding media items in a subset of media item clusters, satisfy a temporal diversity criterion that excludes more than a first predetermined number of the corresponding media items from the episode. A system comprising, in response to the number of corresponding media items in a subset of the cluster being greater than a second predetermined number of media items, excluding one or more of the corresponding media items based on locations such that the subset of the media item cluster satisfies a spatial diversity criterion.

19. Each media item has an associated timestamp. The media items acquired within a specified period are associated with an episode. The system according to claim 18, wherein the selection of the subset of the media item cluster is performed based on corresponding associated timestamps such that the corresponding media items in the subset of the media item cluster satisfy a second temporal diversity criterion which excludes more than a predetermined number of the corresponding media items from a particular episode.

20. The system according to claim 18 or 19, wherein the operation further comprises excluding from the media item collection any media items associated with categories in a prohibited category list before selecting the subset of the media item cluster.

21. The system according to any one of claims 18 to 20, wherein the media item cluster is further determined based on the corresponding media item associated with a label having semantic similarity.

Citation Information

Patent Citations

  • Digital asset search user interface

    CN112088370A

  • Retrieval method of web page and clustering method of web page

    JP2007080061A

  • Array-based media item discovery

    JP2009530741A

  • Method and apparatus for navigating image data set, and program

    JP2011154687A

  • Image selection suggestions

    JP2021504803A