Automatic generation of event using machine learning model

The event machine learning model segments media libraries into events, improving classification and reducing power consumption, enabling efficient organization and access of media based on events.

JP2025148321APending Publication Date: 2025-10-07GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025087703
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-17
Filing Date
2025-05-27
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Users face difficulty in organizing large media libraries containing thousands of images and videos taken over a long period, making it challenging to reminisce about specific events like birthdays or vacations.

Method used

A computer-implemented method using an event machine learning model to segment media libraries into episodes, generate event signals, and determine event importance scores, providing a user interface to display and manage events effectively.

Benefits of technology

The method improves event classification accuracy, reduces power consumption, and enhances efficiency by using a static training set and updating the model only when necessary, allowing users to easily access and manage their media based on events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025148321000001_ABST
    Figure 2025148321000001_ABST
Patent Text Reader

Abstract

To provide an automatic generation system for an event using a machine learning model.SOLUTION: A media application segments a media library associated with a user account into episodes for which each episode is associated with a corresponding time period, and uses an event machine learning model to generate an event signal indicating a likelihood that an event has occurred in each episode. The event machine learning model is a classifier that receives media as input. The media application also generates an event importance score for each episode, determines one or more events from the episode based on an event signal and a corresponding event importance score that exceeds an event importance threshold, and provides a user interface that includes corresponding media from the one or more events.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 17 / 404773, filed August 17, 2021, entitled "Automatic Generation of Events Using Machine Learning Models," which claims priority to both U.S. Provisional Patent Application No. 63 / 187392, filed May 11, 2021, entitled "Automatic Generation of Events Using Machine Learning Models," and U.S. Provisional Patent Application No. 63 / 189657, filed May 17, 201, entitled "Automatic Generation of Events Using Machine Learning Models," each of which is incorporated herein in its entirety. Summary of the Invention [Problem to be solved by the invention]

[0002] background Users of devices such as smartphones or other digital cameras capture and store large amounts of media (e.g., photos and videos) in libraries. Users access their libraries to view their media to reminisce about various events, such as birthdays, weddings, vacations, and trips. However, libraries often contain thousands of images taken over a long period of time and are difficult to organize.

[0003] The discussion of the background art set forth herein is intended to provide a general context for the present disclosure. To the extent that it is set forth in this background art section, the work of the presently named inventors, as well as statements that do not qualify as prior art at the time of filing, are not expressly or implicitly admitted as prior art to the present disclosure. [Means for solving the problem]

[0004] overview The computer-implemented method includes segmenting a media library associated with a user account into episodes, each episode associated with a corresponding time period, and generating an event signal indicating a likelihood that an event occurred in each episode using an event machine learning model, the event machine learning model being a classifier that receives media as input, the method further including generating an event importance score for each episode, determining one or more events from the episode based on the event signal and the corresponding event importance score that exceeds an event importance threshold, and providing a user interface including an event item having corresponding media from a particular event of the one or more events.

[0005] In some embodiments, the user interface is part of a reverse-chronological grid of media including corresponding media, and the method further includes, in response to a user selecting an event item from the reverse-chronological grid, replacing the reverse-chronological grid with a display of corresponding media from the particular event for a predetermined period of time, and in response to completion of the display of the corresponding media, displaying the reverse-chronological grid. In some embodiments, the method further includes generating an event machine learning model offline using a static training set, and updating the event machine learning model in response to an update size being less than a threshold size. In some embodiments, the method further includes receiving a request from a user to remove depictions of people from media in the media library that include depictions of people, and filtering the depictions of people from the media, the filtering being performed before generating the event significance score. In some embodiments, the method further includes determining computations to be performed on individual devices to optimize the computations, and updating the individual device. and implementing the event machine learning model on the plurality of devices based on computations performed on the devices. In some embodiments, the user interface includes an option to hide the event item. In some embodiments, generating the event importance score is based on one or more of at least a threshold number of media items in the corresponding episode, at least a threshold number of face clusters in the corresponding episode, a quality metric for the media items in the corresponding episode, at least one face cluster of a threshold rank, or the presence of a rare face cluster. In some embodiments, the method further includes merging the plurality of episodes into a single event based on one or more of the plurality of episodes each associated with a plurality of time periods, wherein the plurality of time periods generally fall within a 24-hour period, and the single event is a type of event occurring across multiple days, celebrations on different days associated with a single event, multiple episodes associated with the same location, or multiple episodes associated with the same set of face clusters. In some embodiments, the method further includes generating a confidence score indicating a likelihood that the corresponding event is the accurately recognized event type, and adding an automatically generated title to the corresponding event describing the event type in response to the confidence score satisfying the confidence threshold. In some embodiments, the method further includes, in response to the confidence score not meeting a confidence threshold, adding a title to the corresponding event based on the template representation. In some embodiments, the title machine learning model receives corresponding media from one or more events as input, and the title machine learning model generates a title as output. In some embodiments, the user interface includes an option to edit corresponding media from a particular event.In some embodiments, the method further includes receiving new media associated with the media library and replacing a particular event of the one or more events with the new event in response to the new event being associated with a new event importance score higher than the event importance score of the particular event. In some embodiments, the user interface includes an option to change the title of the event item. In some embodiments, the method further includes generating audio for the corresponding media based on a type of the particular event.

[0006] Embodiments may further include a system comprising one or more processors and a memory storing instructions executed by the one or more processors, the instructions including segmenting a media library associated with a user account into episodes, each episode associated with a corresponding time period, generating an event signal indicative of a likelihood that an event occurred in each episode using an event machine learning model, the event machine learning model being a classifier that receives the media as input, generating an event importance score for each episode, determining one or more events from the episode based on the event signal and the corresponding event importance score that exceeds an event importance threshold, and providing a user interface including an event item having corresponding media from a particular event of the one or more events.

[0007] In some embodiments, the user interface is part of a reverse-chronological grid of media including corresponding media, and the method further includes, in response to a user selecting an event item from the reverse-chronological grid, replacing the reverse-chronological grid with a display of corresponding media from a particular event for a predetermined time period, and in response to completion of the display of the corresponding media, displaying the reverse-chronological grid. In some embodiments, the event items are displayed sized based on one or more of an event importance score, a number of media items for a particular event, a total number of events for a time period, or an event type.

[0008] Embodiments may further include a non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform the following operations: segmenting a media library associated with a user account into episodes, each episode associated with a corresponding time period; generating an event signal indicating a likelihood that an event occurred in each episode using an event machine learning model, the event machine learning model being a classifier that receives media as input; generating an event importance score for each episode; determining one or more events from the episode based on the event signal and the corresponding event importance score that exceeds an event importance threshold; and providing a user interface that includes an event item having corresponding media from a particular event of the one or more events.

[0009] In some embodiments, the user interface is part of a reverse-chronological grid of media including corresponding media, and the method further includes, in response to a user selecting an event item from the reverse-chronological grid, replacing the reverse-chronological grid with a display of corresponding media from the particular event for a predetermined period of time, and in response to completion of the display of the corresponding media, displaying the reverse-chronological grid. [Effects of the Invention]

[0010] This specification advantageously describes a method for using an event machine learning model to generate an event signal indicating the likelihood that an event occurred in each episode. In this manner, an improved method for classifying media items into events can be provided. The method can, for example, provide a classification into events that more reliably reflects underlying trends in the data than predefined classifications or categories. Furthermore, the machine learning model can advantageously reduce power consumption and increase efficiency by using a static training set and updating the event machine learning model in response to the update size being less than a threshold size. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary network environment in accordance with certain embodiments described herein. [Figure 2] FIG. 1 is a block diagram illustrating an exemplary computing device in accordance with some embodiments described herein. [Figure 3] FIG. 10 illustrates example titles based on confidence scores indicating the likelihood that the corresponding event is an accurately recognized event type, according to some embodiments. [Figure 4] FIG. 1 illustrates an exemplary reverse-chronological grid of media according to some embodiments. [Figure 5] 10A-10C illustrate exemplary user interfaces for removing media items from an event according to some embodiments. [Figure 6] 10A-10C illustrate exemplary user interfaces including options for removing media items from an event item and providing feedback, according to some embodiments. [Figure 7]1 illustrates an exemplary user interface that includes options for editing titles, deleting events, resizing events in a reverse chronological grid, and changing the importance of events in a reverse chronological grid, according to some embodiments. [Figure 8A] 1A-1C are exemplary block diagrams illustrating different examples for rearranging different events based on changes, according to some embodiments. [Figure 8B] 1A-1C are exemplary block diagrams illustrating different examples for rearranging different events based on changes, according to some embodiments. [Figure 9] 1 is a flow diagram illustrating an example method for displaying event items according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0012] Detailed Description Network Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes media server 101, user device 115a, user device 115n, and network 105. Users 125a, 125n may be associated with user devices 115a, 115n, respectively. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1 or may not include media server 101. In FIG. 1 and other figures, a letter following a reference number, e.g., "115a," indicates a reference to the element with that specific reference number. A reference number in text without a letter following it, e.g., "115," indicates a general reference to an embodiment of the element with that reference number.

[0013] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively connected to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits and receives data to and from one or more user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0014] The media application 103a may include code and routines operable to receive a media library associated with a user account. Media items referred to herein may include images or videos. Images may include digital images having pixels with one or more pixel values ​​(e.g., color values, brightness values, etc.). Images may be still images (e.g., still images, single-frame images, etc.) or dynamic images (images including multiple frames, e.g., videos, animated GIFs, or cinemagraphs in which some portions of the image include movement and other portions are static). Videos referred to herein include multiple frames, with or without audio. In some implementations, one or more camera settings, e.g., zoom level, aperture, etc., may be changed while the video is being captured. In some implementations, the client device capturing the video may be moved while the video is being captured. Text referred to herein may include alphanumeric characters, emojis, symbols, or other characters.

[0015] The media application 103a may include code and routines operable to segment a media library into episodes. The media application 103a may use an event machine learning model to generate event signals indicating the likelihood that an event occurred in each episode. The media application 103a may generate an event importance score for each episode and determine one or more events from the episode based on the event signals and corresponding event importance scores that exceed an event importance threshold. The media application 103a may display event items, including corresponding media, from a library associated with the user account in a user interface.

[0016] In some embodiments, the media application 103a may be implemented using a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), The media application 103a may be implemented using hardware including an application specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.

[0017] The database 199 can store media libraries, training sets for machine learning models, user actions related to media (views, shares, annotations, etc.), and can store media items that are indexed and associated with the identities of users 125 of user devices 115. The database 199 can also store social network data related to users 125, user preferences for users 125, etc.

[0018] The user device 115 may be a computing device that includes a memory and a hardware processor. For example, the user device 115 may include a desktop computer, a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.

[0019] In the illustrated implementation, user device 115a is connected to network 105 via signal line 108, and user device 115n is connected to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a, 115n are utilized by users 125a, 125n, respectively. The user devices 115a, 115n in FIG. 1 are used for illustrative purposes. While FIG. 1 shows two user devices 115a and 115n, the present disclosure applies to system architectures including one or more user devices 115.

[0020] In some embodiments, the media application 103 receives a media library associated with a user. For example, a user captures images and videos from their own camera (e.g., a smartphone or other camera), uploads images from a digital single-lens reflex (DSLR) camera, and adds media captured and shared by other users to the media library. The media application 103 segments the media library into episodes, each associated with a corresponding time period. For example, media associated with trick or treat may be considered an event because the media relates to a period of several hours and is located in a single neighborhood or city. Many episodes may contain types of media that the user does not repeatedly engage with, such as a day at home where the user photographs various broken items to purchase replacement parts at a hardware store.

[0021] The media application 103 may include a trained machine learning model, such as an event machine learning model, that receives media as input and generates an event signal indicating the likelihood that an event occurred for each episode. For example, images that do not have the same theme and are taken in different locations over a 24-hour period may be associated with an event signal indicating that an event is unlikely to have occurred. Conversely, images taken in the same location over a 24-hour period, such as a photo of a cake, a video of a child singing "Happy Birthday," and a room full of balloons, may be associated with an event signal indicating that an event (e.g., a "birthday party") is likely to have occurred. It may be possible.

[0022] The media application 103 can generate an event importance score for each episode. The media application 103 may use a machine learning model or different types of modules for this step. The event importance score may relate to the likelihood that a user will want to engage with the media in the episode. For example, if the event importance score exceeds an event importance threshold, the media application 103 can determine that the episode corresponds to an event in which the user will want to view, share, and send the media to be printed in a photo album. In some embodiments, the event importance score is based on one or more of: at least a threshold number of media items in the corresponding episode; at least a threshold number of face clusters in the corresponding episode; a quality metric for the media items in the corresponding episode; at least one face cluster of a threshold rank; or the presence of rare face clusters.

[0023] In some embodiments, the media application 103 determines events from an episode based on both the event signal output by the event machine learning model and a corresponding event importance score that exceeds an event importance threshold. In some embodiments, the default value for events is 24 hours, but events that appear to be connected in a particular way, such as events that occur over multiple days, can also be combined.

[0024] In some embodiments, the media application 103 provides a user interface that includes corresponding media from events. The user interface may be part of a reverse-chronological grid of media that includes the corresponding media. For example, the user interface may include a "month carousel" of media in which media corresponding to events that occurred during the month of February are organized into a grid. In some embodiments, the user interface limits the number of events displayed in the grid according to a predetermined number. For example, the user interface may include only seven events in a particular month. If the user interface already displays seven events and new media is associated with a new event that has a new event importance score that is higher than the corresponding event importance scores of the other seven events, the user interface may replace the previously displayed event with the new event.

[0025] Computing Device 200 2A is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is user device 115 used to run media application 103. In another example, computing device 200 is media server 101. In yet another example, media application 103 is located partially on user device 115 and partially on media server 101.

[0026] One or more of the methods described herein can be implemented as a standalone program running on any type of computing device, a program running on a web browser, or a mobile application (app) running on a mobile computing device (e.g., a mobile phone, a smartphone, a tablet computer, a wearable device (e.g., a watch, an armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted displays), a laptop computer). In the primary example, all computations are performed by the mobile computing device. The computation is performed by the mobile application (and / or other applications 264) on the mobile computing device. However, a client / server architecture can be used. For example, the mobile computing device sends user input data to a server device and receives and outputs (e.g., displays) final output data from the server. In another example, computation may be shared between the mobile computing device and one or more server devices.

[0027] In some embodiments, computing device 200 includes processor 235, memory 237, I / O interface 239, display 241, camera 243, and storage device 245. Processor 235 may be connected to bus 218 via signal line 222. Memory 237 may be connected to bus 218 via signal line 224. I / O interface 239 may be connected to bus 218 via signal line 226. Display 241 may be connected to bus 218 via signal line 228. Camera 243 may be connected to bus 218 via signal line 230. Storage device 245 may be connected to bus 218 via signal line 232.

[0028] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling basic operations of the computing device 200. A "processor" includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., having a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a dedicated circuit for achieving a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a system with a processor optimized for performing matrix calculations (e.g., matrix multiplication), or other systems. In some implementations, the processor 235 may include one or more coprocessors for performing neural network processing. In some implementations, the processor 235 may be a processor that generates a probabilistic output by processing data. For example, the output generated by the processor 235 may be inaccurate or accurate within a range of expected output values. Processing need not be limited to a particular geographic location, nor need it be limited in time. For example, a processor may perform functions in real time, offline, or batch mode. Portions of processing may be performed by different (or the same) processing systems at different times and in different locations. A computer may be any processor in communication with a memory.

[0029] Memory 237 is typically provided within computing device 200 for use by processor 235 and may be any suitable processor-readable storage medium for storing instructions executed by the processor or set of processors, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc. Memory 237 may be located separately from and / or integrated with processor 235.

[0030] The memory 237 can store software executed on the computing device 200 by the operating system 262, the processor 235, including other applications 264, application data 266, and the media application 103. The other applications 264 may include applications such as a camera application, an image gallery or image library application, a data display engine, a web hosting engine, an image display engine, a notification engine, a social networking engine, etc. In some implementations, the media application 103 and the other applications 264 can each include instructions that enable the processor 235 to perform the functions described herein.

[0031] Application data 266 may also be data generated by other applications 264 or hardware of computing device 200. For example, application data 266 may include images captured by camera 243, user behavior identified by other applications 264 (e.g., social networking applications), etc.

[0032] The I / O interface 239 can provide functionality that allows the computing device 200 to interface with other systems and devices. The interfacing devices may be included as part of the computing device 200 or may be separate but in communication with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or database 199), and input / output devices can communicate through the I / O interface 239. In some embodiments, the I / O interface 239 can connect to interfacing devices, such as input devices (keyboards, pointing devices, touchscreens, microphones, cameras, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.). For example, when a user provides touch input, the I / O interface 239 transmits data to the media application 103.

[0033] Some exemplary interface devices that can be connected to I / O interface 239 may include display 241, which can be used to display the content described herein, e.g., images, video, and / or user interfaces of output applications, and to receive touch (or gesture) input from a user. For example, display 241 can be used to display a user interface including corresponding media from one or more events. Display 241 may include any suitable display device, e.g., a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in an eyeglass form factor or headset device, or a monitor screen of a computing device.

[0034] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.

[0035] Storage device 245 stores data associated with media application 103. For example, storage device 245 may store a media library associated with a user account, a selected media set, a training set for a machine learning model, etc. In embodiments in which media application 103 is part of media server 101, storage device 245 is the same as database 199 of FIG.

[0036] First Exemplary Media Application 103 The media application 103 shown in FIG. 2A includes a segmentation module 202 , an event machine learning module 204 , a scoring module 206 , a titling module 208 , a title machine learning module, and a user interface module 210 .

[0037] The segmentation module 202 segments a media library associated with a user account into episodes. In some embodiments, the segmentation module 202 includes a set of instructions executable by the processor 235 to segment the media library. In some embodiments, the segmentation module 202 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0038] In some embodiments, an episode containing multiple media items is defined as containing media associated with timestamps within a corresponding time period. For example, the segmentation module 202 can define an episode as occurring within a time period of 24 hours, 12 hours, 2 days, etc. In some embodiments, the time period of an episode is determined automatically. In some embodiments, a user can define the time period, for example, via a user interface.

[0039] The event machine learning module 204 generates an event signal that indicates the likelihood that an event occurred in each episode. In some embodiments, the event machine learning module 204 includes a set of instructions executable by the processor 235 to generate the event signal. In some embodiments, the event machine learning module 204 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0040] In some embodiments, the event machine learning module 204 can use training data to generate a trained model, specifically, an event machine learning model. For example, the training data may include any kind of data, such as media organized as events (e.g., images, videos, etc.), reactions to the events (users viewing the event, sharing the event, ordering prints of media in the event, commenting on the media items, etc.), and corresponding features (e.g., labels or tags associated with each media item to identify that an event occurred, the type of event, or the objects within the media item, etc.).

[0041] The training data may be obtained from any source, such as a data repository designated for training, or data that has been granted permission to be used as machine learning training data. In embodiments in which one or more users authorize the use of their user data to train a machine learning model, the training data may include this user data. In some embodiments, the user data may include images / videos or image / video metadata (e.g., images, corresponding features obtained from users providing manual tags or labels, geotags identifying the location of images, timestamps associated with event dates), information (e.g., chat data such as messages on social networks, emails, text messages, audio, video), documents (e.g., spreadsheets, text documents, presentations), etc.

[0042] In some embodiments, the training data may include synthetic data generated for training purposes, e.g., data that is not based on user input or activity in the situation being trained, e.g., data generated from simulations or computer-generated images / videos. In some embodiments, the event machine learning module 204 uses weights obtained and unedited / transferred from another application. For example, in these embodiments, the trained model may be generated, for example, on a different device and provided as part of the media application 103. In various embodiments, the trained model may be provided as a data file that includes a model structure or topology (e.g., defining the number and types of neural network nodes, the connections between the nodes, and the organization of the nodes into layers) and associated weights. The event machine learning module 204 can read the trained model data file and implement a neural network, including node connections, layers, and weights, based on the model structure or topology specified in the trained model.

[0043] The event machine learning module 204 generates a trained model, referred to herein as an event machine learning model. In some embodiments, the event machine learning module 204 is configured to apply the event machine learning model to data, such as application data 266 (e.g., input media), to identify one or more features in input media items and generate an (embedded) feature vector representing the media items.

[0044] In some embodiments, the event machine learning model is a classifier that receives media items along with information about the media items and uses the information to output an event signal indicating the likelihood that an event occurred in each episode. The information can include the results of performing optical character recognition on an image to identify text within the image that indicates a particular event. For example, a dinner menu may include the word "wedding." In some embodiments, the information can include the results of performing object recognition to identify objects associated with the event. For example, at a baby shower, there may be gifts associated with the baby.

[0045] In some implementations, metadata associated with the media item may be provided as an additional input to the gating model if the user permits it. The metadata may include user-permission factors, such as the location and / or time the video was taken, whether the video was shared via a social network, an image-sharing application, a messaging application, etc., depth information associated with one or more video frames, sensor values ​​from one or more sensors (e.g., an accelerometer, gyroscope, light sensor, or other sensors) of the camera that took the video, the user's ID (with user consent), etc. For example, if a video was taken at night in an outdoor location with the camera pointed upward, the metadata may indicate that the camera was pointed toward the sky when the video was taken and therefore is related to a celestial event.

[0046] In some embodiments, the event machine learning module 204 may use a combination of metadata, optical character recognition, and other signals as inputs to an event machine learning model. For example, the metadata may indicate that a person is using a firecracker and the date is July 4th, resulting in the event machine learning module 204 outputting an event signal corresponding to Independence Day. In some embodiments, the event machine learning module 204 outputs an event type for one or more events. Continuing with the example above, the event machine learning module 204 outputs the event signal and a likelihood that the event is Independence Day or a holiday.

[0047] In some embodiments, the event machine learning module 204 can include software code executed by the processor 235. In some embodiments, the event machine learning module 204 can specify a circuit configuration (e.g., a programmable processor, a field programmable gate array (FPGA)) that enables the processor 235 to apply the event machine learning model. In embodiments, the event machine learning module 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the event machine learning module 204 may provide an application programming interface (API). The operating system 262 and / or other applications 264 may utilize this API to call the event machine learning module 204 to determine one or more features of an input image, for example, by applying an event machine learning model to application data 266.

[0048] In some embodiments, the event machine learning model may include one or more model forms or structures. In some embodiments, the event machine learning model can use a support vector machine, but in some embodiments, a convolutional neural network (CNN) is preferred. For example, the model form or structure can include any type of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., a "hidden layer" between the input layer and the output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results obtained from the processing of each tile), or a sequence-to-sequence neural network (e.g., a network that receives sequential data as input, such as words in a sentence or frames in a video, and produces a sequence of results as output).

[0049] The model form or structure may specify the connections between various nodes and the organization of the nodes into layers. For example, nodes in an initial layer (e.g., input layer) may receive data as input data or application data 266. For example, when using an event machine learning model to analyze an input image, e.g., a first image, associated with a user account, such data may include, for example, one or more pixels per node. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connections specified in the model form or structure. These layers are sometimes referred to as hidden layers. The final layer (e.g., output layer) generates the output of the machine learning application. For example, this output may be image features associated with the input image. In some embodiments, the model form or structure specifies the number and / or type of nodes in each layer.

[0050] In some embodiments, the model is a neural network (CNN) that includes network layers, each of which extracts image features at a different level of abstraction. A CNN used to identify features within an image may be used to classify the image. The model architecture may include layer combinations and sequences of multidimensional convolutions, average pooling, max pooling, activation functions, normalization, regularization, and other layers and modules commonly used in applied deep neural networks.

[0051] In different embodiments, the event machine learning model may include one or more models. The one or more models may include multiple nodes arranged in layers according to a neural network structure or morphology, or a filter bank in the case of a CNN. In some embodiments, a node may be, for example, a memoryless computational node configured to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and generating the node output by adjusting the weighted sum with a bias value or an intercept value. Different layers may include different types of inputs related to media items, such as a first layer for metadata, a second layer for optical character recognition, and a third layer for object recognition.

[0052] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, the computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, using individual processing units of a graphical processing unit (GPU), or using dedicated neural circuitry. In some embodiments, the nodes may include memory. The nodes may, for example, store one or more previous inputs and use the one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a long-short-term memory (LSTM) node. The LSTM node can use the memory to maintain state that allows the node to operate like a finite state machine (FSM). Models including such nodes can be used to process sequential data, for example, multiple words in a sentence or paragraph, a series of images, or videos. This may be useful when processing frames in video, speech, or other audio. For example, a heuristics-based model used in a gating model may remember one or more features previously generated for previous images.

[0053] In some embodiments, the event machine learning model may include embeddings or weights for individual nodes. For example, the event machine learning model may be initialized as multiple nodes organized into layers as specified by the model topology or structure. At initialization, a respective weight may be applied to each pair of nodes connected according to the model topology, e.g., the connection between each pair of nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The event machine learning model may then be trained, e.g., using a training set of digital images, to generate results. In some embodiments, a subset of the entire architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.

[0054] For example, training may include applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., media organized as events) and expected outputs corresponding to each input (e.g., event signals for each event). For example, weight values ​​may be automatically adjusted based on a comparison of the output of the event machine learning model with the expected output to increase the probability that, when provided with inputs including correct event identifications provided as training data, the event machine learning model will generate the expected output that identifies an event from the media library by analyzing the pixels and metadata.

[0055] In some embodiments, training may include applying unsupervised learning techniques. For example, only input data (e.g., media organized as events) may be provided, and the event machine learning model may be trained to distinguish between the data, e.g., to identify media that are likely to be classified as events. In another example, the event machine learning model may train itself to generate events by distinguishing between media items within the events.

[0056] In various embodiments, the trained model includes a set of weights corresponding to the model structure. In embodiments that omit a training set of digital images, the event machine learning module 204 may generate the event machine learning model based on prior training, such as by the developer of the event machine learning module 204 or a third party. In some embodiments, the event machine learning model may include a set of fixed weights downloaded from a server that provides the weights.

[0057] In some embodiments, the event machine learning module 204 may be implemented in an offline manner. Implementing the event machine learning module 204 may include using a static training set that does not include updates when data in the static training set changes. This advantageously results in increased efficiency of processing performed by the computing device 200 and reduced power consumption of the processing device 200. In some embodiments, small updates to the event machine learning model may be implemented in an online manner, where updates to the training data are included as part of training the event machine learning model. A small update is an update having a size smaller than a threshold size. The size of the update is related to the number of variables in the machine learning model affected by the update. In such embodiments, an application that invokes the event machine learning module 204 (e.g., the operating system 262, one or more other applications 264) can utilize the feature detections generated by the event machine learning module 204 and generate a system log (e.g., if permitted by a user, actions taken by the user based on the feature detections, or results of further processing, if utilized as input for further processing). The system log may be generated periodically, e.g., hourly, monthly, or quarterly, and, if permitted by a user, may be used to update the event machine learning model, e.g., to update the embeddings of the event machine learning model.

[0058] In some embodiments, the event machine learning module 204 may be implemented in a manner that is compatible with the particular configuration of the computing device 200 on which it executes. For example, the event machine learning module 204 may determine a computation graph that utilizes available computational resources, such as the processor 235. If the event machine learning module 204 is implemented as a distributed application on multiple devices, e.g., if the media server 101 includes multiple media servers 101, the event machine learning module 204 may determine the computations to be performed on each device to optimize the computation. In another example, if the event machine learning module 204 determines that the processor 235 includes a GPU with a particular number (e.g., 1000) of GPU cores, the event machine learning module 204 may implement the event machine learning module 204 (e.g., as 1000 separate processes or threads).

[0059] In some embodiments, the event machine learning module 204 can implement a set of trained models. For example, the event machine learning model can include multiple trained models, each applicable to the same input data. In these embodiments, the event machine learning module 204 can select a particular trained model based on, for example, available computational resources, success rates using previous inferences, etc.

[0060] In some embodiments, the event machine learning module 204 can execute multiple trained models. In these embodiments, the event machine learning module 204 can combine outputs, for example, using a majority vote to score the outputs obtained by applying each trained model, or by selecting one or more specific outputs. In some embodiments, such a selector is part of the model itself and serves as a connecting layer between the trained models. Furthermore, in these embodiments, the event machine learning module 204 can apply a time threshold (e.g., 0.5 ms) for applying individual trained models and only utilize individual outputs that are available within the time threshold. Outputs that are not received within the time threshold may not be utilized, e.g., discarded. For example, such an approach may be appropriate when there is a specified time limit between invocations of the event machine learning module 204, e.g., by the operating system 262 or one or more other applications 264. In this manner, the event machine learning module 204 can combine outputs. The responsiveness of the media application 103 can be improved because the event machine learning module 204 can limit the maximum time it takes to perform a task, such as identifying media that is likely to be classified as an event, thereby allowing the event machine learning module 204 to provide the best classification in real time.

[0061] The scoring module 206 generates an event importance score for each episode. In some embodiments, the scoring module 206 includes a set of instructions executable by the processor 235 to score the episode and determine one or more events from the episode. In some embodiments, the scoring module 206 may be stored in memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0062] In some embodiments, the scoring module 206 generates an event importance score for each episode. The scoring module 206 can generate the event importance score based on data from all users who are authorized to use the user's data. This may be applied when the number of media items in the user's library is below a threshold. In some embodiments, the scoring module 206 generates the event importance score based on data specific to the user, such as if the image includes the user or a person important to the user (e.g., a person who appears in a threshold number of media items related to the user), if the episode occurred in a location important to the user (e.g., a frequently visited or familiar location) or an infrequent location, previous user responses (sharing a media item, printing a media item, showing approval of a media item (e.g., favoring a media item), editing a media item, viewing a media item), or a prediction of the user's response to the episode. For example, if a user consistently views images of a particular person, birthday parties, and Christmas, the scoring module 206 can apply a score to similar images indicating that these images are meaningful to the user. In another example, the scoring module 206 generates an event importance score based on the similarity of the episode to events deemed important.

[0063] In some embodiments, the scoring module 206 further determines the event importance score based on a threshold number of media items in the corresponding episode. For example, an episode may be required to have at least five media items (or seven, or ten, etc.) to qualify as an event. In some embodiments, the scoring module 206 further determines the event importance score based on a greater number of media items in the corresponding episode, the presence of at least one face cluster of a threshold rank in the corresponding episode, or a rare face cluster. In some embodiments, the scoring module 206 further determines the event importance score based on a quality metric of the media items in the corresponding episode. The quality metric may be based, for example, on whether the media items are blurry, oversaturated, have out-of-focus objects, etc.

[0064] In some embodiments, the scoring module 206 merges multiple episodes into a single event based on multiple episodes occurring within a 24-hour period (e.g., multiple birthday celebrations). A single event includes events occurring over multiple days (e.g., Hanukkah, Diwali, Ramadan, Cherry Blossom Festival), celebrations on different days related to a single event (e.g., Indian wedding, Christmas with tree lighting, family dinner, and gift opening), multiple episodes involving the same location (e.g., a work retreat where meals are served at the same location), or multiple episodes involving the same set of people identified by the same set of face clusters. In some embodiments, an event may be scored because it includes events with longer durations (e.g., Ramadan is 30 days, Cherry Blossom Festival is 21 days). The point module 206 merges events into those that are a maximum of 31 days.

[0065] In some embodiments, the scoring module 206 receives an event type (e.g., a birthday) from the event machine learning module 204 and merges multiple events within a predetermined time period based on the event type (e.g., merging all birthday-related celebrations that occurred within five days before and after the birthday). For example, a birthday can include not only a party but also related events, such as dinners, that occurred before and after the birthday.

[0066] In some embodiments, the scoring module 206 receives the event types from the event machine learning module 204 and generates a confidence score indicating the likelihood that the corresponding event is the correctly identified event type. For example, the scoring module 206 receives the same information regarding optical character recognition, object recognition, and metadata as described above with reference to the event machine learning module 204 and can generate a confidence score for the event type based on text within the media item, objects identified from the media item that correspond to the event type (e.g., Santa hats for Christmas, flowers and chocolate hearts for Valentine's Day), dates of particular events that correspond to known holidays (e.g., Christmas, Valentine's Day), etc. The confidence score may be a number (1, 0.4, 500), a percentage (5%, 44%, etc.), or a different metric.

[0067] In some embodiments, the scoring module 206 uses a scoring machine learning model to output an event importance score. In this example, the scoring machine learning model can receive as input user actions as part of a training set for training the scoring machine learning model. Once the scoring machine learning model is trained, the scoring machine learning model can receive as input episodes and output an event importance score.

[0068] In some embodiments, the scoring module 206 determines one or more events from the episode based on the event signals received by the event machine learning module 204 and corresponding event importance scores that exceed the event importance threshold. The scoring module 206 can instruct the user interface module 210 to display in a grid the events that have corresponding event importance scores that exceed the event importance threshold.

[0069] In some embodiments, events may be generated periodically. For example, one or more events may be limited to a predetermined number (e.g., seven) each month to avoid overloading a user with different events. For example, when another user who shares media with a first user receives new media, the first user can associate the new media with their media library. The scoring module 206 generates an event importance score for the episode that includes the new media, and if the event importance score exceeds an event importance threshold, the scoring module 206 determines that the event importance score is greater than the event importance score of a particular event displayed in the user interface, and the scoring module 206 can instruct the user interface 210 to replace the previous event with the new event.

[0070] The titling module 208 generates a title for the event. In some embodiments, the titling module 208 includes a set of instructions executable by the processor 235 to generate the title. In some embodiments, the titling module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.

[0071] In some embodiments, the titling module 208 receives an event signal from the event machine learning module 204 that indicates the likelihood that an event occurred in each episode. The titling module 208 can also receive an identification of an event type from the event machine learning module 204. The titling module 208 can use the event type to determine an event title that describes the event type.

[0072] In some embodiments, the titling module 208 receives a confidence score from the scoring module 206 indicating the likelihood that the corresponding event is the correctly identified event type. If the confidence score meets a confidence threshold, the titling module 208 adds a title to the corresponding event based on the event type. For example, FIG. 3 includes a first example image 600 of a man holding flowers and two people embracing. In this case, because the confidence score meets the threshold, the titling module 208 applies the title "Valentine's Day" to the corresponding event. This may occur, for example, if the scoring module 206 determines that the image was taken as part of a Valentine's Day outing based on the location of the image, which is a restaurant, the date the image was taken, and the man holding the flowers. In another example, the titling module 206 uses a high-confidence title if it determines that more than 50% (or 40%, 90%, etc.) of the media items in the event have a confidence of at least 90% (or 85%, etc.). In some embodiments, if the event has a high confidence value and occurs only once a year, the titling module 208 adds the year to the title (e.g., Diwali 2019).

[0073] If the confidence score does not meet the confidence threshold, the titling module 208 adds a title to the corresponding event based on a template expression that suggests the event type but is not so specific as to be interpreted as an error (e.g., "Celebrating Love and Laughter" for an event occurring shortly after the user's birthday). Continuing with the above example, FIG. 3 includes a second example 350. In this case, because the confidence score does not meet the threshold and the image was taken on Valentine's Day, but at a campsite and the woman is not holding flowers, the titling module 208 applies the template expression "Anniversary of Love." The scoring module 206 can determine the confidence score based on the sum of the media items for the corresponding event.

[0074] Title machine learning module 258 generates a title for the event. In some embodiments, title machine learning module 258 includes a set of instructions executable by processor 235 to generate a title. In some embodiments, title machine learning module 258 may be stored in memory 237 of computing device 200 and accessible and executable by processor 235.

[0075] In some embodiments, the title machine learning module 258 can use training data to generate a trained model, specifically a title machine learning model. For example, the training data may include any type of data, such as media organized as events (e.g., images, videos, etc.), event titles, and corresponding features (e.g., labels or tags associated with each media item to identify that an event occurred, the type of event, objects within the media item, etc.).

[0076] The training data may be obtained from any source, e.g., a data repository designated for training, data that has been given permission to be used as training data for machine learning. In embodiments where one or more users authorize the use of their respective user data to train the machine learning model, the training data may include these user data. In some embodiments, the user data may be images / videos or image / video metadata (e.g., images, corresponding features, tags, labels, etc. obtained from users providing manual titles, etc., and / or image / video metadata). This may include geotags identifying the location of an event, timestamps associated with the date of the event), information (e.g., chat data such as messages on social networks, emails, text messages, audio, video), documents (e.g., spreadsheets, text documents, presentations), etc.

[0077] In some embodiments, the training data may include synthetic data generated for training purposes, e.g., data that is not based on user input or activity in the situation being trained, e.g., data generated from simulations or computer-generated images / videos. In some embodiments, the title machine learning module 258 uses weights obtained and unedited / transferred from another application. For example, in these embodiments, the trained model may be generated, e.g., on a different device and provided as part of the media application 103. In various embodiments, the trained model may be provided as a data file that includes a model structure or form (e.g., defining the number and type of neural network nodes, the connections between the nodes, and the organization of the nodes into multiple layers) and associated weights. The title machine learning module 258 can read the trained model data file and implement a neural network, including node connections, layers, and weights, based on the model structure or form specified in the trained model.

[0078] Title machine learning module 258 generates a trained model, referred to herein as a title machine learning model. In some embodiments, title machine learning module 258 is configured to apply the event machine learning model to data, such as application data 267 (e.g., input media), to identify one or more features in input media items and generate an (embedded) feature vector representing the media items.

[0079] In some embodiments, the media application 103 includes either the titling module 208 or the title machine learning module 258. In some embodiments, the titling module 208 and the title machine learning module 258 are the same module. In some embodiments, the title machine learning module 258 is stored on the media server, and the remaining modules are stored on the user device 115.

[0080] In some implementations, metadata associated with a media item may be provided as additional input to the gating model if the user permits it. The metadata may include the location and / or time the video was taken, whether the video was shared via a social network, an image sharing application, a messaging application, etc., depth information associated with one or more video frames, sensor values ​​from one or more sensors (e.g., an accelerometer, gyroscope, light sensor, or other sensors) of the camera that took the video, and (with the user's consent) user authorization factors such as the user's identity. For example, if a video was taken at night in an outdoor location with the camera pointed upward, the metadata may indicate that the camera was pointed toward the sky when the video was taken and therefore related to a celestial event.

[0081] In some embodiments, title machine learning module 258 may include software code executed by processor 235. In some embodiments, title machine learning module 258 may specify circuitry (e.g., a programmable processor, a field programmable gate array (FPGA)) that enables processor 285 to apply the title machine learning model. In some embodiments, title machine learning module 258 may include software instructions, hardware instructions, or In some embodiments, title machine learning module 258 may include an application programming interface (API) that allows operating system 263 and / or other applications 265 to call title machine learning module 258 to determine one or more features of an input image, for example, by applying an event machine learning model to application data 267.

[0082] In some embodiments, the title machine learning model may include one or more model forms or structures. In some embodiments, the title machine learning model can use a support vector machine, but in some embodiments, a convolutional neural network (CNN) is preferred. For example, the model form or structure can include any type of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., a "hidden layer" between the input layer and the output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results obtained from the processing of each tile), or a sequence-to-sequence neural network (e.g., a network that receives sequential data as input, such as words in a sentence or frames in a video, and produces a sequence of results as output).

[0083] The model form or structure may specify the connections between various nodes and the organization of the nodes into layers. For example, nodes in an initial layer (e.g., input layer) may receive data as input data or application data 266. For example, when using a title machine learning model to analyze an input image, e.g., a first image, associated with a user account, such data may include, for example, one or more pixels per node. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connections specified in the model form or structure. These layers are sometimes referred to as hidden layers. The final layer (e.g., output layer) generates the output of the machine learning application. For example, this output may be a title associated with the input image. In some embodiments, the model form or structure specifies the number and / or type of nodes in each layer.

[0084] In some embodiments, the model is a CNN with network layers, each extracting image features at a different level of abstraction. A CNN used to identify features in an image may be used to classify the image. The model architecture may include layer combinations and sequences of multidimensional convolutions, average pooling, max pooling, activation functions, normalization, regularization, and other layers and modules commonly used in applied deep neural networks.

[0085] In different embodiments, the title machine learning model may include one or more models. The one or more models may include multiple nodes arranged in layers according to a neural network structure or morphology, or, in the case of a CNN, a filter bank. In some embodiments, a node may be, for example, a memoryless computational node configured to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and generating the node output by adjusting the weighted sum with a bias value or intercept value. Different layers may include different types of inputs related to media items, such as a first layer for metadata, a second layer for optical character recognition, and a third layer for object recognition.

[0086] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, the computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, using individual processing units of a graphical processing unit (GPU), or using dedicated neural circuitry. In some embodiments, the nodes may include memory. The nodes may, for example, store one or more previous inputs and use the one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a long-short-term memory (LSTM) node. The LSTM node can use the memory to maintain state that allows the node to operate like a finite state machine (FSM). Models including such nodes can be used to process sequential data, for example, multiple words in a sentence or paragraph, a series of images, or videos. This may be useful when processing frames in video, speech, or other audio. For example, a heuristics-based model used in a gating model may remember one or more features previously generated for previous images.

[0087] In some embodiments, the title machine learning model may include embeddings or weights for individual nodes. For example, the title machine learning model may be initialized as a plurality of nodes organized into layers as specified by the model topology or structure. At initialization, a respective weight may be applied to each pair of nodes connected according to the model topology, e.g., the connection between each pair of nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The title machine learning model may then be trained, e.g., using a training set of digital images, to generate results. In some embodiments, a subset of the entire architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.

[0088] For example, training may include applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., media organized as events) and expected outputs corresponding to each input (e.g., titles for each event). For example, weight values ​​are automatically adjusted based on a comparison of the output of the title machine learning model with the expected output to increase the probability that the title machine learning model generates the expected output for the appropriate tile.

[0089] In some embodiments, training may include applying unsupervised learning techniques. For example, only input data (e.g., media organized as events and titles) may be provided, and the title machine learning model may be trained to distinguish between the data, e.g., to identify events associated with different titles. In another example, the title machine learning model may train itself to generate titles by distinguishing between media items within an event.

[0090] In various embodiments, the trained model includes a set of weights corresponding to the model structure. In embodiments that omit a training set of digital images, title machine learning module 258 may generate the title machine learning model based on prior training, such as by the developer of title machine learning module 258 or a third party. In some embodiments, the title machine learning model may include a set of fixed weights downloaded from a server that provides the weights.

[0091] In some embodiments, the title machine learning module 258 performs offline In some embodiments, small updates to the title machine learning model may be implemented in an online manner. In such embodiments, an application that invokes the title machine learning module 258 (e.g., operating system 263, one or more other applications 265) can utilize the feature detections generated by the title machine learning module 258 and generate a system log (e.g., if permitted by the user, actions the user took based on the feature detections, or results of further processing, if utilized as input for further processing). The system log may be generated periodically, e.g., hourly, monthly, or quarterly, and, if permitted by the user, may be used to update the title machine learning model, e.g., to update the embeddings of the title machine learning model.

[0092] In some embodiments, title machine learning module 258 may be implemented in a manner that is compatible with the particular configuration of computing device 200 on which title machine learning module 258 executes. For example, title machine learning module 258 may determine a computation graph that utilizes available computational resources, such as processor 285. If title machine learning module 258 is implemented as a distributed application on multiple devices, title machine learning module 258 may determine the computations to be performed on each device to optimize the computation. In another example, if title machine learning module 258 determines that processor 235 includes a GPU with a particular number (e.g., 1000) of GPU cores, title machine learning module 258 may implement title machine learning module 258 (e.g., as 1000 separate processes or threads).

[0093] In some embodiments, the title machine learning module 258 can implement a set of trained models. For example, the title machine learning model can include multiple trained models, each applicable to the same input data. In these embodiments, the title machine learning module 258 can select a particular trained model based on, for example, available computational resources, success rates using previous inferences, etc.

[0094] In some embodiments, title machine learning module 258 can execute multiple trained models. In these embodiments, title machine learning module 258 can combine outputs, for example, using a majority vote to score the outputs obtained by applying each trained model, or by selecting one or more specific outputs. In some embodiments, such a selector is part of the model itself and serves as a connecting layer between trained models. Furthermore, in these embodiments, title machine learning module 258 can apply a time threshold (e.g., 0.5 ms) for applying individual trained models and utilize only individual outputs that are available within the time threshold. Outputs that are not received within the time threshold may not be utilized, e.g., discarded. For example, such an approach may be appropriate when there is a specified time limit between invocations of title machine learning module 258, for example, by operating system 263 or one or more other applications 265.

[0095] In some embodiments, the title machine learning module 258 receives title edits from users and incorporates the feedback into the title machine learning model to modify parameters and thereby improve the title output.

[0096] The user interface module 210 generates a user interface. In some embodiments, the user interface module 210 includes a set of instructions executable by the processor 235 to generate a user interface. In some embodiments, the user interface module 210 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235 .

[0097] In some embodiments, the user interface module 210 provides a user interface that includes event items with corresponding media from particular events of the one or more events determined by the scoring module 206. In the user interface, particular events may be featured more prominently than other items. For example, FIG. 4 includes a first example 400 with media items from March and an event item titled "Happy Couple."

[0098] The user interface module 210 may be part of a reverse-chronological grid of media that includes corresponding media, may be a chronological grid, or may not have an organization based on the capture time of the media items. Continuing with the example above, FIG. 4 shows a first example 400 of a reverse-chronological grid 405 that includes media items from March 10, an event item titled "Happy Couple," and media items from March 9. In some embodiments, the event with the highest score is selected by the user interface module 210 to be displayed in the reverse-chronological grid. In some embodiments, a predetermined amount of time (e.g., one week, two days, etc.) must have passed for an event to be displayed in the reverse-chronological grid.

[0099] In some embodiments, the user interface module 210 can generate a user interface that includes events of different salience levels. For example, the first example 400 in FIG. 4 is a larger grid tile in which events are displayed as large tiles interspersed in a reverse-chronological grid with title headers. In some embodiments, the user interface module 210 determines whether to display events in larger or smaller grid tiles based on the event importance score, as described below.

[0100] 4 also includes a second example 425 with smaller grid tiles, where a first media item from a particular event is featured less prominently, a second media item from a particular event is featured more prominently, and the events are interspersed in a reverse-chronological grid. In some embodiments, events in the user interface are displayed below the date header corresponding to the event.

[0101] 4 further includes a third example 450. In this case, a summary of the month's highlights and a list of images from the same month are displayed in the bottom section 455 of the user interface, with specific events being featured most prominently. All events displayed inline in the reverse chronological grid are included in the monthly highlights carousel in the bottom section 455.

[0102] In some embodiments, the user interface module 210 displays the event in the user interface after a predetermined time (e.g., 21 days) has passed. In some embodiments, where an event includes multiple days spanning two different months, the user interface module 210 displays the event in the monthly highlights carousel section based on the month in which the event ended. For example, if an event occurred between October 2 and November 1, the event is shown only in the November highlights carousel. In some embodiments, if multiple events occur in a single day, the user interface module 210 selects and displays the event for that day with the best event importance score.

[0103] In some embodiments, when a user selects an event item from the grid, the user interface module 210 replaces the grid with a display of corresponding media items from the particular event within a predetermined time period. For example, if a user selects the event item "Happy Couple" from FIG. 4, the user interface module 210 displays a video of the proposal, an image of the ring, an image of the park where the proposal took place, etc. The user interface module 210 can display each media item from the particular event at a predetermined interval (e.g., 2 seconds, 5 seconds, etc.). For example, the user interface module 210 can use the entire user interface to display each media item for 2 seconds.

[0104] In some embodiments, the user interface includes elements that allow the user to navigate back and forth through the media items at their own pace, as well as elements that allow actions such as sharing and ordering prints, marking as favorites, etc. Once the viewing of the media items is complete (or the user has left the viewing of the corresponding media items), the user interface module 210 can redisplay the grid.

[0105] In some embodiments, corresponding media from one or more events is displayed in a size that depends on one or more of the following: the event importance score, the number of media items for one or more events, the total number of events in a time period, or the event type (e.g., my wedding vs. a friend's wedding). For example, in example 400 of Figure 4, the event item "Happy Couple" is displayed in a larger size based on its top-ranking event importance score.

[0106] In some embodiments, the user interface module 210 receives a request from a user to remove a depiction of a person from media in a media library that includes the depiction. For example, the user interface may allow the user to remove an image from the media library by clicking the image, the user interface may confirm that the image was removed, and the user interface may provide a page for requesting feedback regarding why the user requested the depiction of the person be removed. The feedback may be used by the user interface module 210 as input to an event machine learning model and / or to modify the event importance score. In some embodiments, the scoring module 206 filters depictions of the person from the media before generating the event importance score. This may advantageously allow the user interface module 210 to avoid displaying media related to events with too few media items after removing the depiction of the person from the media, or to avoid displaying events related to or featuring the hidden depiction of the person.

[0107] The user interface may include options for curating the media grid and library. In some embodiments, the user interface includes options for hiding or adding one or more of media items from one or more events, dates associated with one or more events, or people or pets depicted in media items from one or more events. Figure 5 includes an exemplary user interface 500 that provides a user with options for hiding people and pets or hiding dates from media associated with an event. The exemplary user interface 500 also includes options for selecting memories to include in the reverse chronological grid and options for managing whether to provide notifications related to the memories, such as daily reminders, silent notifications, etc. Includes options.

[0108] In some embodiments, the user interface includes an option for deleting an event item. By deleting an event item, the event item will not be displayed to the user in the future, but the corresponding media items will remain in the media library associated with the user account. In some embodiments, the user interface includes an option for deleting media items from the event item. Deleting a media item from an event item does not delete the media item from the media library associated with the user account. FIG. 6 shows an exemplary user interface. In this case, the user interface includes a request for confirmation 600, a confirmation 625 that the media item has been deleted, and a feedback screen 650 that requests feedback regarding why the user wants to delete the media item from the event item. Exemplary options for feedback include the event item being sensitive, a duplicate, off-topic, low quality, or other.

[0109] In some embodiments, the user interface module 210 sends the feedback to the event machine learning module 204, which uses the feedback to modify parameters of the event machine learning model. In some embodiments, the user interface module 210 sends the feedback to the scoring module 206, which modifies the event importance score in response to the feedback.

[0110] In some embodiments, the user interface includes options for modifying details of the event. For example, the user interface includes an option for editing corresponding media from one or more events. In yet another example, the user interface includes an option for modifying the title of one or more events while the one or more events are displayed in the user interface. In some embodiments, the titling module 208 uses the feedback to improve the generation of the title.

[0111] 7 shows an exemplary user interface 700 with options for editing the title, deleting the event, resizing the event in the grid, and changing the importance of the event in the grid, according to some embodiments. In this example, a user can access these options by right-clicking on an event in the user interface 700 or by some other mechanism. In some embodiments, these edits are available when the event is displayed in the user interface 700 or while the media items within the event are displayed.

[0112] Selecting "Edit Title" changes the title of user interface 700 to something different. Selecting "Delete Memory" leaves the media items associated with the event in the library associated with the user, but the event will not be displayed in the user interface in the future. Selecting "Normal Size" displays the event in user interface 700 at a normal size, like the second example 425 of FIG. 4, instead of the larger size of the first example 400 of FIG. 4 or the smaller size of the third example 450. Selecting "Spotlight" displays the event in user interface 700 at a larger size, like the first example 400 of FIG. 4.

[0113] In some embodiments, the user interface module 210 generates the audio for the corresponding media based on one or more event types. For example, for happy events such as graduations or weddings, the music may be happy music. For more serious events such as a funeral, the music may be more somber.

[0114] In some embodiments, in response to a user editing a feature of an event, such as the title of the event, the user interface module 210 retains the event in the user interface even if the scoring module 206 scores a different event higher than the edited event. This may be referred to as a "frozen" event. In some embodiments, the user interface module 210 retains the event in the user interface until a predetermined period of time (e.g., October) has expired. Conversely, in some embodiments, if a user deletes a media item from the library, the user interface module 210 removes the media item from the event.

[0115] In some embodiments, the user interface module 210 rearranges different events in the user interface based on different modifications. Referring to FIG. 8A, an example block diagram 800 shows different examples in which a grid in a user interface rearranges different events based on a modification that limits the number of events to seven. A first example 805 includes a list of events displayed in the grid whose corresponding event importance scores range from 70 / 100 to 99 / 100.

[0116] When a user uploads a media item from a DSLR camera, the scoring module 206 generates a new event 810 with an event importance score of 81 / 100, which is higher than five of the events in the grid. As a result, the event with the lowest event score (i.e., 70 / 100) is removed from the grid and the new event is added. The user interface module 210 adds the new event 810 to the event list, resulting in a second example 815 in which the new event 810 is added and the event with the lowest event importance score of 70 / 100 is removed from the user interface. In the second example 815, the removed event is indicated by an underline.

[0117] Continuing with the example of FIG. 8A , a user edits the title of an event that has a score of 75 / 100, meaning the event is frozen and kept in the grid. The user then uploads a media item from their DSLR camera, and the scoring module 206 generates a new event 825 with a score of 77 / 100. Because the new event 825 has a higher score of 77 / 100 than the score of 75 / 100, the second event would normally be removed from the grid. However, because the second event is frozen, it remains in the grid. As a result, FIG. 8B shows a block diagram 850 with a third example 835 in which the grid contains eight events.

[0118] 8B , when a user uploads a media item from a DSLR camera, the scoring module 206 generates a new event with a score of 88 / 100. Because the new event 840 has a higher score than the previous event's score of 77 / 100 and the event with a score of 75 / 100 has been frozen, the previous event with a score of 77 / 100 is removed from the grid. A fourth example 845 shows the remaining events in the grid, with the events with event importance scores of 77 / 100 and 70 / 100 removed from the grid (as indicated by the underlining).

[0119] In some embodiments, the user interface module 210 generates a user interface for presenting the remembered event to the user. For example, the user interface module 210 can generate an icon that is displayed at the top of the user's screen. When the icon is selected, the user interface displays a reminder of the event. In some embodiments, the user interface module 210 selects events that occurred a predetermined time (e.g., one year ago), events with the highest event importance score, or random events, etc.

[0120] Exemplary Method 900 9 is a flow diagram illustrating an example method 900 for displaying event items, according to some embodiments. Method 900 may be performed by computing device 200 of FIG. 2A or 2B, such as user device 115 or media server 101 shown in FIG. 1.

[0121] Method 900 may begin at block 902. In block 902, a media library associated with a user account is segmented into episodes. Each episode is associated with a corresponding time period. Block 902 may be followed by block 904.

[0122] In block 904, an event machine learning model generates an event signal indicating the likelihood that an event occurred in each episode. The event machine learning model is a classifier that receives media as input. Block 906 can be executed after block 904.

[0123] In block 906, an event importance score is generated for each episode. In some embodiments, the event importance score is generated based on one or more of at least a threshold number of media items in the corresponding episode, at least a threshold number of face clusters in the corresponding episode, a quality indicator of the media items in the corresponding episode, at least one face cluster of a threshold rank, or the presence of rare face clusters. Block 908 can be performed after block 906.

[0124] At block 908, one or more events are determined from the episode based on the event signal and the corresponding event importance score that exceeds the event importance threshold. Block 908 may be followed by block 910.

[0125] At block 910, a user interface is provided that includes an event item with corresponding media from a particular event of the one or more events. In some embodiments, the user interface is part of a chronological grid of media that includes the corresponding media. When a user selects an event item from the chronological grid, the chronological grid may be replaced with a display of corresponding media items from the particular event for a predetermined period of time. In response to completion of the display of the corresponding media items, the chronological grid is displayed.

[0126] In addition to the above, the systems, programs, or features described herein may provide users with controls to choose whether and when to enable collection of user information (e.g., information about the user's media items, such as photos or videos; the user's social networks, social behavior or activities; occupation; user preferences, such as viewing preferences for image-based creations; settings for hiding people or pets; user interface preferences; or information about the user's current location) and to transmit content or information from the server. Furthermore, certain data may be processed to remove identifiable personal information in one or more ways before being stored or used. For example, the user's ID may be processed so that the user's personal information cannot be identified. Also, location information (e.g., city, zip code, or state level) may be removed so that the user's location cannot be identified. When obtaining information, the user's geographic location can be generalized. Thus, the user has control over what information is collected about them, how the information is used, and what information is provided to them.

[0127] In the above description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments may be practiced without these specific details. In some instances, structures and devices are shown in block diagrams to avoid obscuring the description. For example, the embodiments may be described above primarily with reference to user interfaces and specific hardware. However, the embodiments may apply to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0128] References herein to "some embodiments" or "some instances" mean that a particular feature, structure, or characteristic described in connection with the embodiments or instances may be included in at least one implementation of the description. Appearances of the phrase "in some embodiments" in various places in the specification do not necessarily all refer to the same embodiments.

[0129] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is herein, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It is convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0130] It should be understood that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise stated, or as will be apparent from the discussion, throughout the description, discussions utilizing terms including "processing," "operating," "calculating," "determining," or "displaying," etc., refer to operations and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical quantities in the computer system memory, registers, or other information storage, transmission, or display devices.

[0131] Embodiments herein also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on any type of non-transitory computer-readable storage medium, including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each connected to a computer system bus.

[0132] This specification may include some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements. In some embodiments, this specification may be implemented in software, including but not limited to firmware, resident software, microcode, etc. It is equipped.

[0133] Furthermore, the descriptions may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, machine, or device.

[0134] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, mass storage devices, and cache memory that provides temporary storage of at least some program code to reduce the number of times the code must be retrieved from mass storage devices during execution.

Claims

1. 1. A computer-implemented method comprising: segmenting a media library associated with a user account into episodes, each episode associated with a corresponding time period, the method further comprising: generating an event signal indicative of the likelihood that an event occurred in each episode using an event machine learning model, the event machine learning model being a classifier that receives the media as input, the method further comprising: generating an event importance score for each episode; determining one or more events from the episode based on the event signal and a corresponding event importance score that exceeds an event importance threshold; providing a user interface including an event item having corresponding media from a particular event of the one or more events.

2. the user interface is part of a reverse-chronological grid of media that includes the corresponding media; The method comprises: responsive to a user selecting the event item from the reverse-chronological grid, replacing the reverse-chronological grid with a display of the corresponding media from the particular event for a predetermined period of time; 2. The method of claim 1, further comprising: displaying the reverse-chronological grid in response to completion of the display of the corresponding media.

3. generating the event machine learning model offline using a static training set; responsive to the update size being less than a threshold size, updating the event machine learning model.

4. receiving a request from a user to remove a depiction of a person from the media in the media library that includes the depiction of the person; and filtering the depiction of the person from the media, the filtering being performed prior to generating the event significance score.

5. 10. The method of claim 1, wherein the user interface includes options to hide or add one or more of media items from the one or more events, dates associated with the one or more events, or people or pets depicted in media items from the one or more events.

6. determining the computations to be performed on each individual device to optimize the computations; 10. The method of claim 1, further comprising: implementing the event machine learning model on a plurality of devices based on calculations performed on the individual devices.

7. 2. The method of claim 1 , wherein generating the event importance score is based on one or more of at least a threshold number of media items in a corresponding episode, at least a threshold number of face clusters in the corresponding episode, a quality indicator of the media items in the corresponding episode, at least one face cluster of a threshold rank, or the presence of rare face clusters.

8. based on one or more of a plurality of episodes respectively associated with a plurality of time periods, further comprising merging several episodes into a single event, the plurality of periods being generally comprised within a 24-hour period; 10. The method of claim 1, wherein the single event is a type of event occurring over multiple days, celebrations on different days related to the single event, multiple episodes related to the same location, or multiple episodes related to the same set of face clusters.

9. generating a confidence score indicating the likelihood that the corresponding event is the correctly recognized event type; 10. The method of claim 1, further comprising: in response to the confidence score meeting a confidence threshold, adding an automatically generated title to the corresponding event that describes a type of the event.

10. The method of claim 9 , further comprising, in response to the confidence score not satisfying the confidence threshold, adding the title to the corresponding event based on a template representation.

11. a title machine learning model that receives as input the corresponding media from the one or more events; The method of claim 1 , wherein the title machine learning model generates a title as an output.

12. The method of claim 1 , wherein the user interface includes an option for editing the corresponding media from the particular event.

13. determining the one or more events includes determining the events such that a number of the events is less than or equal to a predetermined number per month; The method comprises: receiving new media associated with the media library; 2. The method of claim 1, further comprising: replacing the particular event of the one or more events with the new event in response to the new event being associated with a new event importance score that is higher than the event importance score of the particular event.

14. The method of claim 1 , wherein the user interface includes an option to change the title of the event item.

15. The method of claim 1 , further comprising generating audio for the corresponding media based on the particular event type.

16. 1. A computing device comprising: a processor; a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform the following operations: The operation is segmenting a media library associated with a user account into episodes, each episode associated with a corresponding time period; generating an event signal indicative of a likelihood that an event occurred in each episode using an event machine learning model, the event machine learning model being a classifier that receives the media as input, the operations further comprising: generating an event importance score for each episode; determining one or more events from the episode based on the event signal and a corresponding event importance score that exceeds an event importance threshold; and providing a user interface including an event item having corresponding media from a particular event of the one or more events.

17. the user interface is part of a reverse-chronological grid of media that includes the corresponding media; The operation is responsive to a user selecting the event item from the reverse-chronological grid, replacing the reverse-chronological grid with a display of the corresponding media from the particular event for a predetermined period of time; 17. The computing device of claim 16, further comprising: displaying the reverse-chronological grid in response to completion of the display of the corresponding media.

18. 17. The computing device of claim 16, wherein the event items are displayed in a size based on one or more of the event importance score, the number of media items for the particular event, the total number of events in a time period, or the event type.

19. 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform the following actions: The operation is segmenting a media library associated with a user account into episodes, each episode associated with a corresponding time period; generating an event signal indicative of a likelihood that an event occurred in each episode using an event machine learning model, the event machine learning model being a classifier that receives the media as input, the operations further comprising: generating an event importance score for each episode; determining one or more events from the episode based on the event signal and a corresponding event importance score that exceeds an event importance threshold; and providing a user interface including an event item having corresponding media from a particular event of the one or more events.

20. the user interface is part of a reverse-chronological grid of media that includes the corresponding media; The operation is responsive to a user selecting the event item from the reverse-chronological grid, replacing the reverse-chronological grid with a display of the corresponding media from the particular event for a predetermined period of time; 20. The computer-readable medium of claim 19, further comprising: displaying the reverse-chronological grid in response to completion of the display of the corresponding media.