Automatically generating events using machine learning models

By using an event machine learning model to segment the media library into episodes and generate event signals, the problem of organizing events in the user's media library is solved, achieving efficient and accurate event recognition and management.

CN116235167BActive Publication Date: 2026-07-31GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2021-12-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Users' media libraries are difficult to organize images and videos effectively, and existing technologies cannot efficiently identify and classify events, making it difficult to find and recall media items for specific events.

Method used

The media library is segmented into rounds using an event-based machine learning model, and event signals are generated. Events are identified by an event importance score, and a user interface is provided to display and manage event items.

Benefits of technology

It improves the efficiency of media item classification, reduces power consumption, and enhances the accuracy of event recognition and user experience, while simplifying the process of finding event media.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116235167B_ABST
    Figure CN116235167B_ABST
Patent Text Reader

Abstract

A media application divides a media library associated with a user's account into rounds, where each round is associated with a corresponding time period. The media application uses an event machine learning model to generate event signals indicating the probability of an event occurring in each round; the event machine learning model is a classifier that receives media as input. The media application generates an event importance score for each round. Based on the event signals and the corresponding event importance scores exceeding a threshold event importance value, the media application determines one or more events from the round. The media application provides a user interface that includes the corresponding media from the one or more events.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 404773, filed August 17, 2021, entitled "Automatic Generation of Events Using a Machine-Learning Model," which claims priority to U.S. Provisional Patent Application No. 63 / 187392, filed May 11, 2021, entitled "Automatic Generation of Events Using a Machine-Learning Model," and U.S. Provisional Patent Application No. 63 / 189657, filed May 17, 2021, entitled "Automatic Generation of Events Using a Machine-Learning Model," each of which is incorporated herein by reference in its entirety. Background Technology

[0003] Users of devices such as smartphones or other digital cameras capture and store large amounts of media (such as photos and videos) in their libraries. Users can access the library to view their media to recall various events such as birthdays, weddings, vacations, trips, etc. However, libraries often contain thousands of images taken over a long period of time that are difficult to organize.

[0004] The background description provided herein is for the purpose of generally presenting the context of this disclosure. Within the scope of the description in this background section, neither the work of the inventors currently identified nor aspects that may not at the time of filing conform to the description as prior art are expressly or impliedly acknowledged as prior art relative to this disclosure. Summary of the Invention

[0005] A computer-implemented method includes: segmenting a media library associated with a user account into episodes, wherein each episode is associated with a corresponding time period; generating an event signal using an event machine learning model that indicates the likelihood of an event occurring in each episode, wherein the event machine learning model is a classifier that receives media as input; generating an event importance score for each episode; determining one or more events from the episode based on the event signal and the corresponding event importance score exceeding a threshold event importance value; and providing a user interface that includes event items having corresponding media from specific events of the one or more episodes.

[0006] In some embodiments, the user interface is part of an inverse time grid of media including corresponding media, and the method further includes: replacing the inverse time grid with the display of the corresponding media from the specific event for a predetermined duration in response to a user selecting an event item from the inverse time grid; and displaying the inverse time grid in response to completion of the display of the corresponding media. In some embodiments, the method further includes: generating an event machine learning model offline using a static training set; and updating the event machine learning model in response to an update size less than a threshold size. In some embodiments, the method further includes: receiving a request from a user to remove human depictions from media in a media library including human depictions; and filtering human depictions from the media, wherein the filtering is performed before generating an event importance score. In some embodiments, the method further includes: determining computations to be performed on individual devices to optimize the computations; and implementing the event machine learning model on multiple devices based on the computations to be performed on individual devices. In some embodiments, the user interface includes an option to hide event items. In some embodiments, generating an event importance score is based on one or more of the following: at least a threshold number of media items in a corresponding round, at least a threshold number of facial clusters in a corresponding round, a quality indicator of media items in a corresponding round, at least one facial cluster of a threshold level, or the presence of a rare facial cluster. In some embodiments, the method further includes: combining multiple rounds into a single event based on one or more of the following: the multiple rounds are associated with corresponding time periods all falling within a 24-hour cycle; the single event is an event type occurring over multiple days; celebrations on different dates associated with the single event; the multiple rounds involve the same location; or the multiple rounds involve the same facial cluster. In some embodiments, the method further includes: generating a confidence score indicating the probability that a corresponding event is an accurately identified event type; and adding an automatically generated title describing the event type to the corresponding event in response to the confidence score meeting a threshold confidence value. In some embodiments, the method further includes: adding a title to the corresponding event based on a template phrase in response to the confidence score failing to meet the threshold confidence value. In some embodiments, a title machine learning model receives corresponding media as input from one or more events, and the title machine learning model generates a title as output. In some embodiments, the user interface includes options for editing corresponding media from a specific event. In some embodiments, the method further includes: determining one or more events, including determining events such that the number of events per month is less than or equal to a predetermined number; receiving new media to associate with a media library; and replacing a specific event among the one or more events with the new event in response to the new event being associated with a new event importance score higher than that of the specific event. In some embodiments, the user interface includes an option to change the title of the event item. In some embodiments, the method further includes: generating audio for the corresponding media based on the type of the specific event.

[0007] The embodiment may also include a system comprising one or more processors and a memory storing instructions executed by the one or more processors, the instructions including: dividing a media library associated with a user account into rounds, wherein each round is associated with a corresponding time period; generating an event signal using an event machine learning model that indicates the probability of an event occurring in each round, wherein the event machine learning model is a classifier that receives media as input; generating an event importance score for each round; determining one or more events from the round based on the event signal and corresponding event importance scores that exceed a threshold event importance value; and providing a user interface including event items having corresponding media from a specific event from the one or more events.

[0008] In some embodiments, the user interface is part of a reverse time grid of media including corresponding media, and the method further includes: in response to a user selecting an event item from the reverse time grid, replacing the reverse time grid with a display of the corresponding media from the specific event for a predetermined duration; and in response to completion of the display of the corresponding media, displaying the reverse time grid. In some embodiments, event items are displayed based on one or more of the following: an event importance score, the number of media items for a specific event, the total number of events over a period of time, or the size of event types.

[0009] The embodiments may also include a non-transitory computer-readable medium including instructions stored thereon that, when executed by one or more computers, cause one or more desktop computers to perform operations including: dividing a media library associated with a user account into rounds, wherein each round is associated with a corresponding time period; generating event signals using an event machine learning model that indicates the probability of an event occurring in each round, wherein the event machine learning model is a classifier that receives media as input; generating an event importance score for each round; determining one or more events from the round based on the event signals and corresponding event importance scores that exceed a threshold event importance value; and providing a user interface including event items having corresponding media from specific events from the one or more events.

[0010] In some embodiments, the user interface is part of a reverse time grid of media including corresponding media, and the method further includes: in response to a user selecting an event item from the reverse time grid, replacing the reverse time grid with a display of the corresponding media from the specific event for a predetermined duration; and in response to completion of the display of the corresponding media, displaying the reverse time grid.

[0011] This specification advantageously describes a method for generating event signals using an event machine learning model, which indicates the probability of an event occurring in each round. In this way, an improved method for classifying media items as events can be provided, for example, to provide a classification of events that more reliably reflects underlying trends in the data than predefined classifications or categories. Furthermore, by using a static training set and updating the event machine learning model in response to updates smaller than a threshold size, the machine learning model can advantageously reduce power consumption and improve efficiency. Attached Figure Description

[0012] Figure 1 This is a block diagram of an example network environment based on some embodiments described herein.

[0013] Figure 2 This is a block diagram of an example computing device according to some embodiments described herein.

[0014] Figure 3 Example headers based on confidence scores, according to some embodiments, are shown. These confidence scores indicate the likelihood that the corresponding event is an event type that has been accurately identified.

[0015] Figure 4 An example inverse time grid of media according to some embodiments is shown.

[0016] Figure 5 An example user interface for removing media items from an event is shown, according to some embodiments.

[0017] Figure 6 An example user interface according to some embodiments is shown, which has the option to remove media items from event items and provide feedback.

[0018] Figure 7 An example user interface according to some embodiments is shown, which has options for editing the title, removing events, changing the size of events in the reverse time grid, and changing the importance of events in the reverse time grid.

[0019] Figures 8A-8B Example block diagrams are shown, illustrating different examples of reorganizing different events based on changes, according to some embodiments.

[0020] Figure 9 This is a flowchart illustrating an example method for displaying event items according to some embodiments. Detailed Implementation

[0021] Network environment 100

[0022] Figure 1A block diagram of an example environment 100 is shown. In some embodiments, environment 100 includes a media server 101, user equipment 115a, user equipment 115n, and network 105. Users 125a and 125n may be associated with corresponding user equipment 115a and 115n. In some embodiments, environment 100 may include... Figure 1 Other servers or devices not shown, or media server 101 may not be included. Figure 1 In the other figures, letters following the reference numerals, such as "115a", indicate a reference to an element having that particular reference numeral. Reference numerals in the text without the following letter, such as "115", indicate a general reference to an embodiment of the element having that reference numeral.

[0023] Media server 101 may include a processor, memory, and network communication hardware. In some embodiments, media server 101 is a hardware server. Media server 101 is communicatively coupled to network 105 via signal line 102. Signal line 102 may be a wired connection or a wireless connection, such as Ethernet, coaxial cable, fiber optic cable, etc., and the wireless connection may be such as Wi-Fi®, Bluetooth®, or other wireless technologies. In some embodiments, media server 101 sends and receives data to and from one or more user devices 115a, 115n via network 105. Media server 101 may include media application 103a and database 199.

[0024] Media application 103a may include code and routines operable to receive a media library associated with a user account. Media items as referred to herein may include images or videos. Images may include digital images having pixels with one or more pixel values ​​(e.g., color values, brightness values, etc.). Images may be still images (e.g., still photographs, images with a single frame, etc.) or moving images (e.g., images containing multiple frames, such as animations, animated GIFs, movie images (where a portion of the image includes motion while other portions are static), etc.). Videos as referred to herein include multiple frames with or without audio. In some implementations, one or more camera settings may be modified during video capture, such as zoom level, aperture, etc. In some implementations, the client device capturing the video may be moved during video capture. Text as referred to herein may include alphanumeric characters, emojis, symbols, or other characters.

[0025] Media application 103a may include code and routines operable to divide a media library into rounds. Media application 103a may use an event-based machine learning model to generate event signals indicating the likelihood of an event occurring in each round. Media application 103a may generate an event importance score for each round and identify one or more events from these rounds based on the event signals and corresponding event importance scores exceeding a threshold event importance value. Media application 103a may enable a user interface to display event items with corresponding media from a library associated with a user account.

[0026] In some embodiments, media application 103a may be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, media application 103a may be implemented using a combination of hardware and software.

[0027] Database 199 can store media libraries, training sets for machine learning models, and user actions associated with media (viewing, sharing, commenting, etc.). Database 199 can store media items associated with the identity index of user 125 on user device 115. Database 199 can also store social network data associated with user 125, user preferences of user 125, etc.

[0028] User equipment 115 may be a computing device including memory and a hardware processor. For example, user equipment 115 may include a desktop computer, mobile device, tablet computer, mobile phone, wearable device, head-mounted display, mobile email device, portable game console, portable music player, reader device, or another electronic device capable of accessing network 105.

[0029] In the illustrated implementation, user equipment 115a is coupled to network 105 via signal line 108, while user equipment 115n is coupled to network 104 via signal line 110. Media application 103 can be stored as media application 103b on user equipment 115a or media application 103c on user equipment 115n. Signal lines 108 and 110 can be wired or wireless connections, such as Ethernet, coaxial cable, fiber optic cable, etc., and wireless connections such as Wi-Fi®, Bluetooth®, or other wireless technologies. User equipment 115a and 115n are accessed by users 125a and 125n, respectively. Figure 1 User equipment 115a and 115n are used as examples. Although Figure 1 Two user equipments 115a and 115n are shown, but this disclosure applies to system architectures having one or more user equipments 115.

[0030] In some embodiments, media application 103 receives a media library associated with a user. For example, a user captures images and videos from their camera (e.g., a smartphone or other camera), uploads images from a digital single-lens reflex (DSLR) camera, adds media captured by another user to their media item library, and so on. Media application 103 divides the media library into rounds, where each round is associated with a corresponding time period. For example, media associated with trick-or-treating might be considered an event because the media is associated with a time period of several hours and the location is a single neighborhood or town. Many rounds may include media types that the user would not repeatedly encounter. For example, a day at home where the user takes pictures of various broken things in order to buy replacement parts at a hardware store.

[0031] Media application 103 may include trained machine learning models, such as event machine learning models that receive media as input and generate event signals indicating the probability of an event occurring in each round. For example, images taken at different locations during a 24-hour period without similar subjects may be associated with event signals indicating a low probability of an event occurring. Conversely, images taken at the same location during a 24-hour period, such as a picture of a cake, a video of children singing "happy birthday," and a room full of balloons, may be associated with event signals indicating a high probability of an event (e.g., a "birthday party").

[0032] Media application 103 can generate an event importance score for each round. Media application 103 may use a machine learning model or different types of modules for this step. The event importance score may be associated with the likelihood that the user will want to engage with media in that round. For example, if the event importance score exceeds a threshold event importance value, media application 103 may determine that the round corresponds to an event that the user may want to view, share, send to media for printing in an album, etc. In some embodiments, the event importance score is based on one or more of the following: at least a threshold number of media items in the corresponding round, at least a threshold number of face clusters in the corresponding round, a quality indicator of media items in the corresponding round, at least a threshold level of face clusters, or the presence of rare face clusters.

[0033] In some embodiments, media application 103 determines events from a round based on event signals and corresponding event importance scores output by an event machine learning model that exceed a threshold event importance value. In some embodiments, the default value for an event is 24 hours, but events can also be combined if they appear to be connected in some way, such as events occurring over multiple days.

[0034] In some embodiments, media application 103 provides a user interface that includes corresponding media from events. The user interface may be part of a reverse-time grid of media including the corresponding media. For example, the user interface may include a "monthly wheel" of media corresponding to events occurring during February, where the media are organized in a grid. In some embodiments, the user interface limits the number of events displayed in the grid according to a predetermined number. For example, the user interface may include only seven events in a particular month. If seven events are already displayed in the user interface, and new media is associated with a new event that has a higher new event importance score than the corresponding events of the other seven events, the user interface may replace the previously displayed events with the new event.

[0035] Computing equipment 200

[0036] Figure 2 This is a block diagram of an example computing device 200 that can be used to implement one or more features described herein. The computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a user device 115 for implementing media application 103. In another example, the computing device 200 is a media server 101. In yet another example, media application 103 is partially on user device 115 and partially on media server 101.

[0037] One or more methods described herein can run on a standalone program that can execute on any type of computing device, a program that runs on a web browser, or a mobile application (“app”) that runs on a mobile computing device (e.g., a mobile phone, smartphone, smart display, tablet, wearable device (watch, armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted display, etc.), laptop, etc.). In the primary example, all computation is performed within the mobile application (and / or other application 264) on the mobile computing device. However, a client / server architecture can also be used, whereby the mobile computing device sends user input data to a server device and receives final output data from the server for output (e.g., for display). In another example, computation can be split between the mobile computing device and one or more server devices.

[0038] In some embodiments, the computing device 200 includes a processor 235, a memory 237, an I / O interface 239, a display 241, a camera 243, and a storage device 245. The processor 235 can be coupled to the bus 218 via signal line 222, the memory 237 can be coupled to the bus 218 via signal line 224, the I / O interface 239 can be coupled to the bus 212 via signal line 226, the display 241 can be coupled to the bus 218 via signal line 228, the camera 243 can be coupled to the bus 218 via signal line 230, and the storage device 245 can be coupled to the bus 218 via signal line 232.

[0039] Processor 235 may be one or more processors and / or processing circuitry to execute program code and control the basic operations of computing device 200. "Processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. Processors may include: systems with a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration); multiple processing units (e.g., in a multi-processor configuration); graphics processing unit (GPU); field-programmable gate array (FPGA); application-specific integrated circuit (ASIC); complex programmable logic device (CPLD); dedicated circuitry for implementing functionality; dedicated processors for implementing processing based on neural network models; neural circuitry; processors optimized for matrix computations (e.g., matrix multiplication); or other systems. In some embodiments, processor 235 may include one or more coprocessors implementing neural network processing. In some embodiments, processor 235 may be a processor that processes data to produce probabilistic outputs; for example, the output produced by processor 235 may be imprecise or may be accurate within a range of the expected output. Processing is not limited to a specific geographical location or has time constraints. For example, a processor may perform its functions in real-time, offline, in batch mode, etc. This refers to parts of the process that can be performed by different (or the same) processing systems at different times and locations. A computer can be any processor that communicates with memory.

[0040] The memory 237 is typically located in the computing device 200 for access by the processor 235 and can be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., which is suitable for storing instructions executed by the processor or processor set and is located separately from and / or integrated with the processor 235.

[0041] Memory 237 may store software operated on computing device 200 by processor 235, including operating system 262, other applications 264, application data 266, and media application 103. Other applications 264 may include applications such as camera applications, image gallery or image library applications, data display engines, web hosting engines, image display engines, notification engines, social networking engines, etc. In some implementations, each of media application 103 and other applications 264 may include instructions that enable processor 235 to perform the functions described herein.

[0042] Application data 266 may be data generated by other applications 264 or hardware of computing device 200. For example, application data 266 may include images captured by camera 243, user actions recognized by other applications 264 (e.g., social networking applications), etc.

[0043] I / O interface 239 provides functionality that enables computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be independent and communicate with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or database 199), and input / output devices may communicate via I / O interface 239. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, camera, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.). For example, when a user provides touch input, I / O interface 239 transmits data to media application 103.

[0044] Some examples of docking devices that can be connected to I / O interface 239 may include display 241, which can be used to display content, such as images, videos, and / or user interfaces of output applications as described herein, and to receive touch (or gesture) input from a user. For example, display 241 may be used to display a user interface including corresponding media from one or more events. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), a light-emitting diode (LED) or plasma display, a cathode ray tube (CRT), a television, a monitor, a touch screen, a 3D display, or other visual display device. For example, display 241 may be a flat panel display provided on a mobile device, multiple displays embedded in a shape-factor eyeglass or headphone device, or a monitor screen for a computer device.

[0045] Camera 243 can be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video transmitted to media application 103 via I / O interface 239.

[0046] Storage device 245 stores data related to media application 103. For example, storage device 245 may store media libraries associated with user accounts, selected media sets, training sets for machine learning models, etc. In embodiments where media application 103 is part of media server 101, storage device 240 and... Figure 1 The database 199 is the same.

[0047] First example media application 103

[0048] Figure 2 The media application 103 shown includes a segmentation module 202, an event machine learning module 204, a rating module 206, a title module 208, a title machine learning module, and a user interface module 210.

[0049] The segmentation module 202 segments the media library associated with a user account into rounds. In some embodiments, the segmentation module 202 includes a set of instructions executable by the processor 235 to segment the media library. In some embodiments, the segmentation module 202 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0050] In some embodiments, a round comprising a set of media items is defined as including media associated with timestamps within a corresponding time period. For example, the segmentation module 202 may define a round as occurring within a 24-hour time period, 12 hours, two days, etc. In some embodiments, the time period of a round is determined automatically. In some embodiments, a user may define the time period, for example, via a user interface.

[0051] Event machine learning module 204 generates event signals that indicate the probability of an event occurring in each round. In some embodiments, event machine learning module 204 includes a set of instructions executable by processor 235 to generate event signals. In some embodiments, event machine learning module 204 is stored in memory 237 of computing device 200 and can be accessed and executed by processor 235.

[0052] In some embodiments, the event machine learning module 204 may use training data to generate a training model, specifically an event machine learning model. For example, the training data may include any type of data, such as media organized as events (e.g., images, videos, etc.), reactions to events (users watching events, sharing events, ordering print media during events, commenting on media items, etc.), and corresponding features (e.g., tags or labels associated with each media item that identifies the event, event type, objects in the media item, etc.).

[0053] Training data can be obtained from any source, such as a data repository specifically tagged for training, data that provides permission to use as training data for machine learning, etc. In embodiments where one or more users are permitted to use their respective user data to train a machine learning model, the training data may include such user data. In some embodiments, user data may include images / videos or image / video metadata (e.g., images, corresponding features that may originate from users who provide manually tagged or labeled data, geotags identifying the image location, timestamps associated with event dates, etc.), communications (e.g., messages on social networks; emails; chat data such as text messages, voice, video, etc.), documents (e.g., spreadsheets, text documents, presentations, etc.), etc.

[0054] In some embodiments, training data may include synthetic data generated for training purposes, such as data not based on user input or activity in the context of training, for example, data generated from simulated or computer-generated images / videos. In some embodiments, the event machine learning module 204 uses weights obtained from another application and not edited / transmitted. For example, in these embodiments, a training model may be generated, for example, on different devices and provided as part of media application 103. In various embodiments, the training model may be provided as a data file including a model structure or form (e.g., defining the number and type of neural network nodes, the connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. The event machine learning module 204 may read the data file of the training model and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the training model.

[0055] Event machine learning module 204 generates a training model, referred to herein as the event machine learning model. In some embodiments, event machine learning module 204 is configured to apply the event machine learning model to data, such as application data 266 (e.g., input media), to identify one or more features in the input media items and generate feature vectors (embeddings) representing the media items.

[0056] In some embodiments, the event machine learning model is a classifier that receives media items and information about those items, and uses that information to output an event signal indicating the likelihood of an event occurring in each round. This information may include the result of performing optical character recognition on an image to identify text within the image that indicates a specific event. For example, a dinner menu might include the word "wedding." In some embodiments, the information may include the result of performing object recognition to identify objects associated with the event. For example, a baby shower might have gifts associated with a baby.

[0057] In some implementations, metadata associated with the media item can be provided as additional input to the gating model, provided the user grants permission. This metadata can include user-permitted factors such as the video's capture location and / or time; whether the video was shared via social networks, image-sharing applications, messaging applications, etc.; depth information associated with one or more video frames; sensor values ​​from one or more sensors of the camera that captured the video, such as accelerometers, gyroscopes, light sensors, or other sensors; and the user's identity (if user consent has been obtained). For example, if the video was captured outdoors at night using an upward-pointing camera, this metadata could indicate that the camera was pointed towards the sky when capturing the video, thus associating it with astronomical events.

[0058] In some embodiments, the event machine learning module 204 may use a combination of metadata, optical character recognition, and other signals as input to the event machine learning model. For example, metadata might indicate that a person is setting off firecrackers and that the date is July 4th, causing the event machine learning module 204 to output an event signal corresponding to Independence Day. In some embodiments, the event machine learning module 204 may also output the event type of one or more events. Continuing with the example above, the event machine learning module 204 outputs the event signal and the probability that the event is Independence Day or a holiday.

[0059] In some embodiments, the event machine learning module 204 may include software code to be executed by the processor 235. In some embodiments, the event machine learning module 204 may specify a circuit configuration (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.) that enables the processor 235 to apply an event machine learning model. In some embodiments, the event machine learning module 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the event machine learning module 204 may provide an application programming interface (API) that can be used by the operating system 262 and / or other applications 264 to invoke the event machine learning module 204 to, for example, apply an event machine learning model to application data 266 to determine one or more features of an input image.

[0060] In some embodiments, an event-based machine learning model may include one or more model forms or structures. In some embodiments, an event-based machine learning model may use a support vector machine; however, in some embodiments, a convolutional neural network (CNN) is preferred. For example, the model form or structure may include any type of neural network, such as a linear network, a deep neural network that implements multiple layers (e.g., "hidden layers" between the input and output layers, each of which is a linear network), a convolutional neural network (CNN) (e.g., a network that splits or divides input data into multiple parts or blocks, processes each block individually using one or more neural network layers, and aggregates the processing results from each block), a sequence-to-sequence neural network (e.g., a network that takes sequence data (e.g., words in a sentence, frames in a video, etc.) as input and produces a sequence of results as output), etc.

[0061] The model form or structure can specify the connections between various nodes and the organization of nodes into layers. For example, nodes in the first layer (e.g., the input layer) can receive data as input data or application data. For example, when an event-driven machine learning model is used to analyze, for example, an input image (e.g., a first image associated with a user account), such data may include one or more pixels per node. Subsequent intermediate layers can receive the outputs of nodes in previous layers as input, depending on the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. The final layer (e.g., the output layer) produces the output for the machine learning application. For example, the output may be image features associated with the input image. In some embodiments, the model form or structure also specifies the number and / or type of nodes in each layer.

[0062] In some embodiments, the model form is a CNN with network layers, where each network layer extracts image features at a different level of abstraction. A CNN for identifying features in an image can be used for image classification. The model architecture can include combinations and orders of layers consisting of multidimensional convolutions, average pooling, max pooling, activation functions, normalization, regularization, and other layers and modules of deep neural networks used in practice for applications.

[0063] In various embodiments, the event-based machine learning model may include one or more models. These models may include multiple nodes, or, in the case of a CNN, a filter bank arranged in layers according to a model structure or form. In some embodiments, a node may be a computational node without memory, for example, configured to process an input unit to produce an output unit. The computation performed by the node may include, for example, multiplying each of the multiple node inputs by a weight to obtain a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. Different layers may include different types of inputs associated with media items, such as a first layer for metadata, a second layer for optical character recognition, and a third layer for object recognition.

[0064] In some embodiments, the computation performed by a node may further include applying a step / activation function to an adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, a single processing unit using a graphics processing unit (GPU), or a dedicated neural circuit. In some embodiments, a node may include memory, for example, being able to store and use one or more earlier inputs while processing subsequent inputs. For example, a node with memory may include a Long Short-Term Memory (LSTM) node. LSTM nodes may use memory to maintain the state that allows the node to act like a finite state machine (FSM). Models with such nodes can be used to process sequential data, such as words in a sentence or paragraph, a series of images, frames in a video, speech or other audio, etc. For example, a heuristic-based model used in a gating model may store one or more previously generated features corresponding to previous images.

[0065] In some embodiments, an event machine learning model may include embeddings or weights of individual nodes. For example, an event machine learning model may be initialized with multiple nodes organized into layers specified by a model form or structure. During initialization, appropriate weights may be applied to the connections between each pair of nodes connected according to the model form (e.g., nodes in consecutive layers of a neural network). For example, the appropriate weights may be randomly assigned or initialized to default values. The event machine learning model can then be trained, for example, using a training set of digital images to produce results. In some embodiments, a subset of the overall architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.

[0066] For example, training might involve applying supervised learning techniques. In supervised learning, training data could include multiple inputs (e.g., media organized as events) and a corresponding expected output for each input (e.g., an event signal for each event). Based on a comparison between the output of the event machine learning model and the expected output, the values ​​of the weights are automatically adjusted, for example, in a way that increases the probability that, when provided with inputs including real-valued event recognition provided as training data, the event machine learning model generates the expected output recognizing events from a media library by analyzing pixels and metadata.

[0067] In some embodiments, training may include applying unsupervised learning techniques. For example, only input data (e.g., media organized as events) may be provided, and an event machine learning model may be trained to distinguish the data, for example, to identify media that might be classified as events. In another example, an event machine learning model may train itself to generate events by distinguishing media items within an event.

[0068] In various embodiments, the training model includes a weight set corresponding to the model structure. In embodiments where the training set of digital images is omitted, the event machine learning module 204 may generate an event machine learning model based on previous training, for example, through the developer of the event machine learning module 204, through a third party, etc. In some embodiments, the event machine learning model may include a fixed weight set, for example, downloaded from a server that provides weights.

[0069] In some embodiments, the event machine learning module 204 can be implemented offline. Implementing the event machine learning module 204 may include using a static training set that does not include updates when the data in the static training set changes. This advantageously leads to increased processing efficiency and reduced power consumption of the computing device 200. In some embodiments, small updates to the event machine learning model can be implemented online, wherein updates to the training data are included as part of training the event machine learning model. A small update is an update smaller than a size threshold. The size of the update is related to the number of variables in the machine learning model affected by the update. In such embodiments, an application that invokes the event machine learning module 204 (e.g., operating system 262, one or more other applications 264, etc.) can utilize feature detections generated by the event machine learning unit 204 and can generate system logs (e.g., actions taken by the user based on the feature detections if permitted by the user; or results of further processing if used as input for further processing). The system logs can be generated periodically (e.g., hourly, monthly, quarterly, etc.) and can be used, with user permission, to update the event machine learning model, for example, to update the embeddings of the event machine learning model.

[0070] In some embodiments, the event machine learning module 204 may be implemented in a manner adaptable to a specific configuration of the computing device 200, and executed on the computing device 200. For example, the event machine learning module 204 may determine a computation graph utilizing available computing resources (e.g., processor 235). For example, if the event machine learning module 204 is implemented as a distributed application across multiple devices, such as media server 101 comprising multiple instances of media server 101, the event machine learning module 204 may determine, in a computationally optimized manner, the computations to be performed on each device. In another example, the event machine learning module 204 may determine that processor 235 includes a GPU with a specific number of GPU cores (e.g., 1000), and implement the event machine learning module 204 accordingly (e.g., as 1000 separate processes or threads).

[0071] In some embodiments, the event machine learning module 204 can integrate trained models. For example, the event machine learning model may include multiple trained models, each suitable for the same input data. In these embodiments, the event machine learning module 204 may select a particular trained model, for example, based on available computing resources, the success rate of prior inference, etc.

[0072] In some embodiments, the event machine learning module 204 may execute multiple training models. In these embodiments, the event machine learning module 204 may, for example, use a voting technique that scores the individual outputs from each applied training model, or combine the outputs from the applied individual models by selecting one or more specific outputs. In some embodiments, such a selector is part of the model itself and serves as a connection layer between the training models. Furthermore, in these embodiments, the event machine learning module 204 may apply a time threshold (e.g., 0.5 ms) to the applied individual training models and utilize only those individual outputs available within the time threshold. Outputs not received within the time threshold may not be utilized, for example, and may be discarded. This approach may be suitable, for example, when a time limit is specified when the event machine learning module 204 is invoked, for example, by the operating system 262 or one or more other applications 264. In this way, the maximum time spent by the event machine learning module 204 performing a task (e.g., identifying media that may be classified as an event) can be constrained, which improves the responsiveness of the media application 103 and results in the event machine learning unit 204 providing a real-time guarantee for best-effort classification.

[0073] The scoring module 206 generates an event importance score for each round. In some embodiments, the scoring module 206 includes a set of instructions executable by the processor 235 to score the round and determine one or more events from the round. In some embodiments, the scoring module 206 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0074] In some embodiments, the rating module 206 generates an event importance score for each round. The rating module 206 may generate the event importance score based on data from all users who have been granted access to its user data. This can be applied when a user has a lower-than-threshold number of media items in the library. In some embodiments, the rating module 206 generates the event importance score based on user-specific data (e.g., if the image includes the user or someone important to the user (e.g., someone appearing in a threshold number of media items associated with the user), if the round occurs in a place important to the user (e.g., frequently visited or familiar) or rare), based on previous user responses (sharing media items, printing media items, providing indications of approval for media items (e.g., liking media items), editing media items, viewing media items), or predictions of how the user will respond to the round. For example, if a user consistently views images of specific people, birthday parties, and Christmas, the rating module 206 may apply a score indicating that similar images are meaningful to the user to those similar images. In another example, the rating module 206 generates the event importance score based on the similarity of the round to an event considered important.

[0075] In some embodiments, the scoring module 206 further determines an event importance score based on the threshold number of media items in a corresponding round. For example, a round may require at least five media items (or seven or ten, etc.) to constitute an event. In some embodiments, the scoring module 206 further determines an event importance score based on a higher number of media items in a corresponding round, the presence of at least one face cluster at the threshold level in the corresponding round, or the presence of rare face clusters. In some embodiments, the scoring module 206 further determines an event importance score based on the quality indicators of the media items in the corresponding round. For example, the quality indicator may be based on whether the media item is blurry, oversaturated, or has out-of-focus objects, etc.

[0076] In some embodiments, the scoring module 206 combines multiple events into a single event based on the following: multiple events occurring within a 24-hour time period (e.g., numerous birthday celebrations); a single event occurring over multiple days (e.g., Hanukkah, Diwali, Ramadan, Cherry Blossom Festival); celebrations on different days associated with the single event (e.g., Indian weddings; Christmas, including lighting trees, family dinners, and opening gifts); multiple rounds involving the same location (e.g., workplaces where meals are held at the same location); or multiple rounds involving the same set of people identified by the same face cluster. In some embodiments, the scoring module 206 synthesizes events with a day array of up to 31 days because the event contains longer events (e.g., Ramadan is 30 days, Cherry Blossom Festival is 21 days).

[0077] In some embodiments, the scoring module 206 receives an event type (e.g., birthday) from the event machine learning module 204 and combines multiple events of that event type within a predetermined time period (e.g., combining all birthday-related celebrations that occur within five days of the birthday date plus or minus five days). For example, a birthday may include parties, but also related events such as dinners that occur around the birthday.

[0078] In some embodiments, the scoring module 206 receives the event type from the event machine learning module 204 and generates a confidence score indicating the likelihood that the corresponding event is an accurately identified event type. For example, the scoring module 206 may receive the same information described above regarding optical character recognition, object recognition, and metadata as referred to in the event machine learning module 204, and generate a confidence score for the event type based on text in a media item, objects identified in the media item corresponding to the event type (e.g., a Santa hat for Christmas, flowers and chocolate hearts for Valentine's Day), dates of specific events corresponding to known holidays (e.g., Christmas, Valentine's Day), etc. The confidence score can be a single value (1, 0.4, 500), a percentage (5%, 44%, etc.), or different metrics.

[0079] In some embodiments, the rating module 206 uses a rating machine learning model to output an event importance score. In this example, the rating machine learning model can receive user actions as input as part of a training set to train the model. Once trained, the rating machine learning model can receive rounds as input and output an event importance score.

[0080] In some embodiments, the scoring module 206 determines one or more events from these rounds based on event signals received by the event machine learning module 204 and corresponding event importance scores exceeding a threshold event importance value. The scoring module 206 may instruct the user interface module 210 to display events with corresponding event importance scores exceeding the threshold event importance value in a grid.

[0081] In some embodiments, events can be generated periodically. For example, one or more events can be limited to a predetermined number per month (e.g., seven) to avoid overwhelming the user with a variety of events. When new media is received, for example, by another user sharing media with a first user, the first user can associate the new media with their media library. The scoring module 206 generates an event importance score for rounds that include new media, and if the event importance score exceeds a threshold event importance value, the scoring module 206 determines whether the event importance score is greater than the event importance score of a specific event displayed in the user interface. The scoring module 206 can then instruct the user interface 210 to replace the previous event with the new event.

[0082] The heading module 208 generates a title for the event. In some embodiments, the heading module 208 includes a set of instructions executable by the processor 235 to generate the title. In some embodiments, the heading module 208 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0083] In some embodiments, the title module 208 receives event signals from the event machine learning module 204, which indicate the probability of an event occurring in each round. The title module 208 may also receive event type identification from the event machine learning module 204. The title module 208 can use the event type to determine a title describing the event type.

[0084] In some embodiments, the heading module 208 receives a confidence score from the scoring module 206, which indicates the probability that the corresponding event is an accurately identified event type. If the confidence score meets a threshold confidence value, the heading module 208 adds a title to the corresponding event based on the event type. For example, Figure 3The first example 600 includes an image of two people embracing, with the man holding flowers, where the confidence score meets a threshold, and the captioning module 208 applies the caption "Valentine's Day" to the corresponding event. This might happen, for example, if the scoring module 206 determines the image was taken as part of a Valentine's Day outing based on the image's location (a restaurant), the date the image was captured, and the man holding flowers. In another example, the captioning module 206 uses a high-confidence caption if it determines that more than 50% (or 40%, 90%, etc.) of the media items in the event have at least 90% (or 85%, etc.) confidence. In some embodiments, if the event has a high confidence value and occurs only once a year, the captioning module 208 appends the year to the caption (e.g., Diwali 2019).

[0085] If the confidence score fails to meet the threshold confidence value, the title module 208 adds a title to the corresponding event based on a template phrase. This template phrase implies the event type but is not specific enough to be interpreted as an error, such as "Celebrate love and laughter" for an event occurring shortly after the user's birthday. Continuing with the example above, Figure 3 Including the second example 350, where the confidence score fails to meet the threshold, the title module 208 applies the template phrase "Day of Hearts" because, although the image was taken on Valentine's Day, the location is a campsite and the woman is not holding flowers. The scoring module 206 can determine the confidence score based on the overall media items for the corresponding event.

[0086] The title machine learning module 258 generates a title for an event. In some embodiments, the title machine learning module 258 includes a set of instructions executable by the processor 235 to generate the title. In some embodiments, the title machine learning module 258 is stored in the memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0087] In some embodiments, the title machine learning module 258 may use training data to generate a training model, particularly a title machine learning model. For example, the training data may include any type of data, such as media organized as events (e.g., images, videos, etc.), the titles of the events, and corresponding features (e.g., tags or labels associated with each media item that identifies the occurrence of the event, the type of the event, objects in the media items, etc.).

[0088] Training data can be obtained from any source, such as data repositories specifically tagged for training, permitted data that provides training data for machine learning, etc. In embodiments where one or more users are permitted to use their respective user data to train machine learning models, training data may include such user data. In some embodiments, user data may include images / videos or image / video metadata (e.g., images, corresponding features that may originate from user-provided manual captions, tags, markers, geotags identifying image locations, timestamps associated with event dates, etc.), communications (e.g., messages on social networks; emails; chat data such as text messages, voice, video, etc.), documents (e.g., spreadsheets, text documents, presentations, etc.), etc.

[0089] In some embodiments, training data may include synthetic data generated for training purposes, such as data not based on user input or activity in the context of training, for example, data generated from simulated or computer-generated images / videos. In some embodiments, the title machine learning module 258 uses weights obtained from another application and not edited / transmitted. For example, in these embodiments, the training model may be generated, for example, on different devices and provided as part of the media application 103. In various embodiments, the trained model may be provided as a data file including a model structure or form (e.g., defining the number and type of neural network nodes, the connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. The title machine learning module 258 may read the data file of the training model and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the training model.

[0090] The title machine learning module 258 generates a training model referred to herein as the title machine learning model. In some embodiments, the title machine learning module 258 is configured to apply the title machine learning model to data, such as application data 267 (e.g., input media), to identify one or more features in the input media items and generate feature vectors (embeddings) representing the media items.

[0091] In some embodiments, media application 103 includes a title module 208 or a title machine learning module 258. In some embodiments, the title module 208 and the title machine learning module 258 are the same module. In some embodiments, the title machine learning module 258 is stored on a media server, while the remaining modules are stored on user device 115.

[0092] In some implementations, metadata associated with the media item can be provided as additional input to the gating model, provided the user grants permission. This metadata can include user-permitted factors such as the video's capture location and / or time; whether the video was shared via social networks, image-sharing applications, messaging applications, etc.; depth information associated with one or more video frames; sensor values ​​from one or more sensors (e.g., accelerometers, gyroscopes, light sensors, or other sensors) of the camera that captured the video; and the user's identity (if user consent has been obtained). For example, if the video was captured outdoors at night using an up-pointing camera, this metadata could indicate that the camera was pointed towards the sky when capturing the video, thus associating it with astronomical events.

[0093] In some embodiments, the title machine learning module 258 may include software code to be executed by the processor 285. In some embodiments, the title machine learning module 258 may specify a circuit configuration (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.) that enables the processor 285 to apply a title machine learning model. In some embodiments, the title machine learning module 258 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the title machine learning module 258 may provide an application programming interface (API) that can be used by the operating system 263 and / or other applications 265 to invoke the title machine learning module 258 to, for example, apply a title machine learning model to application data 267 to determine one or more features of the input image.

[0094] In some embodiments, the title machine learning model may include one or more model forms or structures. In some embodiments, the title machine learning model may use a support vector machine; however, in some embodiments, a convolutional neural network (CNN) is preferred. For example, the model form or structure may include any type of neural network, such as a linear network, a deep neural network implementing multiple layers (e.g., "hidden layers" between the input and output layers, each of which is a linear network), a convolutional neural network (CNN) (e.g., a network that splits or divides input data into multiple parts or blocks, processes each block individually using one or more neural network layers, and aggregates the results of the processing from each block), a sequence-to-sequence neural network (e.g., a network that takes sequence data such as words in a sentence, frames in a video, etc., as input and produces a sequence of results as output), and so on.

[0095] The model form or structure can specify the connections between various nodes and the organization of nodes into layers. For example, nodes in the first layer (e.g., the input layer) can receive data as input data or application data. For example, when a title machine learning model is used to analyze, for example, an input image (such as a first image associated with a user account), such data may include one or more pixels per node. Subsequent intermediate layers may receive the outputs of nodes in previous layers as input, depending on the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. The final layer (e.g., the output layer) produces the output for the machine learning application. For example, the output may be a title associated with the input image. In some embodiments, the model form or structure also specifies the number and / or type of nodes in each layer.

[0096] In some embodiments, the model form is a CNN with network layers, where each network layer extracts image features at a different level of abstraction. A CNN for identifying features in an image can be used for image classification. The model architecture can include combinations and orders of layers consisting of multidimensional convolutions, average pooling, max pooling, activation functions, normalization, regularization, and other layers and modules of deep neural networks used in practice for applications.

[0097] In various embodiments, the title machine learning model may include one or more models. These models may include multiple nodes, or, in the case of a CNN, a filter bank arranged in layers according to a model structure or form. In some embodiments, a node may be a computational node without memory, for example, configured to process an input unit to produce an output unit. The computation performed by the node may include, for example, multiplying each of the multiple node inputs by a weight to obtain a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. Different layers may include different types of inputs associated with media items, such as a first layer for metadata, a second layer for optical character recognition, and a third layer for object recognition.

[0098] In some embodiments, the computation performed by a node may further include applying a step / activation function to an adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, a single processing unit using a graphics processing unit (GPU), or a dedicated neural circuit. In some embodiments, a node may include memory, for example, being able to store and use one or more earlier inputs while processing subsequent inputs. For example, a node with memory may include a Long Short-Term Memory (LSTM) node. LSTM nodes may use memory to maintain the state that allows the node to act like a finite state machine (FSM). Models with such nodes can be used to process sequential data, such as words in a sentence or paragraph, a series of images, frames in a video, speech or other audio, etc. For example, a heuristic-based model used in a gating model may store one or more previously generated features corresponding to previous images.

[0099] In some embodiments, a title machine learning model may include embeddings or weights of individual nodes. For example, a title machine learning model may be initialized with multiple nodes organized into layers specified by a model form or structure. During initialization, appropriate weights may be applied to the connections between each pair of nodes connected according to the model form (e.g., nodes in consecutive layers of a neural network). For example, the appropriate weights may be randomly assigned or initialized to default values. The title machine learning model can then be trained, for example, using a training set of digital images to produce results. In some embodiments, a subset of the overall architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.

[0100] For example, training might involve applying supervised learning techniques. In supervised learning, training data could include multiple inputs (e.g., media organized as events) and a corresponding expected output for each input (e.g., a title for each event). The values ​​of the weights are automatically adjusted, for example, by comparing the output of the title machine learning model with the expected output, in a way that increases the probability that the title machine learning model will produce the expected output of appropriate blocks.

[0101] In some embodiments, training may include applying unsupervised learning techniques. For example, only input data (e.g., media organized as events and headlines) may be provided, and a headline machine learning model may be trained to differentiate the data, for example, to identify events associated with different headlines. In another example, the headline machine learning model may train itself to generate headlines by differentiating media items within an event.

[0102] In various embodiments, the training model includes a weight set corresponding to the model structure. In embodiments where the training set of digital images is omitted, the title machine learning module 258 may generate an event machine learning model based on previous training, for example, through the developer of the title machine learning module 258, through a third party, etc. In some embodiments, the title machine learning model may include a fixed weight set, for example, downloaded from a server that provides weights.

[0103] In some embodiments, the title machine learning module 258 can be implemented offline. In some embodiments, minor updates to the title machine learning model can be implemented online. In such embodiments, applications that invoke the title machine learning module 258 (e.g., operating system 263, one or more other applications 265, etc.) can utilize feature detections generated by the title machine learning module 258 and can generate system logs (e.g., actions taken by the user based on the feature detections if permitted by the user; or the results of further processing if used as input for further processing). The system logs can be generated periodically (e.g., hourly, monthly, quarterly, etc.) and can be used, with user permission, to update the title machine learning model, for example, to update the embeddings of the title machine learning model.

[0104] In some embodiments, the title machine learning module 258 may be implemented in a manner that adapts to a specific configuration of the computing device 200 on which the title machine learning module 258 is executed. For example, the title machine learning module 258 may determine a computation graph utilizing available computing resources (e.g., processor 285). For example, if the title machine learning module 258 is implemented as a distributed application across multiple devices, the title machine learning module 258 may determine, in a computationally optimized manner, the computations to be performed on individual devices. In another example, the title machine learning module 258 may determine that the processor 285 includes a GPU with a specific number of GPU cores (e.g., 1000), and implement the title machine learning module 258 accordingly (e.g., as 1000 separate processes or threads).

[0105] In some embodiments, the title machine learning module 258 can integrate trained models. For example, the title machine learning model may include multiple trained models, each adapted to the same input data. In these embodiments, the title machine learning module 258 may select a particular trained model, for example, based on available computing resources, the success rate of prior inference, etc.

[0106] In some embodiments, the title machine learning module 258 may execute multiple trained models. In these embodiments, the title machine learning module 258 may, for example, use a voting technique that scores the individual outputs from each applied trained model, or combine the outputs from the applied models by selecting one or more specific outputs. In some embodiments, such a selector is part of the model itself and serves as a connection layer between the trained models. Furthermore, in these embodiments, the title machine learning module 258 may apply a time threshold (e.g., 0.5 ms) to each individual trained model and utilize only those individual outputs available within the time threshold. Outputs not received within the time threshold may not be utilized, for example, and may be discarded. This approach may be suitable, for example, when a time limit is specified when the title machine learning module 258 is invoked, for example, via operating system 263 or one or more other applications 265.

[0107] In some embodiments, the title machine learning module 258 receives title edits from the user. The title machine learning module 258 incorporates the feedback into the title machine learning model to modify parameters to improve the title output.

[0108] User interface module 210 generates a user interface. In some embodiments, user interface module 210 includes a set of instructions executable by processor 235 to generate the user interface. In some embodiments, user interface module 210 is stored in memory 237 of computing device 200 and can be accessed and executed by processor 235.

[0109] In some embodiments, the user interface module 210 provides a user interface that includes event items with corresponding media from one or more events determined by the scoring module 206. A specific event can be highlighted more prominently in the user interface than other items in the user interface. For example, Figure 4 This includes the first example 400, which has a media item from March and an event item titled "The Happy Couple".

[0110] User interface module 210 may be part of a reverse time grid of media, which includes the corresponding media, a time grid, or an organization not based on the capture time of media items. Continuing the example above, Figure 4 A first example 400 of a reverse time grid 405 is shown, which includes media items from March 10, an event item titled "The Happy Couple," and media items from March 9. In some embodiments, the event with the highest score is selected by the user interface module 210 to appear in the reverse time grid. In some embodiments, a predetermined amount of time (e.g., a week, two days, etc.) must have elapsed for the event to appear in the reverse time grid.

[0111] In some embodiments, the user interface module 210 can generate user interfaces with different levels of prominence for events. For example, Figure 4 The first example 400 is a larger grid block where events are displayed as a large block scattered across a reverse-time grid with headers. In some embodiments, the user interface module 210 determines whether to display events as a larger grid block or a smaller grid block based on an event importance score, as described below.

[0112] Figure 4 A second example 425 is also included, featuring smaller grid blocks, in which a first media item from a particular event is less prominently highlighted, while a second media item from the same particular event is highlighted, and the event is displayed as scattered across a reverse-time grid. In some embodiments, the event in the user interface appears below the date header corresponding to that event.

[0113] Figure 4 It also includes a third example 450, in which a summary of the month's highlights and a list of images from the same month are displayed in the bottom portion 455 of the user interface, with specific events highlighted as such. Any events displayed inline in the reverse time grid are also included in the monthly highlights carousel in the bottom portion 455.

[0114] In some embodiments, the user interface module 210 displays the event in the user interface after a predetermined amount of time (e.g., 21 days). In some embodiments where the event includes multiple days spanning two different months, the user interface module 210 displays the event in a monthly highlight wheel based on the month the event ends. For example, if the event occurs between October 2nd and November 1st, the event will only be displayed in the November highlight wheel. In some embodiments, where multiple events occur in a single day, the user interface module 210 selects the event with the best event importance score to display for that day.

[0115] In some embodiments, when a user selects an event item from the grid, the user interface module 210 replaces the grid with the display of the corresponding media item from the specific event for a predetermined duration. For example, if the user selects an event item from the grid... Figure 4 If the event item "The Happy Couple" is selected, the user interface module 210 will display the proposal video, ring image, and image of the park where the proposal took place. The user interface module 210 can display each media item from a specific event for a predetermined time period (e.g., two seconds, five seconds, etc.). For example, the user interface module 210 can use the entire user interface to display each media item individually for up to two seconds.

[0116] In some embodiments, the user interface includes elements that allow the user to navigate forward and backward between media items at their own speed, and also includes elements that implement actions such as sharing, command printing, and marking as favorites. Once the display of media items is complete (or the user exits the display of the corresponding media item), the user interface module 210 can display the grid again.

[0117] In some embodiments, corresponding media from one or more events is displayed based on the magnitude of one or more of the following: event importance score, the number of media items for one or more events, the total number of events over a period of time, or event type (e.g., my wedding and my friend's wedding). For example, in Figure 4 In Example 400, the event item "The Happy Couple" is displayed in a larger size based on the event importance score indicating the highest ranking.

[0118] In some embodiments, the user interface module 210 receives a request from a user to remove human depictions from media in a media library that includes human depictions. For example, the user interface may allow the user to click on an image to remove it from the media library, the user interface may confirm that the image has been removed, and the user interface may provide a page with feedback on why the user requested the removal of the human depiction. The user interface module 210 may use the feedback as input to an event machine learning model and / or modify the event importance score. In some embodiments, the scoring module 206 filters human depictions from the media before generating the event importance score. This can advantageously prevent the user interface module 210 from displaying media associated with events having too few media items after removing human depictions from the media, or from displaying events featuring hidden human depictions or events characterized by hidden human depictions.

[0119] The user interface may include options for manipulating the media grid and media library. In some embodiments, the user interface includes options for hiding or adding one or more media items from one or more events, dates associated with one or more events, or people or pets depicted in media items from one or more events. Figure 5 Includes a sample user interface 500 that provides users with options to hide people and pets or to hide dates from media associated with events. The sample user interface 500 also includes options for selecting which memories to facilitate in reverse time-grid processing and options for managing memory-related notifications, such as whether to provide daily reminders, mute notifications, etc.

[0120] In some embodiments, the user interface includes an option to remove an event item. Removing an event item ensures that it will not be displayed to the user in the future, but the corresponding media item remains in the media library associated with the user account. In some embodiments, the user interface includes an option to remove a media item from an event item. Removing a media item from an event item does not remove the media item from the media library associated with the user account. Figure 6 This includes an example where the user interface comprises a confirmation request 600, a confirmation 625 indicating the media item has been removed, and a feedback screen 650 requesting feedback from the user regarding why the media item was removed from the event. Example feedback options include event item sensitive, duplicate, off-topic, poor quality, or others.

[0121] In some embodiments, the user interface module 210 transmits feedback to the event machine learning module 204, which uses the feedback to modify the parameters of the event machine learning model. In some embodiments, the user interface module 210 transmits feedback to the scoring module 206, which accordingly modifies the event importance score.

[0122] In some embodiments, the user interface includes options for modifying event details. For example, the user interface includes options for editing corresponding media from one or more events. In yet another example, the user interface includes options for changing the titles of one or more events when they are displayed in the user interface. In some embodiments, the title module 208 uses feedback to improve title generation.

[0123] Figure 7 An example user interface 700 according to some embodiments is shown, which has options for editing titles, removing events, changing the size of events in a grid, and changing the importance of events in the grid. In this example, the user can access these options by right-clicking an event in the user interface 700 or by some other mechanism. In some embodiments, these edits are available when an event is displayed in the user interface 700 or during the display of media items within an event.

[0124] Selecting "edit title" will cause the title to differ in user interface 700. Selecting "remove memory" will cause the media item associated with the event to remain in the user-associated library, but the event will not appear in the user interface in the future. Selecting "normal size" will cause the event to be displayed in user interface 700 at a normal size, such as... Figure 4 The second example 425, instead of... Figure 4The larger size in the first example 400 or the smaller size in the third example 450 is displayed in the user interface 700. Selecting "spotlight" will cause the event to be displayed in the user interface 700 at a larger size, such as... Figure 4 The first example is 400.

[0125] In some embodiments, the user interface module 210 generates audio for the corresponding media based on the type of one or more events. For example, for happy events such as graduation or a wedding, the music might be upbeat. For more serious events such as a funeral, the music might be more somber.

[0126] In some embodiments, in response to a user editing an aspect of an event, such as editing the event title, the user interface module 210 retains the event in the user interface even if the rating module 206 rates a different event higher than the edited event. This may be referred to as "freezing" the event. In some embodiments, the user interface module 210 retains the event in the user interface until the end of a predetermined period of time (e.g., October). Conversely, in some embodiments, if the user removes a media item from the library, the user interface module 210 removes that media item from the event.

[0127] In some embodiments, the user interface module 210 reorganizes different events in the user interface based on different changes. Go to Figure 8A Example block diagram 800 is shown, illustrating different examples of reorganizing different events based on variations in the number of events limited to seven using a grid in the user interface. First example 805 includes a list of events displayed in a grid where the corresponding event importance scores range from 70 / 100 to 99 / 100.

[0128] The user uploads media items from a DSLR camera, causing the scoring module 206 to generate a new event 810 with an event importance score of 81 / 100, which is higher than the five events in the grid. As a result, the event with the lowest event score (i.e., 70 / 100) is removed from the grid, and the new event is added. The user interface module 210 adds the new event 810 to the event list, resulting in a second example 815, where the new event 810 is added, and the event with the lowest event importance score of 70 / 100 is removed from the user interface. In the second example 815, the removed event is shown with a strikethrough.

[0129] continue Figure 8AIn the example, the user edits the title of an event with a score of 75 / 100, meaning the event is frozen and remains in the grid. The user then uploads a media item from a DSLR camera, and the scoring module 206 generates a new event 825 with a score of 77 / 100. Because the event with a score of 77 / 100 is higher than the event with a score of 75 / 100, the new event 825 would normally cause the second event to be removed from the grid. However, because the second event is frozen, it also remains in the grid. Therefore, Figure 8B A block diagram 850 with a third example 835 is shown, in which the grid now includes eight events.

[0130] continue Figure 8B In the example below, a user uploads a media item from a DSLR camera, and the scoring module 206 generates a new event with a score of 88 / 100. Because the new event 840 has a higher score than the previous event with a score of 77 / 100, and because the event with a score of 75 / 100 is frozen, the previous event with a score of 77 / 1000 is removed from the grid. A fourth example 845 shows the remaining events in the grid and the events removed from the grid with event importance scores of 77 / 100 and 70 / 100 (as indicated by strikethrough).

[0131] In some embodiments, the user interface module 210 generates a user interface to provide the user with event memories. For example, the user interface module 210 may generate an icon that appears at the top of the user's screen, which, when selected, causes the user interface to display media items associated with the event. In some embodiments, the user interface module 210 selects events that occurred a predetermined amount of time ago (e.g., one year ago), events with the highest event importance score, random events, etc.

[0132] Example method 900

[0133] Figure 9 This is a flowchart of an example method 900 for displaying event items according to some embodiments. Method 900 can be provided by... Figure 2 The computing device 200 (such as Figure 1 The user device 115 or media server 101 shown in the figure is executed.

[0134] Method 900 can begin with box 902. In box 902, the media library associated with the user account is divided into rounds, where each round is associated with a corresponding time period. Box 902 can be followed by box 904.

[0135] In box 904, the event machine learning model generates an event signal indicating the probability of an event occurring in each round, where the event machine learning model is a classifier that receives media as input. Box 904 may be followed by box 906.

[0136] In box 906, an event importance score is generated for each round. In some embodiments, the event importance score is based on one or more of the following: at least a threshold number of media items in the corresponding round, at least a threshold number of face clusters in the corresponding round, a quality indicator for the media items in the corresponding round, the presence of at least one face cluster of a threshold level, or the presence of rare face clusters. Box 906 may be followed by box 908.

[0137] In box 908, one or more events are determined from a round based on the event signal and the corresponding event importance score that exceeds a threshold event importance value. Box 908 may be followed by box 910.

[0138] In box 910, a user interface is provided, which includes event items of corresponding media for a specific event from one or more events. In some embodiments, the user interface is part of a time grid of media including the corresponding media. If a user selects an event item from the time grid, the time grid can be replaced by the display of the corresponding media item from the specific event for a predetermined duration, and the time grid is displayed in response to the completion of the display of the corresponding media item.

[0139] In addition to the above description, users can be given control over whether and when the system, program, or feature described herein can collect user information (e.g., information about the user's media items (such as photos or videos); the user's social networks; social behaviors or activities; occupation; user preferences, such as viewing preferences for image-based creations, settings for hiding people or pets, user interface preferences, etc.; or the user's current location); and whether content or communications are sent to the user from the server. Furthermore, some data may be processed in one or more ways before it is stored or used to remove personally identifiable information. For example, a user's identity may be processed so that personally identifiable information about the user cannot be determined, or the user's geographic location may be generalized (e.g., to the city, zip code, or state level) if location information is available, making it impossible to determine the user's specific location. Therefore, users can control what information about themselves is collected, how that information is used, and what information is provided to them.

[0140] In the foregoing description, numerous specific details have been set forth for purposes of explanation in order to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these specific details. In some cases, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described above primarily with reference to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, as well as any peripheral devices providing services.

[0141] References to "certain embodiments" or "certain examples" in the specification mean that a particular feature, structure, or characteristic described in connection with an embodiment or example may be included in at least one implementation described. The phrase "in some embodiments" appearing in various places throughout the specification does not necessarily refer to the same embodiment.

[0142] Some of the foregoing detailed descriptions have already been presented regarding the algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the most efficient way for those skilled in the art of data processing to communicate the essence of their work to others skilled in the art. An algorithm here is generally considered to be a self-consistent sequence of steps that leads to a desired result. This step is one that requires the physical manipulation of physical quantities. Typically, although not always necessary, these quantities take the form of electrical or magnetic data that can be stored, transmitted, combined, compared, and otherwise manipulated. Sometimes, primarily for general reasons, it has proven convenient to refer to these data as bits, values, elements, symbols, characters, items, or numbers, etc.

[0143] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise explicitly stated from the discussion above, it should be understood that throughout the description, the discussion using terms such as “processing” or “calculating” or “determining” or “displaying” refers to the actions and processes of computer systems or similar electronic computing devices that manipulate and convert data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the computer system's memory or registers or other such information storage, transmission, or display devices.

[0144] The embodiments of the specification may also relate to a processor for performing one or more steps of the methods described above. This processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including but not limited to any type of disk, including optical discs, ROMs, CD-ROMs, magnetic disks, RAM, EPROMs, EEPROMs, magnetic cards or optical cards, flash memory (including USB keys with non-volatile memory), or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0145] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0146] Furthermore, this description may take the form of a computer program product accessible from a computer-usable or computer-readable medium, which provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium may be any means that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0147] A data processing system suitable for storing or executing program code will include at least one processor directly or indirectly coupled to a memory element via a system bus. The memory element may include local memory used during the actual execution of the program code, mass storage, and a cache memory that provides temporary storage for at least some of the program code to reduce the number of times code must be retrieved from the mass storage during execution.

Claims

1. A computer-implemented method, comprising: The media library associated with a user account is divided into rounds, where each round is associated with a corresponding time period; An event signal is generated using an event machine learning model, which indicates the probability of an event occurring in each round, wherein the event machine learning model is a classifier that receives the media as input; Generate an event importance score for each round; One or more events are determined from the round based on the event signals and the corresponding event importance scores that exceed the threshold event importance value; Provide a user interface, the user interface including event items having corresponding media from specific events of the one or more events; Generate a confidence score, which indicates the probability that the corresponding event is an event type that has been accurately identified; and In response to the confidence score meeting the threshold confidence value, an automatically generated title describing the event type is added to the corresponding event.

2. The method according to claim 1, wherein, The user interface is a portion of the media's inverse time grid, including the corresponding media, and also includes: In response to a user selecting an event item from the inverse time grid, the inverse time grid is replaced with a display of the corresponding media from the specific event for a predetermined duration; and In response to the completion of the display of the corresponding media, the reverse time grid is displayed.

3. The method according to claim 1, further comprising: The event machine learning model was generated offline using a static training set. as well as The event machine learning model is updated in response to an update size that is less than a size threshold.

4. The method according to claim 1, further comprising: Receive a request from a user to remove the depiction of a person from the media in the media library that includes depictions of people; as well as The description of the person is filtered from the media, wherein the filtering is performed before the event importance score is generated.

5. The method according to claim 1, wherein, The user interface includes options to hide one or more of the following or to add one or more of the following: media items from the one or more events or dates associated with the one or more events.

6. The method according to claim 1, wherein, The user interface includes options for hiding or adding the following: people or pets depicted in media items from one or more of the events.

7. The method according to claim 1, further comprising: Determine which computations to perform on separate devices to optimize computation; as well as The event machine learning model is implemented on multiple devices based on the computation to be performed on the individual device.

8. The method according to claim 1, wherein, The event importance score is generated based on one or more of the following: at least a threshold number of media items in the corresponding round, at least a threshold number of facial clusters in the corresponding round, the quality indicator of the media items in the corresponding round, the presence of at least one facial cluster of the threshold level, or the presence of rare facial clusters.

9. The method according to claim 1, further comprising: Multiple rounds are combined into a single event based on one or more of the following: the multiple rounds are associated with corresponding time periods all falling within a 24-hour cycle; the single event is an event type that occurs over multiple days; celebrations on different days are associated with the single event; the multiple rounds involve the same location; or the multiple rounds involve the same facial cluster.

10. The method according to claim 1, further comprising: In response to the confidence score failing to meet the threshold confidence value, a title is added to the corresponding event based on a template phrase.

11. The method according to claim 1, wherein, The title machine learning model receives the corresponding media as input from the one or more events, and generates a title as output.

12. The method according to claim 1, wherein, The user interface includes options for editing the corresponding media from the specific event.

13. The method according to claim 1, wherein, Determining the one or more events includes determining events such that the number of events in each month is less than or equal to a predetermined number, and the method further includes; Receive new media to associate with the media library; as well as In response to a new event being associated with a new event importance score that is higher than that of the specific event, the specific event is replaced with the new event.

14. The method according to claim 1, wherein, The user interface includes an option to change the title of the event item.

15. The method of claim 1, further comprising generating audio for the corresponding media based on the type of the specific event.

16. A computing device, comprising: processor; as well as A memory coupled to the processor stores instructions thereon, which, when executed by the processor, cause the processor to perform operations, the operations including: The media library associated with a user account is divided into rounds, where each round is associated with a corresponding time period; An event signal is generated using an event machine learning model, which indicates the probability of an event occurring in each round, wherein the event machine learning model is a classifier that receives the media as input; Generate an event importance score for each round; Based on the event signals and the corresponding event importance scores that exceed the threshold event importance value, one or more events are determined from the round; A user interface is provided, the user interface including event items of corresponding media having specific events from the one or more events, wherein the user interface is a portion of the media comprising the corresponding media in reverse time grid; In response to a user selecting an event item from the inverse time grid, the inverse time grid is replaced with a display of the corresponding media from the specific event for a predetermined duration; and In response to the completion of the display of the corresponding media, the reverse time grid is displayed.

17. The computing device according to claim 16, wherein, The operation also includes: Receive a request from a user to remove the depiction of a person from media in the media library that includes depictions of people; and The description of the person is filtered from the media, wherein the filtering is performed before the event importance score is generated.

18. The computing device according to claim 16, wherein, The event items are displayed based on one or more of the following: the event importance score, the number of media items for the specific event, the total number of events in a time period, or the event type.

19. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the one or more computers to perform operations when executed by the one or more computers, the operations including: The media library associated with a user account is divided into rounds, where each round is associated with a corresponding time period; An event signal is generated using an event machine learning model, which indicates the probability of an event occurring in each round, wherein the event machine learning model is a classifier that receives the media as input; Generate an event importance score for each round; Based on the event signals and the corresponding event importance scores that exceed the threshold event importance value, events are determined from the rounds such that the number of events in each month is less than or equal to a predetermined number; Provide a user interface, the user interface including event items having corresponding media from a specific event of the event; Receive new media to associate with the media library; and In response to a new event being associated with a new event importance score that is higher than that of the specific event, the specific event in the event is replaced with the new event.

20. The computer-readable medium of claim 19, wherein, The user interface is a portion of the media's inverse time grid, including the corresponding media, and the operation further includes: In response to a user selecting an event item from the inverse time grid, the inverse time grid is replaced with a display of the corresponding media from the specific event for a predetermined duration; and In response to the completion of the display of the corresponding media, the reverse time grid is displayed.

21. The computer-readable medium of claim 19, wherein, The user interface includes options to hide one or more of the following or to add one or more of the following: media items from one or more events or dates associated with the one or more events.

22. The computer-readable medium of claim 19, wherein, The user interface includes options for hiding or adding the following: people or pets depicted in media items from one or more of the events.