Automatic event generation using machine learning models
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2021-12-13
- Publication Date
- 2026-07-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
【0010】 本明細書は、有利には、イベント機械学習モデルを用いて、イベントが各エピソードに発生した可能性を示すイベント信号を生成する方法を記載する。このようにして、メディアアイテムをイベントに分類するための改良方法を提供することができる。この方法は、例えば、事前に定義された分類またはカテゴリよりも、データ中の基礎的傾向をより確実に反映するイベントへの分類を提供することができる。さらに、機械学習モデルは、静的訓練セットを用いて、更新サイズが閾値サイズ未満であることに応答してイベント機械学習モデルを更新することによって、電力消費を有利に低減し、効率を高めることができる。
Smart Images

Figure 0007898441000001 
Figure 0007898441000002 
Figure 0007898441000003
Abstract
Description
Technical Field
[0001] Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 187,392, filed May 11, 2021, titled "Automated Generation of Events Using Machine Learning Models" and U.S. Provisional Patent Application No. 63 / 189,657, filed May 17, 2021, titled "Automated Generation of Events Using Machine Learning Models", and claims the benefit of priority of U.S. Patent Application No. 17 / 404,773, filed Aug. 17, 2021, titled "Automated Generation of Events Using Machine Learning Models", the entire contents of each application being incorporated herein by reference.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Background Users of devices such as smartphones or other digital cameras take and store a large amount of media (e.g., photos and videos) in a library. By accessing the library and viewing their media, users recall various events such as birthdays, weddings, vacations, and trips. However, libraries often contain thousands of images taken over a long period of time and are difficult to organize.
[0003] The description of the background art set forth herein is for the purpose of generally indicating the context of the disclosure. Within the scope described in this background art section, the research of the inventors, as presently named, is not expressly or implicitly admitted as prior art to the disclosure, any more than is the description that is not regarded as prior art at the time of the application.
Means for Solving the Problems
[0004] Overview A computer-implemented method includes segmenting a media library associated with a user account into episodes, each episode relating to a corresponding period, and generating event signals using an event machine learning model that indicate the likelihood of an event occurring in each episode, the event machine learning model being a classifier that takes media as input, and the method further includes generating an event importance score for each episode, determining one or more events from the episode based on the event signals and corresponding event importance scores that exceed an event importance threshold, and providing a user interface that includes event items having corresponding media from a particular event among the one or more events.
[0005] In some embodiments, the user interface is part of an inverse time-series grid of media containing the corresponding media, and the method further includes replacing the inverse time-series grid with a display of the corresponding media from a particular event over a predetermined period in response to a user selecting an event item from the inverse time-series grid, and displaying the inverse time-series grid in response to the completion of the display of the corresponding media. In some embodiments, the method further includes generating an event machine learning model offline using a static training set, and updating the event machine learning model in response to the update size being less than a threshold size. In some embodiments, the method further includes receiving a request from the user to remove depictions of people from media in a media library containing depictions of people, and filtering out depictions of people from the media, the filtering being performed before generating an event importance score. In some embodiments, the method further includes determining computations to be performed on individual devices to optimize the computation, and implementing the event machine learning model on multiple devices based on the computations performed on individual devices. In some embodiments, the user interface includes an option to hide event items. In some embodiments, generating an event importance score is based on one or more of the following: at least a threshold number of media items in a corresponding episode, at least a threshold number of face clusters in a corresponding episode, a quality metric for media items in a corresponding episode, at least one face cluster of a threshold rank, or the presence of rare face clusters. In some embodiments, the method further comprises merging multiple episodes into a single event based on one or more episodes each associated with multiple periods, where the multiple periods comprise a 24-hour period as a whole, and the single event is an event type occurring over multiple days, celebrations on different days associated with a single event, multiple episodes associated with the same location, or multiple episodes associated with the same set of face clusters.In some embodiments, the method further includes generating a confidence score indicating the likelihood that a corresponding event is of an accurately recognized event type, and adding an automatically generated title describing the event type to the corresponding event in response to the confidence score meeting a confidence threshold. In some embodiments, the method further includes adding a title based on a template representation to the corresponding event in response to the confidence score not meeting a confidence threshold. In some embodiments, the title machine learning model receives corresponding media from one or more events as input, and the title machine learning model generates a title as output. In some embodiments, the user interface includes an option for editing corresponding media from a particular event. In some embodiments, determining one or more events includes determining events such that the number of events is less than or equal to a predetermined number per month, receiving new media related to the media library, and replacing a particular event among one or more events with the new event in response to the new event being associated with a new event importance score higher than the event importance score of a particular event. In some embodiments, the user interface includes an option for changing the title of an event item. In some embodiments, the method further includes generating audio for the corresponding media based on a particular event type.
[0006] Embodiments may further include a system comprising one or more processors and memory for storing instructions executed by the one or more processors. The instructions include segmenting a media library associated with a user account into episodes, each episode relating to a corresponding period, and generating an event signal using an event machine learning model that indicates the likelihood that an event occurred in each episode, the event machine learning model being a classifier that takes media as input, generating an event importance score for each episode, determining one or more events from the episode based on the event signal and corresponding event importance scores that exceed an event importance threshold, and providing a user interface that includes an event item having corresponding media from a particular event among the one or more events.
[0007] In some embodiments, the user interface is part of a reverse time-series grid of media including the corresponding media, and the method further includes replacing the reverse time-series grid with a display of the corresponding media from a particular event over a given period in response to the user selecting an event item from the reverse time-series grid, and displaying the reverse time-series grid in response to the completion of the display of the corresponding media. In some embodiments, event items are displayed by size based on one or more of the following: an event importance score, the number of media items for a particular event, the total number of events in one period, or the event type.
[0008] The embodiment may further include a non-temporary computer-readable medium that, when executed by one or more computers, stores instructions causing one or more computers to perform the following actions: The actions include segmenting a media library associated with a user account into episodes, each episode relating to a corresponding period, and generating event signals using an event machine learning model that indicate the likelihood that an event occurred in each episode, the event machine learning model being a classifier that takes media as input; the actions further include generating an event importance score for each episode, determining one or more events from the episode based on the event signals and corresponding event importance scores that exceed an event importance threshold, and providing a user interface that includes an event item having corresponding media from a particular event among the one or more events.
[0009] In some embodiments, the user interface is part of a reverse time-series grid of media including the corresponding media, and the method further includes replacing the reverse time-series grid with a display of the corresponding media from a particular event over a predetermined period of time in response to the user selecting an event item from the reverse time-series grid, and displaying the reverse time-series grid in response to the completion of the display of the corresponding media. [Effects of the Invention]
[0010] This specification advantageously describes a method for generating event signals that indicate the likelihood of an event occurring in each episode, using an event machine learning model. In this way, an improved method for classifying media items into events can be provided. This method can, for example, provide a classification to events that more reliably reflects underlying trends in the data than a predefined classification or category. Furthermore, the machine learning model can be made more efficient and power-efficient by updating the event machine learning model in response to the update size being below a threshold size using a static training set. [Brief explanation of the drawing]
[0011] [Figure 1] This block diagram shows an exemplary network environment according to some embodiments described herein. [Figure 2] A block diagram illustrating an exemplary computing device according to some embodiments described herein. [Figure 3] This figure shows exemplary titles based on confidence scores indicating the likelihood that the corresponding event is of the precisely recognized event type, according to several embodiments. [Figure 4] This figure shows an exemplary inverse time-series grid of media according to several embodiments. [Figure 5] This figure shows an exemplary user interface for removing media items from an event, according to several embodiments. [Figure 6] This figure shows an exemplary user interface that, according to several embodiments, removes media items from event items and includes options for providing feedback. [Figure 7] According to several embodiments, exemplary user interfaces are illustrated that include options for editing titles, deleting events, resizing events in an inverse time series grid, and changing the importance of events in an inverse time series grid. [Figure 8A] This is an exemplary block diagram illustrating different examples of reorganizing different events based on changes, according to several embodiments. [Figure 8B] This is an exemplary block diagram illustrating different examples of reorganizing different events based on changes, according to several embodiments. [Figure 9] This flowchart illustrates exemplary methods for displaying event items according to several embodiments. [Modes for carrying out the invention]
[0012] Detailed explanation Network environment 100 Figure 1 shows a block diagram of an exemplary environment 100. In some embodiments, the environment 100 includes a media server 101, a user device 115a, a user device 115n, and a network 105. Users 125a and 125n may be associated with user devices 115a and 115n, respectively. In some embodiments, the environment 100 may include other servers or devices not shown in Figure 1, or may not include the media server 101. In Figure 1 and other drawings, letters following a reference number, such as "115a," indicate a reference to the element having that particular reference number. Reference numbers in the text without letters following them, such as "115," indicate a general reference to embodiments of the element having that reference number.
[0013] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicably connected to the network 105 via a signal line 102. The signal line 102 may be a wired connection such as Ethernet®, coaxial cable, or fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.
[0014] Media application 103a may include code and routines that can operate to receive a media library associated with a user account. Media items referred to herein may include images or videos. Images may include digital images having pixels with one or more pixel values (e.g., color values, luminance values, etc.). Images may be still images (e.g., still images, single-frame images, etc.) or dynamic images (images containing multiple frames, e.g., videos, animated GIFs, or cinemagraphs where part of the image contains motion and other parts are static, etc.). Videos referred to herein may contain multiple frames, with or without sound. In some implementations, one or more camera settings, e.g., zoom level, aperture, etc., may be changed during video recording. In some implementations, the client device recording the video may be moved during video recording. Text referred to herein may include alphanumeric characters, emojis, symbols, or other characters.
[0015] The media application 103a may include code and routines that can operate to segment the media library into episodes. The media application 103a may use an event machine learning model to generate event signals indicating the likelihood that an event occurred in each episode. The media application 103a may generate an event importance score for each episode and determine one or more events from an episode based on the event signals and corresponding event importance scores that exceed an event importance threshold. The media application 103a may display event items, including corresponding media from libraries associated with a user account, in the user interface.
[0016] In some embodiments, media application 103a may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any other type of processor, or a combination thereof. In some embodiments, media application 103a may be implemented using a combination of hardware and software.
[0017] Database 199 can store a media library, a training set of a machine learning model, and user actions related to media (such as browsing, sharing, annotating, etc.). Database 199 can store indexed media items associated with the ID of user 125 of user device 115. Also, database 199 can store social network data related to user 125, user preferences of user 125, and the like.
[0018] User device 115 may be a computing device including a memory and a hardware processor. For example, user device 115 may include a desktop computer, a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or any other electronic device capable of accessing network 105.
[0019] In the illustrated implementation, user device 115a is connected to network 105 via signal line 108, and user device 115n is connected to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections such as Ethernet (registered trademark), coaxial cable, fiber optic cable, or wireless connections such as Wi-Fi (registered trademark), Bluetooth (registered trademark), or other wireless technologies. User devices 115a, 115n are each utilized by users 125a, 125n. User devices 115a, 115n in FIG. 1 are used as an example. FIG. 1 shows two user devices 115a and 115n, but the present disclosure applies to system architectures including one or more user devices 115.
[0020] In some embodiments, media application 103 receives a media library related to the user. For example, the user may obtain images and videos from their camera (e.g., smartphone or other camera), upload images from a digital single-lens reflex (DSLR) camera, and add media taken and shared by other users to the media library. Media application 103 segments the media library into episodes, and each episode is related to a corresponding period. For example, media related to trick or treat may be considered an event because the media is related to a period of several hours and the location is a single neighborhood or city. Many episodes may include media of a type that the user does not repeatedly engage with. For example, a day at home where the user takes pictures of various broken items in order to purchase replacement parts at a hardware store.
[0021] The media application 103 may include a trained machine learning model, such as an event machine learning model, which receives media as input and generates event signals indicating the likelihood that an event occurred in each episode. For example, images taken in different locations over a 24-hour period that do not share the same theme may be associated with an event signal indicating that an event is unlikely to have occurred. Conversely, images taken in the same location over a 24-hour period, such as a picture of a cake, a video of children singing "Happy Birthday," and a room full of balloons, may be associated with an event signal indicating that an event (e.g., "birthday party") is likely to have occurred.
[0022] The media application 103 can generate an event importance score for each episode. The media application 103 may use a machine learning model or a different type of module for this step. The event importance score may be related to the likelihood that a user will want to engage with the media in the episode. For example, if the event importance score exceeds an event importance threshold, the media application 103 may determine that the episode corresponds to an event in which a user will want to view, share, and send printed media in a photo album. In some embodiments, the event importance score is based on one or more of the following: at least a threshold number of media items in the corresponding episode, at least a threshold number of face clusters in the corresponding episode, a quality metric for the media items in the corresponding episode, at least one face cluster of a threshold rank, or the presence of rare face clusters.
[0023] In some embodiments, the media application 103 determines events from an episode based on both the event signal output by the event machine learning model and the corresponding event importance score that exceeds the event importance threshold. In some embodiments, the default value for an event is 24 hours, but events that appear to be connected in a particular way, such as events that occurred over multiple days, can also be combined.
[0024] In some embodiments, the media application 103 provides a user interface that includes corresponding media from events. The user interface may be part of an inverse time-series grid of media that includes corresponding media. For example, the user interface may include a "monthly carousel" of media in which media corresponding to events that occurred during February are organized in the grid. In some embodiments, the user interface limits the number of events displayed in the grid according to a predetermined number. For example, the user interface may include only seven events in a particular month. If seven events are already displayed in the user interface and a new media relates to a new event that has a higher new event importance score than the corresponding event importance scores of the other seven events, the user interface may replace the previously displayed events with the new event.
[0025] Computing device 200 Figure 2A is a block diagram of an exemplary computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a user device 115 used to run a media application 103. In another example, the computing device 200 is a media server 101. In yet another example, the media application 103 is partially located on the user device 115 and partially on the media server 101.
[0026] One or more of the methods described herein can be performed as a standalone program running on any type of computing device, a program running on a web browser, or a mobile application (app) running on a mobile computing device (e.g., a mobile phone, smartphone, tablet computer, wearable device (e.g., a watch, armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted display), or laptop computer). In the primary example, all computations are performed in a mobile application (and / or other application 264) on the mobile computing device. However, a client / server architecture can be used. For example, the mobile computing device sends user input data to a server device and receives and outputs (e.g., displays) the final output data from the server. In another example, the computations may be shared between the mobile computing device and one or more server devices.
[0027] In some embodiments, the computing device 200 includes a processor 235, memory 237, I / O interface 239, display 241, camera 243, and storage device 245. The processor 235 may be connected to a bus 218 via a signal line 222. The memory 237 may be connected to a bus 218 via a signal line 224. The I / O interface 239 may be connected to a bus 218 via a signal line 226. The display 241 may be connected to a bus 218 via a signal line 228. The camera 243 may be connected to a bus 218 via a signal line 230. The storage device 245 may be connected to a bus 218 via a signal line 232.
[0028] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling the basic operation of the computing device 200. “Processor” includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. The processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configurations), multiple processing units (e.g., having a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for achieving a function, a dedicated processor for performing processing based on a neural network model, a neural circuit, a system with a processor optimized for performing matrix calculations (e.g., matrix multiplication), or other systems. In some implementations, the processor 235 may include one or more coprocessors for performing neural network processing. In some implementations, the processor 235 may be a processor that generates a probabilistic output by processing data. For example, the output generated by the processor 235 may be inaccurate or accurate within the range of an expected output. The processing does not need to be limited to a specific geographical location or time. For example, a processor can perform functions in real time, offline, or batch mode. Parts of the processing may be performed by different (or the same) processing systems at different times and locations. The computer may be any processor that communicates with memory.
[0029] Memory 237 is typically located within the computing device 200 for use by the processor 235 and may be any suitable processor-readable storage medium for storing instructions executed by the processor or set of processors, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory. Memory 237 may be located separately from the processor 235 and / or integrated with it.
[0030] Memory 237 can store software, including other applications 264, application data 266, and media applications 103, which are executed on the computing device 200 by the operating system 262 and processor 235. Other applications 264 may include applications such as camera applications, image gallery or image library applications, data display engines, web hosting engines, image display engines, notification engines, and social networking engines. In some implementations, media applications 103 and other applications 264 may each include instructions that enable the processor 235 to perform the functions described herein.
[0031] The application data 266 may also be data generated by other applications 264 or hardware of the computing device 200. For example, the application data 266 may include images captured by the camera 243, user behavior identified by other applications 264 (e.g., a social networking application), and so on.
[0032] The I / O interface 239 can provide functionality that enables the computing device 200 to interface with other systems and devices. Interface devices may be included as part of the computing device 200 or may be separate but capable of communicating with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or database 199), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can be connected to interface devices, such as input devices (keyboard, pointing device, touchscreen, microphone, camera, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.). For example, when a user provides touch input, the I / O interface 239 transmits data to the media application 103.
[0033] Some exemplary interface connection devices that can be connected to the I / O interface 239 may include a display 241 that can be used to display the user interface of the content described herein, e.g., images, video and / or output applications, and to receive touch (or gesture) input from the user. For example, the display 241 can be used to display a user interface that includes corresponding media from one or more events. The display 241 may include any suitable display device, e.g., a liquid crystal display (LCD), a light-emitting diode (LED), or a plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen on a computer device.
[0034] Camera 243 may be any type of image capture device capable of capturing images and / or videos. In some embodiments, camera 243 captures images or videos that the I / O interface 239 transmits to the media application 103.
[0035] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store a media library associated with a user account, a selected media set, a training set for a machine learning model, and so on. In embodiments where the media application 103 is part of the media server 101, the storage device 245 is the same as the database 199 in Figure 1.
[0036] First exemplary media application 103 The media application 103 shown in Figure 2A includes a segmentation module 202, an event machine learning module 204, a scoring module 206, a title generation module 208, a title machine learning module, and a user interface module 210.
[0037] The segmentation module 202 segments the media library associated with a user account into episodes. In some embodiments, the segmentation module 202 includes a set of instructions that can be executed by the processor 235 to segment the media library. In some embodiments, the segmentation module 202 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0038] In some embodiments, an episode containing multiple media items is defined as containing media related to timestamps within a corresponding period. For example, the segmentation module 202 may define an episode as something that occurred within a period such as 24 hours, 12 hours, or 2 days. In some embodiments, the duration of an episode is determined automatically. In some embodiments, the user can define the duration, for example, through a user interface.
[0039] The event machine learning module 204 generates event signals that indicate the likelihood of an event occurring in each episode. In some embodiments, the event machine learning module 204 includes a set of instructions that can be executed by the processor 235 to generate event signals. In some embodiments, the event machine learning module 204 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0040] In some embodiments, the event machine learning module 204 can use training data to generate a trained model, specifically an event machine learning model. For example, the training data may include any type of data, such as media organized as events (e.g., images, videos, etc.), reactions to events (e.g., users viewing events, sharing events, ordering prints of media within events, commenting on media items, etc.), and corresponding features (e.g., labels or tags associated with each media item to identify that an event occurred, the type of event, or objects within the media item, etc.).
[0041] Training data may be obtained from any source, e.g., a data repository specifically designated for training, or data for which permission has been granted to be used as training data for machine learning. In embodiments where one or more users permit the use of their respective user data to train a machine learning model, the training data may include this user data. In some embodiments, user data may include images / videos or image / video metadata (e.g., images, corresponding features obtained from users providing manual tags or labels, geotags identifying the location of images, timestamps related to the date of an event), information (e.g., chat data such as messages on social networks, emails, text messages, audio, video), documents (e.g., spreadsheets, text documents, presentations), etc.
[0042] In some embodiments, the training data may include synthetic data generated for training purposes, such as data not based on user input or activity in the situation being trained, such as data generated from simulations or computer-generated images / videos. In some embodiments, the event machine learning module 204 uses weights obtained from another application and not edited / transferred. For example, in these embodiments, the trained model may be generated on a different device, for example, and provided as part of the media application 103. In various embodiments, the trained model may be provided as a data file containing the model structure or form (defining, for example, the number and types of neural network nodes, the connections between nodes, and organizing the nodes into multiple layers) and the associated weights. The event machine learning module 204 can read the data file of the trained model and implement a neural network, including node connections, layers, and weights, based on the model structure or form specified in the trained model.
[0043] The event machine learning module 204 generates a trained model, referred to herein as an event machine learning model. In some embodiments, the event machine learning module 204 is configured to identify one or more features within input media items and generate (embedded) feature vectors representing the media items by applying the event machine learning model to data such as application data 266 (e.g., input media).
[0044] In some embodiments, the event machine learning model is a classifier that receives media items along with information about the media items and uses that information to output an event signal indicating the likelihood that the event occurred in each episode. The information may include the results of optical character recognition performed on an image to identify text within the image that indicates a particular event. For example, a meal menu might contain the word "wedding." In some embodiments, the information may include the results of object recognition performed to identify objects associated with the event. For example, a baby shower might have gifts associated with the baby.
[0045] In some implementations, metadata associated with media items may be provided as additional input to the gating model if permitted by the user. The metadata may include user permission factors, such as the location and / or time the video was taken, whether the video was shared via a social network, image sharing application, messaging application, etc., depth information associated with one or more video frames, sensor values from one or more sensors of the camera that took the video (e.g., accelerometer, gyroscope, light sensor, or other sensors), and (with the user's consent) the user's ID. For example, if a video is taken outdoors at night with the camera pointed upwards, the metadata may indicate that the camera was pointed towards the sky when the video was taken, and therefore related to an astronomical event.
[0046] In some embodiments, the event machine learning module 204 may use a combination of metadata, optical character recognition, and other signals as input to the event machine learning model. For example, the metadata may indicate that a person is using firecrackers and the date is July 4th, and as a result, the event machine learning module 204 outputs an event signal corresponding to Independence Day. In some embodiments, the event machine learning module 204 outputs the event type for one or more events. Continuing the above example, the event machine learning module 204 outputs an event signal and the likelihood that the event is Independence Day or a holiday.
[0047] In some embodiments, the event machine learning module 204 may include software code executed by the processor 235. In some embodiments, the event machine learning module 204 may specify a circuit configuration (e.g., a programmable processor, a field-programmable gate array (FPGA)) that enables the processor 235 to apply an event machine learning model. In some embodiments, the event machine learning module 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the event machine learning module 204 may provide an application programming interface (API). The operating system 262 and / or other applications 264 can use this API to call the event machine learning module 204 and determine one or more features of an input image, for example, by applying an event machine learning model to application data 266.
[0048] In some embodiments, an event machine learning model may include one or more model forms or structures. In some embodiments, an event machine learning model may use a support vector machine, but in some embodiments, a convolutional neural network (CNN) is preferred. For example, the model form or structure may include any type of neural network, e.g., a linear network, a deep neural network implementing multiple layers (e.g., "hidden layers" between input and output layers, each of which is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results obtained from processing each tile), or a sequence-to-sequence neural network (e.g., a network that takes sequential data as input, such as words in a sentence or frames in a video, and produces a result sequence as output).
[0049] The model configuration or structure can specify the connections between various nodes and the organization of nodes into layers. For example, the nodes in the first layer (e.g., the input layer) can receive data as input data or application data. For example, if an event machine learning model is used to analyze input images associated with user accounts, e.g., the first image, such data may include, for example, one or more pixels per node. Subsequent intermediate layers can receive the outputs of the nodes in the previous layer as input, according to the connections specified in the model configuration or structure. These layers are sometimes called hidden layers. The final layer (e.g., the output layer) generates the output of the machine learning application. For example, this output may be image features associated with the input image. In some embodiments, the model configuration or structure specifies the number and / or types of nodes in each layer.
[0050] In some embodiments, the model configuration is a CNN including network layers, each network layer extracting image features at a different level of abstraction. The CNN used to identify features in an image may also be used to classify the image. The model architecture may include combinations and sequences of layers consisting of multidimensional convolution, mean pooling, max pooling, activation functions, normalization, regularization, and other layers and modules actually used in applied deep neural networks.
[0051] In different embodiments, an event machine learning model may include one or more models. One or more models may include multiple nodes arranged in layers according to a del structure or morphology, or, in the case of a CNN, a filter bank. In some embodiments, a node may be a memoryless computation node configured to process one unit of input and produce one unit of output. The computation performed by the node may include, for example, the steps of multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and producing a node output by adjusting the weighted sum with a bias value or intercept value. Different layers may include different types of inputs related to media items, such as a first layer for metadata, a second layer for optical character recognition, and a third layer for object recognition.
[0052] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to a tuned weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multicore processor, using individual processing units of a graphical processing unit (GPU), or using dedicated neural circuits. In some embodiments, nodes may include memory. A node may, for example, remember one or more previous inputs and use one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long-short-term memory (LSTM) node. An LSTM node can use memory to maintain state that allows the node to behave like a finite state machine (FSM). Models including such nodes would be useful when processing sequential data, such as multiple words in a sentence or paragraph, a series of images, frames in a video, conversation, or other audio. For example, a heuristics-based model used in a gating model can remember one or more previously generated features for a previous image.
[0053] In some embodiments, the event machine learning model may include embeddings or weights for individual nodes. For example, the event machine learning model may be initialized as a group of nodes organized into layers as specified by the model morphology or structure. During initialization, each weight can be applied to the connections between each pair of nodes connected according to the model morphology, for example, between each pair of nodes in a continuous layer of a neural network. For example, each weight may be assigned randomly or initialized to a default value. The event machine learning model can then be trained, for example, with a training set of digital images to produce results. In some embodiments, a subset of the entire architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.
[0054] For example, training may involve applying supervised learning techniques. In supervised learning, training data can include multiple inputs (e.g., media organized as events) and expected outputs corresponding to each input (e.g., event signals for each event). For example, given inputs containing ground truth event identification provided as training data, the event machine learning model analyzes pixels and metadata to automatically adjust weight values based on a comparison of the event machine learning model's output with the expected output, thereby increasing the probability that it will generate an expected output identifying an event from a media library.
[0055] In some embodiments, training may involve applying unsupervised learning techniques. For example, only input data (e.g., media organized as events) may be provided, and the event machine learning model may be trained to distinguish the data, for example, to identify media that are likely to be classified as events. In another example, the event machine learning model may be trained to generate events by distinguishing media items within events.
[0056] In various embodiments, the trained model includes a set of weights corresponding to the model structure. In embodiments where the training set of digital images is omitted, the event machine learning module 204 may generate an event machine learning model based on prior training by, for example, the developer of the event machine learning module 204 or a third party. In some embodiments, the event machine learning model may include a set of fixed weights downloaded from a server that provides the weights.
[0057] In some embodiments, the event machine learning module 204 may be implemented offline. Implementing the event machine learning module 204 may include using a static training set that does not include updates when data in the static training set changes. This is advantageous as it improves the efficiency of processing performed by the computing device 200 and reduces the power consumption of the processing device 200. In some embodiments, small updates to the event machine learning model may be implemented online, where updates to the training data are included as part of training the event machine learning model. A small update is an update with a size smaller than a threshold size. The size of the update is related to the number of variables in the machine learning model affected by the update. In such embodiments, an application calling the event machine learning module 204 (e.g., an operating system 262, one or more other applications 264) can utilize the feature detections generated by the event machine learning module 204 and can generate a system log (e.g., actions taken by the user based on the feature detections, if permitted by the user, or the results of further processing, if used as input for further processing). The system log may be generated periodically, for example, every hour, every month, or every three months, and may be used to update the event machine learning model, for example, to update the embedding of the event machine learning model, if permitted by the user.
[0058] In some embodiments, the event machine learning module 204 may be implemented in a manner that conforms to a specific configuration of the computing device 200 on which the event machine learning module 204 is executed. For example, the event machine learning module 204 can determine a computation graph that utilizes available computing resources, such as the processor 235. If the event machine learning module 204 is implemented as a distributed application across multiple devices, for example, if the media server 101 includes multiple media servers 101, the event machine learning module 204 can determine which computations are performed on each device to optimize the computation. In another example, if the event machine learning module 204 determines that the processor 235 includes a GPU with a certain number (e.g., 1000) GPU cores, the event machine learning module 204 can be implemented (e.g., as 1000 separate processes or threads).
[0059] In some embodiments, the event machine learning module 204 can implement a set of trained models. For example, the event machine learning model may include multiple trained models, each applicable to the same input data. In these embodiments, the event machine learning module 204 can select a particular trained model based, for example, on available computational resources, the success rate when using previous inferences, etc.
[0060] In some embodiments, the event machine learning module 204 can run multiple trained models. In these embodiments, the event machine learning module 204 can synthesize outputs, for example, by using a majority vote to score the outputs obtained by applying each trained model, or by selecting one or more specific outputs. In some embodiments, such selectors are part of the model itself and function as a connecting layer between the trained models. Furthermore, in these embodiments, the event machine learning module 204 can apply a time threshold (e.g., 0.5 ms) for applying each trained model and utilize only the individual outputs available within the time threshold. Outputs not received within the time threshold are not utilized and may be discarded, for example. For example, such a technique would be appropriate when there is a specified time limit between calling the event machine learning module 204, for example, by the operating system 262 or one or more other applications 264. In this way, the maximum time required for the event machine learning module 204 to perform a task, for example, to identify media that are likely to be classified as events, can be limited, thereby improving the responsiveness of the media application 103, and as a result, the event machine learning module 204 can provide the best classification in real time.
[0061] The scoring module 206 generates an event importance score for each episode. In some embodiments, the scoring module 206 includes a set of instructions that can be executed by the processor 235 to score an episode and determine one or more events from the episode. In some embodiments, the scoring module 206 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0062] In some embodiments, the scoring module 206 generates an event importance score for each episode. The scoring module 206 can generate an event importance score based on data from all users who are permitted to use user data. This may apply if the number of media items in a user's library falls below a threshold. In some embodiments, the scoring module 206 generates an event importance score based on specific data for a user, for example, if an image contains the user or a person important to that user (e.g., a person appearing in a threshold number of media items related to the user), if the episode takes place in a place important to the user (e.g., a frequently visited or familiar place) or a rare place, previous user reactions (sharing media items, printing media items, indicating approval of media items (e.g., cherishing a media item), editing media items, viewing media items), or predictions of the user's reaction to the episode. For example, if a user consistently views images of a particular person, a birthday party, and Christmas, the scoring module 206 can apply a score to similar images indicating that these images are meaningful to the user. In another example, the scoring module 206 generates an event importance score based on the similarity between an episode and an event deemed important.
[0063] In some embodiments, the scoring module 206 further determines the event importance score based on a threshold number of media items in the corresponding episode. For example, an episode may need to have at least five media items (or seven, or ten, etc.) to be recognized as an event. In some embodiments, the scoring module 206 further determines the event importance score based on a larger number of media items in the corresponding episode, at least one face cluster of a threshold rank in the corresponding episode, or the presence of rare face clusters. In some embodiments, the scoring module 206 further determines the event importance score based on quality metrics of the media items in the corresponding episode. Quality metrics may be based on, for example, whether the media items are blurry, oversaturated, or have out-of-focus objects.
[0064] In some embodiments, the scoring module 206 merges multiple episodes into a single event based on multiple episodes occurring within a 24-hour period (e.g., many birthday celebrations). A single event includes events occurring over multiple days (e.g., Hanukkah, Diwali, Ramadan, Cherry Blossom Festival), celebrations on different days related to a single event (e.g., Christmas including an Indian wedding, tree lighting, family dinner, and gift opening), multiple episodes involving the same location (e.g., work treats where meals are taken at the same location), or multiple episodes involving the same set of people identified by the same set of face clusters. In some embodiments, since an event includes events of even longer durations (e.g., Ramadan is 30 days, Cherry Blossom Festival is 21 days), the scoring module 206 merges the events into one that is up to 31 days long.
[0065] In some embodiments, the scoring module 206 receives an event type (e.g., birthday) from the event machine learning module 204 and merges multiple events within a predetermined period based on the event type (e.g., merge all birthday-related celebrations that occurred within 5 days before or after the birthday). For example, a birthday may include not only a party but also related events that occurred around the birthday, such as dinner.
[0066] In some embodiments, the scoring module 206 receives event types from the event machine learning module 204 and generates a confidence score indicating the likelihood that the corresponding event is an accurately identified event type. For example, the scoring module 206 receives the same information regarding optical character recognition, object recognition, and metadata as described above with reference to the event machine learning module 204, and can generate a confidence score for the event type based on text within media items, objects identified from media items corresponding to the event type (e.g., a Christmas Santa hat, Valentine's Day flowers, and heart-shaped chocolates), and the date of a particular event corresponding to a known holiday (e.g., Christmas, Valentine's Day). The confidence score may be numerical (1, 0.4, 500), a percentage (5%, 44%, etc.), or a different metric.
[0067] In some embodiments, the scoring module 206 outputs an event importance score using a scoring machine learning model. In this example, the scoring machine learning model can receive user behavior as input as part of a training set for training the scoring machine learning model. Once the scoring machine learning model is trained, it can receive episodes as input and output an event importance score.
[0068] In some embodiments, the scoring module 206 determines one or more events from an episode based on event signals received by the event machine learning module 204 and corresponding event importance scores that exceed an event importance threshold. The scoring module 206 can instruct the user interface module 210 to display events with corresponding event importance scores that exceed an event importance threshold in a grid.
[0069] In some embodiments, events may be generated periodically. For example, one or more events may be limited to a predetermined number (e.g., seven) each month to avoid overloading the user with different events. For example, if another user sharing media with a first user receives new media, the first user can associate the new media with the media library. The scoring module 206 generates an event importance score for the episode containing the new media, and if the event importance score exceeds an event importance threshold, the scoring module 206 determines that the event importance score is greater than the event importance score for a particular event displayed in the user interface, and the scoring module 206 can instruct the user interface 210 to replace the previous event with the new event.
[0070] The titling module 208 generates a title for an event. In some embodiments, the titling module 208 includes a set of instructions that can be executed by the processor 235 to generate the title. In some embodiments, the titling module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0071] In some embodiments, the titling module 208 receives an event signal from the event machine learning module 204 indicating the possibility that an event occurred in each episode. The titling module 208 can also receive an identification of the event type from the event machine learning module 204. Using the event type, the titling module 208 can determine an event title that describes the event type.
[0072] In some embodiments, the titling module 208 receives a confidence score from the scoring module 206 indicating the likelihood that a corresponding event is of an accurately recognized event type. If the confidence score meets a confidence threshold, the titling module 208 adds a title to the corresponding event based on the event type. For example, Figure 3 includes a first example 600 of an image of a man holding flowers and two people embracing. In this case, since the confidence score meets the threshold, the titling module 208 applies the title "Valentine's Day" to the corresponding event. This could happen, for example, if the scoring module 206 determines, based on the location of the image, which is a restaurant, the date the image was taken, and the man holding the flowers, that the image was taken as part of a Valentine's Day outing. In another example, the titling module 206 uses a high-confidence title if it determines that more than 50% (or 40%, 90%, etc.) of the media items in an event have at least 90% (or 85%, etc.) confidence. In some embodiments, if an event has a high confidence value and occurs only once a year, the titling module 208 adds the year to the title (e.g., Diwali 2019).
[0073] If the confidence score does not meet the confidence threshold, the titling module 208 adds a title to the corresponding event (for example, "Celebrating Love and Laughter" for an event that occurs immediately after the user's birthday) based on a template expression that suggests an event type that is not clear enough to be interpreted as a mistake. Continuing the above example, Figure 3 includes a second example 350. In this case, the confidence score does not meet the threshold, and although the image was taken on Valentine's Day, it is at a campsite and the woman is not holding flowers, so the titling module 208 applies the template expression "Anniversary of Love". The scoring module 206 can determine the confidence score based on the sum of media items for the corresponding event.
[0074] The title machine learning module 258 generates a title for an event. In some embodiments, the title machine learning module 258 includes a set of instructions that can be executed by the processor 235 to generate the title. In some embodiments, the title machine learning module 258 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0075] In some embodiments, the title machine learning module 258 can use training data to generate a trained model, specifically a title machine learning model. For example, the training data may include any type of data, such as media organized as events (e.g., images, videos, etc.), titles of the events, and corresponding features (e.g., labels or tags associated with each media item to identify that an event occurred, the type of event, objects within the media item, etc.).
[0076] Training data may be obtained from any source, for example, a data repository specifically designated for training, or data for which permission has been granted to be used as training data for machine learning. In embodiments where one or more users permit the use of their respective user data to train a machine learning model, the training data may include this user data. In some embodiments, user data may include images / videos or image / video metadata (e.g., images, corresponding features obtained from users providing manual titles, tags, labels, geotags identifying the location of the images, timestamps related to the date of the event), information (e.g., chat data such as messages on social networks, emails, text messages, audio, video), documents (e.g., spreadsheets, text documents, presentations), etc.
[0077] In some embodiments, the training data may include synthetic data generated for training purposes, such as data not based on user input or activity in the situation being trained, such as data generated from simulations or computer-generated images / videos. In some embodiments, the Title Machine Learning Module 258 uses weights obtained from another application and not edited / transferred. For example, in these embodiments, the trained model may be generated on a different device, for example, and provided as part of the media application 103. In various embodiments, the trained model may be provided as a data file containing the model structure or morphology (defining, for example, the number and types of neural network nodes, the connections between nodes, and organizing the nodes into multiple layers) and the associated weights. The Title Machine Learning Module 258 can read the data file of the trained model and implement a neural network, including node connections, layers, and weights, based on the model structure or morphology specified in the trained model.
[0078] The title machine learning module 258 generates a trained model, referred to herein as the title machine learning model. In some embodiments, the title machine learning module 258 is configured to identify one or more features within input media items and generate (embedded) feature vectors representing the media items by applying the event machine learning model to data such as application data 267 (e.g., input media).
[0079] In some embodiments, the media application 103 includes either a title generation module 208 or a title machine learning module 258. In some embodiments, the title generation module 208 and the title machine learning module 258 are the same module. In some embodiments, the title machine learning module 258 is stored in the media server, and the remaining modules are stored in the user device 115.
[0080] In some implementations, metadata associated with media items may be provided as additional input to the gating model if permitted by the user. The metadata may include the location and / or time the video was taken, whether the video was shared via a social network, image sharing application, messaging application, etc., depth information associated with one or more video frames, sensor values from one or more sensors of the camera that took the video (e.g., accelerometer, gyroscope, light sensor, or other sensors), and user permission factors such as the user's ID (if the user consents). For example, if a video is taken outdoors at night with the camera pointed upwards, the metadata may indicate that the camera was pointed towards the sky when the video was taken, and therefore related to an astronomical event.
[0081] In some embodiments, the title machine learning module 258 may include software code executed by the processor 235. In some embodiments, the title machine learning module 258 may specify a circuit configuration (e.g., a programmable processor, a field-programmable gate array (FPGA)) that enables the processor 285 to apply a title machine learning model. In some embodiments, the title machine learning module 258 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the title machine learning module 258 may provide an application programming interface (API). The operating system 263 and / or other applications 265 can use this API to call the title machine learning module 258 and determine one or more features of an input image by, for example, applying an event machine learning model to application data 267.
[0082] In some embodiments, the title machine learning model may include one or more model forms or structures. In some embodiments, the title machine learning model may use a support vector machine, but in some embodiments, a convolutional neural network (CNN) is preferred. For example, the model form or structure may include any type of neural network, e.g., a linear network, a deep neural network implementing multiple layers (e.g., a "hidden layer" between the input and output layers, each of which is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results obtained from processing each tile), or a sequence-to-sequence neural network (e.g., a network that takes sequential data as input, such as words in a sentence or frames in a video, and produces a result sequence as output).
[0083] The model configuration or structure can specify the connections between various nodes and the organization of nodes into layers. For example, the nodes in the first layer (e.g., the input layer) can receive data as input data or application data. For example, if a title machine learning model is used to analyze input images related to user accounts, e.g., the first image, such data may include, for example, one or more pixels per node. Subsequent intermediate layers can receive the outputs of the nodes in the previous layer as input, according to the connections specified in the model configuration or structure. These layers are sometimes called hidden layers. The final layer (e.g., the output layer) generates the output of the machine learning application. For example, this output may be a title related to the input image. In some embodiments, the model configuration or structure specifies the number and / or types of nodes in each layer.
[0084] In some embodiments, the model architecture is a CNN having network layers, each network layer extracting image features at a different level of abstraction. The CNN used to identify features in an image may also be used to classify the image. The model architecture may include combinations and sequences of layers consisting of multidimensional convolution, mean pooling, max pooling, activation functions, normalization, regularization, and other layers and modules actually used in applied deep neural networks.
[0085] In different embodiments, the title machine learning model may include one or more models. One or more models may include multiple nodes arranged in layers according to a delta structure or morphology, or, in the case of a CNN, a filter bank. In some embodiments, a node may be a memoryless computation node configured to process one unit of input and produce one unit of output. The computation performed by the node may include, for example, the steps of multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and producing a node output by adjusting the weighted sum with a bias value or intercept value. Different layers may include different types of inputs related to media items, such as a first layer for metadata, a second layer for optical character recognition, and a third layer for object recognition.
[0086] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to a tuned weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multicore processor, using individual processing units of a graphical processing unit (GPU), or using dedicated neural circuits. In some embodiments, nodes may include memory. A node may, for example, remember one or more previous inputs and use one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long-short-term memory (LSTM) node. An LSTM node can use memory to maintain state that allows the node to behave like a finite state machine (FSM). Models including such nodes would be useful when processing sequential data, such as multiple words in a sentence or paragraph, a series of images, frames in a video, conversation, or other audio. For example, a heuristics-based model used in a gating model can remember one or more previously generated features for a previous image.
[0087] In some embodiments, the title machine learning model may include embeddings or weights for individual nodes. For example, the title machine learning model may be initialized as a group of nodes organized into layers as specified by the model morphology or structure. During initialization, each weight may be applied to the connections between each pair of nodes connected according to the model morphology, for example, between each pair of nodes in a continuous layer of a neural network. For example, each weight may be assigned randomly or initialized to a default value. The title machine learning model can then be trained, for example, with a training set of digital images to produce results. In some embodiments, a subset of the entire architecture may be reused from other machine learning applications as a transfer learning method to leverage the pre-trained weights.
[0088] For example, training may involve applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., media organized as events) and expected outputs corresponding to each input (e.g., titles for each event). For example, the weight values may be automatically adjusted based on a comparison between the output of the title machine learning model and the expected output so that the title machine learning model increases the probability of generating the expected output for the appropriate tile.
[0089] In some embodiments, training may involve applying unsupervised learning techniques. For example, only input data (e.g., media organized as events and titles) may be provided, and the title machine learning model may be trained to distinguish the data, for example, to identify events associated with different titles. In another example, the title machine learning model may be trained to generate titles by distinguishing media items within events.
[0090] In various embodiments, the trained model includes a set of weights corresponding to the model structure. In embodiments where the training set of digital images is omitted, the title machine learning module 258 may generate a title machine learning model based on prior training by, for example, the developer of the title machine learning module 258 or a third party. In some embodiments, the title machine learning model may include a set of fixed weights downloaded from a server that provides the weights.
[0091] In some embodiments, the title machine learning module 258 may be implemented offline. In some embodiments, minor updates to the title machine learning model may be implemented online. In such embodiments, an application calling the title machine learning module 258 (e.g., the operating system 263, one or more other applications 265) can utilize the feature detections generated by the title machine learning module 258 and generate a system log (e.g., actions taken by the user based on the feature detections, if permitted by the user, or the results of further processing, if used as input for further processing). The system log may be generated periodically, for example, every hour, every month, or every three months, and may be used to update the title machine learning model, for example, to update the embedding of the title machine learning model, if permitted by the user.
[0092] In some embodiments, the Title Machine Learning Module 258 may be implemented in a manner that conforms to a specific configuration of the computing device 200 on which the Title Machine Learning Module 258 is executed. For example, the Title Machine Learning Module 258 can determine the computation graph that utilizes available computing resources, such as the processor 285. If the Title Machine Learning Module 258 is implemented as a distributed application across multiple devices, the Title Machine Learning Module 258 can determine which computations are performed on each device to optimize the computation. In another example, if the Title Machine Learning Module 258 determines that the processor 235 contains a GPU with a certain number (e.g., 1000) GPU cores, the Title Machine Learning Module 258 can be implemented (e.g., as 1000 separate processes or threads).
[0093] In some embodiments, the title machine learning module 258 can implement a set of trained models. For example, the title machine learning module may include multiple trained models, each applicable to the same input data. In these embodiments, the title machine learning module 258 can select a particular trained model based, for example, on available computational resources, the success rate when using previous inferences, etc.
[0094] In some embodiments, the Title Machine Learning Module 258 can run multiple trained models. In these embodiments, the Title Machine Learning Module 258 can synthesize outputs, for example, by using a majority vote to score the outputs obtained by applying each trained model, or by selecting one or more specific outputs. In some embodiments, such selectors are part of the model itself and function as a connecting layer between the trained models. Furthermore, in these embodiments, the Title Machine Learning Module 258 can apply a time threshold (e.g., 0.5 ms) for applying each trained model and utilize only the individual outputs available within the time threshold. Outputs not received within the time threshold are not utilized and may be discarded, for example. For example, such a technique would be appropriate when there is a specified time limit between calling the Title Machine Learning Module 258, for example, by the operating system 263 or one or more other applications 265.
[0095] In some embodiments, the title machine learning module 258 receives edits to the title from the user. The title machine learning module 258 incorporates the feedback into the title machine learning model and improves the title output by changing parameters.
[0096] The user interface module 210 generates the user interface. In some embodiments, the user interface module 210 includes a set of instructions that can be executed by the processor 235 to generate the user interface. In some embodiments, the user interface module 210 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0097] In some embodiments, the user interface module 210 provides a user interface that includes event items having corresponding media from a particular event among one or more events determined by the scoring module 206. In the user interface, a particular event may be more prominently characterized than other items. For example, Figure 4 includes a first example 400 having a media item from March and an event item titled "Happy Couple".
[0098] The user interface module 210 may be part of a reverse time-series grid of media containing the corresponding media, or it may be a time-series grid, or it may not have an organization based on the capture time of the media items. Continuing the above example, Figure 4 shows a first example 400 of a reverse time-series grid 405 containing media items from March 10, an event item titled "Happy Couple," and media items from March 9. In some embodiments, the event with the highest score is selected by the user interface module 210 to be displayed in the reverse time-series grid. In some embodiments, a certain amount of time (e.g., one week, two days, etc.) must have elapsed for the event to be displayed in the reverse time-series grid.
[0099] In some embodiments, the user interface module 210 can generate a user interface that includes events with different significance levels. For example, the first example 400 in Figure 4 is a larger grid tile where events are displayed as large tiles scattered across an inverse time-series grid with a title header. In some embodiments, the user interface module 210 determines whether to display an event in a larger grid tile or a smaller grid tile based on an event importance score, as will be described later.
[0100] Figure 4 also includes a second example 425 with smaller grid tiles. In this case, the first media item from a particular event is less prominently featured, the second media item from a particular event is more prominently featured, and the events are scattered across an inverse time-series grid. In some embodiments, events within the user interface are displayed below the date header corresponding to the event.
[0101] Figure 4 further includes a third example 450. In this case, a summary of the monthly highlights and a list of images from the same month are displayed in the lower section 455 of the user interface, with specific events being characterized in the least prominent way. All events displayed inline in the reverse time-series grid are included in the monthly highlights carousel in the lower section 455.
[0102] In some embodiments, the user interface module 210 displays the event in the user interface after a predetermined time (e.g., 21 days) has elapsed. In some embodiments, where the event spans multiple days across two different months, the user interface module 210 displays the event in the monthly highlight carousel section based on the month in which the event ended. For example, if the event occurred between October 2 and November 1, the event would only be shown in the November highlight carousel. In some embodiments, if multiple events occur on a single day, the user interface module 210 selects and displays the event for that day that has the best event importance score.
[0103] In some embodiments, when a user selects an event item from a grid, the user interface module 210 replaces the grid with a display of corresponding media items from a specific event within a predetermined period. For example, if a user selects the event item "Happy Couple" from Figure 4, the user interface module 210 displays a video of the proposal, an image of the ring, an image of the park where the proposal took place, and so on. The user interface module 210 can display each media item from a specific event at predetermined intervals (e.g., 2 seconds, 5 seconds, etc.). For example, the user interface module 210 can use the entire user interface to display each media item over 2 seconds.
[0104] In some embodiments, the user interface includes elements that allow the user to navigate back and forth through media items at their own pace, and elements that enable print sharing and ordering actions, and actions to mark items as favorites. Once the viewing of a media item is complete (or the user exits the viewing of a corresponding media item), the user interface module 210 can display the grid again.
[0105] In some embodiments, corresponding media from one or more events are displayed in a size that depends on one or more of the following: the event importance score, the number of media items for one or more events, the total number of events within a single period, or the event type (e.g., my wedding vs. a friend's wedding). For example, in Example 400 of Figure 4, the event item "Happy Couple" is displayed in a larger size based on the event importance score indicating a top ranking.
[0106] In some embodiments, the user interface module 210 receives a request from the user to remove a depiction of a person from media in the media library that contains such depictions. For example, the user interface may allow the user to remove an image from the media library by clicking on the image, the user interface may be able to confirm that the image has been removed, and the user interface may provide a page to request feedback on why the user requested the removal of the depiction of a person. The feedback may be used by the user interface module 210 as input to an event machine learning model and / or to modify the event importance score. In some embodiments, the scoring module 206 filters out depictions of people from the media before generating the event importance score. This can favorably avoid the user interface module 210 displaying media related to events that have too few media items after the depictions of people have been removed from the media, or favorably avoid displaying events with or featuring hidden depictions of people.
[0107] The user interface may include options for curating media grids and libraries. In some embodiments, the user interface may include options for hiding or adding one or more media items from one or more events, dates associated with one or more events, or people or pets depicted within media items from one or more events. Figure 5 includes an exemplary user interface 500 that provides the user with options for hiding people and pets or for hiding dates from media associated with events. The exemplary user interface 500 also includes options for selecting memories to place in a reverse chronological grid, and options for managing whether or not to provide notifications associated with the memories, such as daily reminders, silent notifications, etc.
[0108] In some embodiments, the user interface includes an option to delete an event item. By deleting an event item, the event item will no longer be visible to the user in the future, but the corresponding media item remains in the media library associated with the user account. In some embodiments, the user interface includes an option to delete a media item from an event item. Deleting a media item from an event item does not delete the media item from the media library associated with the user account. Figure 6 shows an exemplary user interface. In this case, the user interface includes a request 600 for confirmation, a confirmation 625 that the media item has been deleted, and a feedback screen 650 that requests feedback from the user about why they want to delete the media item from the event item. Exemplary feedback options include the event item being sensitive, duplicate, off-topic, low quality, or otherwise.
[0109] In some embodiments, the user interface module 210 sends feedback to the event machine learning module 204, which uses the feedback to modify the parameters of the event machine learning model. In some embodiments, the user interface module 210 sends feedback to the scoring module 206, which modifies the event importance score in response to this feedback.
[0110] In some embodiments, the user interface includes options for modifying event details. For example, the user interface includes options for editing corresponding media from one or more events. In yet another example, the user interface includes options for changing the titles of one or more events while one or more events are displayed in the user interface. In some embodiments, the titling module 208 uses feedback to improve title generation.
[0111] Figure 7 shows an exemplary user interface 700 that, according to several embodiments, has options for editing titles, deleting events, resizing events in a grid, and changing the importance of events in a grid. In this example, the user can access these options by right-clicking on an event in the user interface 700 or by some other mechanism. In some embodiments, these edits are available while the event is displayed in the user interface 700 or while media items within the event are displayed.
[0112] Selecting "Edit Title" changes the title of User Interface 700 to something different. Selecting "Delete Memory" leaves media items associated with the event in the user's associated library, but the event will not be displayed in the User Interface in the future. Selecting "Normal Size" displays the event in User Interface 700 at a normal size, like Example 425 in Figure 4, rather than the larger size of Example 400 in Figure 4 or the smaller size of Example 450 in Figure 4. Selecting "Spotlight" displays the event in User Interface 700 at a larger size, like Example 400 in Figure 4.
[0113] In some embodiments, the user interface module 210 generates audio in a corresponding media based on one or more event types. For example, in the case of a happy event such as a graduation ceremony or a wedding, the music may be happy music. In the case of a more serious event such as a funeral, the music may be more somber.
[0114] In some embodiments, if the scoring module 206 scores a different event higher than the edited event in response to the user editing an event's characteristics, such as the event's title, the user interface module 210 retains the event in the user interface. This is sometimes referred to as a “frozen” event. In some embodiments, the user interface module 210 retains the event in the user interface until a predetermined period (e.g., October) has ended. Conversely, in some embodiments, if the user deletes a media item from the library, the user interface module 210 removes the media item from the event.
[0115] In some embodiments, the user interface module 210 reorganizes different events within the user interface based on different changes. Referring to Figure 8A, exemplary block diagram 800 shows a different example in which the user interface grid reorganizes different events based on a change that limits the number of events to "7". The first example 805 includes a list of events displayed in a grid where the corresponding event severity scores range from 70 / 100 to 99 / 100.
[0116] When a user uploads a media item from a DSLR camera, the scoring module 206 generates a new event 810 with an event importance score of 81 / 100 that is higher than five of the events in the grid. As a result, the event with the lowest event score (i.e., 70 / 100) is removed from the grid and the new event is added. The user interface module 210 adds the new event 810 to the event list. This results in a second example 815 where the new event 810 is added and the event with the lowest event importance score of 70 / 100 is removed from the user interface. In the second example 815, the removed event is indicated by a line.
[0117] Continuing with the example in Figure 8A, the user edits the title of an event with a score of 75 / 100. This means the event is frozen and retained in the grid. The user then uploads a media item from a DSLR camera, and the scoring module 206 generates a new event 825 with a score of 77 / 100. Since the new event 825 has a higher score of 77 / 100 than the previous event, the second event would normally be removed from the grid. However, because the second event is frozen, it remains in the grid. As a result, Figure 8B shows block diagram 850 with a third example 835 in which the grid contains eight events.
[0118] Continuing with the example in Figure 8B, when the user uploads a media item from a DSLR camera, the scoring module 206 generates a new event with a score of 88 / 100. Since the new event 840 has a higher score than the previous event's score of 77 / 100, and the event with a score of 75 / 100 is frozen, the previous event with a score of 77 / 100 is removed from the grid. The fourth example 845 shows the remaining events in the grid, along with the events with event importance scores of 77 / 100 and 70 / 100 that have been removed from the grid (as indicated by the lines).
[0119] In some embodiments, the user interface module 210 generates a user interface for providing the user with memory events. For example, the user interface module 210 can generate an icon that appears at the top of the user's screen. When this icon is selected, it causes the user interface to display media items related to the event. In some embodiments, the user interface module 210 selects events that occurred at a predetermined time (e.g., one year ago), events with the highest event importance score, or random events, etc.
[0120] Exemplary method 900 Figure 9 is a flowchart illustrating an exemplary method 900 for displaying event items according to several embodiments. Method 900 may be performed by a computing device 200 in Figure 2A or Figure 2B, for example, a user device 115 or media server 101 shown in Figure 1.
[0121] Method 900 can begin with block 902. In block 902, the media library associated with the user account is segmented into episodes. Each episode is associated with a corresponding period. Block 904 can be executed after block 902.
[0122] In block 904, the event machine learning model generates event signals indicating the likelihood that an event occurred in each episode. This event machine learning model is a classifier that receives media as input. Block 906 can be executed after block 904.
[0123] In block 906, an event importance score is generated for each episode. In some embodiments, the event importance score is generated based on one or more of the following: at least a threshold number of media items in the corresponding episode, at least a threshold number of face clusters in the corresponding episode, a quality metric for media items in the corresponding episode, at least one face cluster of threshold rank, or the presence of rare face clusters. Block 908 can be executed after block 906.
[0124] In block 908, one or more events are determined from the episode based on the event signal and the corresponding event importance score that exceeds the event importance threshold. Block 910 can be executed after block 908.
[0125] In block 910, a user interface is provided that includes event items having corresponding media from a specific event among one or more events. In some embodiments, the user interface is part of a time-series grid of media containing the corresponding media. When a user selects an event item from the time-series grid, the time-series grid may be replaced with a display of corresponding media items from the specific event over a predetermined period. In response to the completion of the display of the corresponding media items, the time-series grid is displayed.
[0126] In addition to the above description, the system, program, or function described herein may give the user control over whether and when it is possible to collect user information (e.g., information about the user's media items such as photographs or videos, the user's social networks, social behavior or activities, occupation, viewing preferences for image-based creations, settings for hiding people or pets, user interface preferences, or information about the user's current location) and whether content or information is transmitted from the server. Furthermore, before storing or using certain data, it may be processed to remove personally identifiable information in one or more ways. For example, the user's ID may be processed so that the user's personal information cannot be identified. Also, when obtaining location information (e.g., city, zip code, or state level) so that the user's location cannot be identified, the user's geographic location may be generalized. Thus, the user can control what user information is collected, how the information is used, and what information is provided to the user.
[0127] In the above description, for the purpose of explanation, many specific details are provided to give a complete understanding of the various embodiments described. However, it will be apparent to those skilled in the art that the various embodiments described can be carried out even without these specific details. In some cases, structures and devices are shown in block diagrams to avoid obscuring the description. For example, embodiments can be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.
[0128] In this specification, any reference to “some embodiments” or “some instances” means that certain features, structures, or characteristics described in relation to an embodiment or instance may be included in at least one implementation of the description. The phrase “in some embodiments” in various places in this specification does not necessarily refer to the same embodiment.
[0129] Some parts of the detailed description above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic descriptions and representations are means used by those skilled in the field of data processing to communicate the nature of their work to others skilled in the field in the most effective way. In this specification, an algorithm is generally considered to be a set of consistent steps that produce a desired result. These steps require the physical manipulation of physical quantities. These quantities usually take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and other manipulated, though not always. In some cases, and primarily for reasons of common use, it is convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.
[0130] It is important to understand that all these and similar terms are merely convenient labels associated with and applied to appropriate physical quantities. Unless otherwise specified or as evident from the discussion, any discussion using terms including “process,” “operate,” “calculate,” “determine,” or “display” throughout the description refers to the operation and processes of a computer system or similar electronic computing device that processes and transforms data represented as physical quantities in computer system memory, registers, or other information storage devices, transmission devices, or display devices.
[0131] Embodiments of this specification also relate to processors for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively started or reconfigured by a computer program stored in the computer. Such computer programs may be stored in non-temporary computer-readable storage media, including, but not limited to, optical discs, ROMs, CD-ROMs, magnetic disks, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory including USB keys with non-volatile memory, or any type of medium suitable for storing electronic instructions, each of which is connected to a computer system bus.
[0132] This specification may include several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, and microcode.
[0133] Furthermore, the description may take the form of a computer program product accessible from a computer-enabled or computer-readable medium that provides program code used by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-enabled or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program used by or in connection with an instruction execution system, machine, or apparatus.
[0134] A data processing system suitable for storing or executing program code includes at least one processor directly or indirectly connected to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, mass storage, and cache memory providing temporary storage for at least some program code to reduce the number of times the code must be retrieved from mass storage during execution.
Claims
1. A method implemented by a computer, The method includes segmenting the media library associated with a user account into media sets by time period, each media set containing one or more media related to the corresponding time period, and the method further includes: The method includes generating event signals indicating the likelihood of an event occurring in each media set using an event machine learning model, wherein the event machine learning model is a classifier that receives media in the media library as input, and the method further includes The method includes generating an event importance score for each media set using a scoring machine learning model, wherein the scoring machine learning model is trained to receive each media set as input and output the event importance score, and the method further includes: Based on the event signal and the corresponding event importance score that exceeds the event importance threshold, one or more events are determined from the media set. A method comprising providing a user interface that includes an event item having media corresponding to a specific event among the one or more events.
2. The user interface is part of an inverse time-series grid of media, including the corresponding media. The aforementioned method, In response to a user selecting an event item from the inverse time series grid, the inverse time series grid is replaced with a display of the corresponding media from the specific event during a predetermined period of time. The method according to claim 1, further comprising displaying the inverse time-series grid in response to the completion of the display of the corresponding media.
3. Receiving a request from a user to delete the depiction of a person from the media in the media library that contains the depiction of a person, The method according to claim 1 or 2, further comprising filtering the description of the person from the media, wherein the filtering is performed before generating the event importance score.
4. The method according to any one of claims 1 to 3, wherein the user interface includes an option to hide or add one or more media items from the one or more events, dates associated with the one or more events, or people or pets depicted in the media items from the one or more events.
5. To optimize the calculations, determine which calculations are performed on each individual device, The method according to any one of claims 1 to 4, further comprising implementing the event machine learning model on a plurality of devices based on calculations performed on the individual devices.
6. The process further includes merging multiple media sets into a single event based on one or more media sets each associated with multiple periods, wherein the multiple periods as a whole comprise a 24-hour period. The method according to any one of claims 1 to 5, wherein the single event is a type of event that occurs over multiple days, celebrations on different days related to the single event, multiple media sets related to the same location, or multiple media sets related to the same set of face clusters.
7. To generate a confidence score indicating the likelihood that the corresponding event is of the correctly identified event type, The method according to any one of claims 1 to 6, further comprising adding an automatically generated title describing the type of event to the corresponding event in response to the confidence score satisfying a confidence threshold.
8. The method according to claim 7, further comprising adding the automatically generated title based on a template representation to the corresponding event in response to the confidence score not meeting the confidence threshold.
9. The title machine learning model receives the corresponding media as input from one or more events, The method according to any one of claims 1 to 8, wherein the title machine learning model generates a title as an output.
10. The method according to any one of claims 1 to 9, wherein the user interface includes an option for editing the corresponding media from the specific event.
11. Determining one or more events includes determining the events such that the number of events is less than or equal to a predetermined number each month. The aforementioned method, Receiving new media related to the aforementioned media library, The method according to any one of claims 1 to 10, further comprising replacing one or more of the specific events with the new event in response to the new event being associated with a new event importance score higher than the event importance score of the specific event.
12. The method according to any one of claims 1 to 11, wherein the user interface includes an option for changing the title of the event item.
13. The method according to any one of claims 1 to 12, further comprising generating audio for the corresponding media based on the type of the specific event.
14. A computer program for causing one or more computers to perform the method described in any one of claims 1 to 13.
15. Processor and A computing device comprising a memory connected to the processor and storing the computer program described in claim 14.
16. The computing device according to claim 15, wherein the event item is displayed in size based on one or more of the following: the event importance score, the number of media items for the specific event, the total number of events during a single period, or the event type.