Machine learning-based image compression settings that reflect user preferences
Through feature detection and user-specific machine learning models, personalized compression settings are determined based on user preferences and image features, solving the quality degradation problem caused by image compression, improving user experience and optimizing storage space utilization.
Patent Information
- Application Number
- CN202080009025.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-08-17
AI Technical Summary
The image compression methods in the prior art lead to degradation of image quality and poor user experience. In particular, compression loss of important images for users is difficult to control.
It uses feature detection machine learning models and user-specific machine learning models to determine personalized compression settings based on user preferences and image features, prioritizing the quality of users' important images and reducing storage space.
Through personalized image compression settings, the user experience is improved, storage space is effectively utilized, and the high quality of important images is maintained.
Smart Images

Figure CN114080615B_ABST
Abstract
Description
Background Art
[0001] As smartphones and other portable cameras grow in popularity, users are capturing more and more images. However, on-device storage, as well as cloud or server storage, is a limited resource. Image compression is an effective way to reduce the amount of storage space required to store images. However, lossy compression can significantly degrade the quality of compressed images, resulting in a poor user experience.
[0002] The purpose of the background description provided herein is generally to present the context of the present disclosure. To the extent described in this background section, the work currently claimed as the inventor's work, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, is not admitted, either explicitly or implicitly, as prior art to the present disclosure. Summary of the Invention
[0003] Embodiments described herein relate to methods, devices, and computer-readable media for generating compression settings. The method may include obtaining an input image, the input image associated with a user account; determining one or more features of the input image using a feature detection machine learning model; determining a compression setting for the input image based on the one or more features in the input image using a user-specific machine learning model personalized for the user account; and compressing the input image based on the compression setting.
[0004] In some embodiments, a feature detection machine learning model is generated by obtaining a training set of digital images and corresponding features and training the feature detection machine learning model based on the training set and the corresponding features, wherein, after training, the feature detection machine learning model is capable of identifying image features in an input image provided to the feature detection machine learning model. In some embodiments, the feature detection machine learning model comprises a convolutional neural network (CNN) having a plurality of network layers, wherein each network layer extracts one or more image features at a different level of abstraction. In some embodiments, a user-specific machine learning model is generated by obtaining a training set of user-specific features associated with a user, the user-specific features indicating user actions with respect to one or more prior images, and training the user-specific machine learning model based on the user-specific features and the one or more prior images, wherein, after training, the user-specific machine learning model determines a rank of an input image provided to the user-specific machine learning model. In some embodiments, the corresponding image features of the one or more prior images are obtained by applying the feature detection machine learning model to the one or more prior images.
[0005] In some embodiments, the method may further include: providing a first user interface having two or more versions of a sample image to a user associated with the user account, each version compressed with a different compression setting; obtaining user input from the user identifying a specific version of the sample image; and selecting the compression setting associated with the specific version of the sample image as a baseline compression setting for the user account. In some embodiments, determining the compression setting includes: determining a rating of the input image using a user-specific machine learning model and mapping the rating to the compression setting, wherein the mapping is based on the baseline compression setting.
[0006] In some embodiments, the method may further include: determining that the rating of the input image satisfies an importance threshold and, in response to determining that the rating satisfies the importance threshold, performing one or more of the following: providing a suggestion to the user to share the input image, prioritizing backup of the input image over backups of other images associated with the user account that do not satisfy the importance threshold, or providing a second user interface including instructions for capturing a subsequent image if the scene depicted in the subsequent image has at least one of the one or more features of the input image.
[0007] Some embodiments may include a computing device comprising a processor and a memory having instructions stored thereon, the instructions, when executed by the processor, causing the processor to perform operations comprising: obtaining an input image, the input image being associated with a user account; determining one or more features of the input image using a feature detection machine learning model; determining a compression setting for the input image based on one or more features in the input image using a user-specific machine learning model personalized for the user account; and compressing the input image based on the compression setting.
[0008] In some embodiments, the feature detection machine learning model is generated by obtaining a training set of digital images and corresponding features and training the feature detection machine learning model based on the training set and the corresponding features, wherein, after training, the feature detection machine learning model is able to identify image features in an input image provided to the feature detection machine learning model. In some embodiments, the user-specific machine learning model is generated by obtaining a training set of user-specific features associated with a user, the user-specific features indicating user actions with respect to one or more prior images, and training the user-specific machine learning model based on the user-specific features and the one or more prior images, wherein, after training, the user-specific machine learning model determines a rank of an input image provided to the user-specific machine learning model.
[0009] In some embodiments, the memory has further instructions stored thereon that, when executed by the processor, cause the processor to perform further operations, the operations comprising: providing a first user interface to a user associated with the user account using two or more versions of a sample image, each version compressed with a different compression setting; obtaining user input from the user identifying a particular version of the sample image; and selecting the compression setting associated with the particular version of the sample image as a baseline compression setting for the user account. In some embodiments, determining the compression setting comprises: determining a rating of the input image using a user-specific machine learning model and mapping the rating to the compression setting, wherein the mapping is based on the baseline compression setting.
[0010] Some embodiments may include a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations comprising: obtaining an input image, the input image being associated with a user account; determining one or more features of the input image using a feature detection machine learning model; determining a compression setting for the input image based on one or more features in the input image using a user-specific machine learning model personalized for the user account; and compressing the input image based on the compression setting.
[0011] In some embodiments, a feature detection machine learning model is generated by obtaining a training set of digital images and corresponding features and training the feature detection machine learning model based on the training set and the corresponding features, wherein after training, the feature detection machine learning model is able to identify image features in an input image provided to the feature detection machine learning model. In some embodiments, a user-specific machine learning model is generated by obtaining a training set of user-specific features associated with a user, the user-specific features indicating user actions with respect to one or more prior images, and training the user-specific machine learning model based on the user-specific features and the one or more prior images, wherein after training, the user-specific machine learning model determines a rank for an input image provided to the user-specific machine learning model. In some embodiments, the training set further includes corresponding image features for the one or more prior images.
[0012] In some embodiments, the operations further include: providing a first user interface having two or more versions of a sample image to a user associated with the user account, each version compressed with a different compression setting; obtaining user input from the user identifying a particular version of the sample image; and selecting the compression setting associated with the particular version of the sample image as a baseline compression setting for the user account. In some embodiments, determining the compression setting includes: determining a rating of the input image via a user-specific machine learning model and mapping the rating to the compression setting, wherein the mapping is based on the baseline compression setting. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a block diagram of an example network environment that can be used with one or more embodiments described herein.
[0014] Figure 2 is a block diagram of an example device that can be used with one or more embodiments described herein.
[0015] Figure 3 is a flow chart illustrating an example method of using a feature detection machine learning model to identify one or more features in an input image and using a user-specific machine learning model to determine compression settings from the one or more features of the image, in accordance with some embodiments.
[0016] Figure 4 is a flowchart illustrating an example method of creating a training model according to some embodiments.
[0017] Figure 5 is a flow chart illustrating an example method of applying a model to an input image according to some embodiments. DETAILED DESCRIPTION
[0018] A user captures an image using a camera (such as through a smartphone or other device). For example, the image may include a still image, a dynamic picture / motion picture, or an image frame from a video. The user may store the image on a client device or a server (e.g., a server providing an image hosting service). An application may be provided via the user's client device and / or a server that enables the user to manage images by, for example, viewing and / or editing the images; generating image-based creations such as slideshows, collages, etc.; sharing the images; posting the images to a social network or chat application where other users provide indications of approval for the image, such as by liking the image or coming on the image; and the like.
[0019] Storage space on client devices or servers is limited. One way to gain additional storage space without having to delete images is to use image compression to reduce the file size of images. However, lossy compression can result in a degradation of image quality, leading to a poor user experience when all images associated with a user undergo image compression.
[0020] When a user account has a large number of images, a subset of these images is likely to be images that the user particularly likes. For example, some users may only have strong feelings about the quality of their landscape photos because they like to print them out. Others may have strong feelings about high-resolution portraits of people because they like to share these portraits with family or may run a photography business. The image quality of receipts, screenshots, or other functional images (such as meme images, photos of business cards, newspaper articles, etc.) may not be as important. Therefore, the user's perception of the loss from image compression can depend on the type of image and the user account associated with the image. Therefore, it is helpful to identify which images may be important to the user. In some embodiments, the image management application generates and utilizes a feature detection machine learning model that identifies features in the first image. The image management application can also generate a user-specific machine learning model that is personalized for the user account. The image management application can use the user-specific machine learning model to determine compression settings for the first image. The image management application can compress the first image based on the compression settings, thereby freeing up storage space because the final compressed image has a smaller file size than the original image.
[0021] The figures use the same reference numerals to identify the same elements. A letter following a reference numeral (such as "103a") indicates that the text refers exclusively to the element having that particular reference numeral. A reference numeral in the text without a following letter (such as "103") refers to any or all elements in the figures bearing that reference numeral (e.g., "103" in the text refers to reference numerals "103a" and / or "103b" in the figures).
[0022] Sample network environment 100
[0023] Figure 1 1 illustrates a block diagram of an example network environment 100 that may be used in some embodiments described herein. In some embodiments, the network environment 100 includes one or more server systems, e.g., Figure 1 In the example of server system 101, for example, server system 101 can communicate with network 105. Server system 101 can include server device 104 and database 199 or other storage device. Database 199 can store one or more images and / or videos and metadata associated with the one or more images and / or videos. In some embodiments, server device 104 can provide image management application 103a. Image management application 103a can access images stored in database 199.
[0024] The network environment 100 may also include one or more client devices, such as client devices 115a, 115n, which may communicate with each other and / or with the server system 101 via the network 105. The network 105 may be any type of communication network, including one or more of the Internet, a local area network (LAN), a wireless network, a switch or hub connection, etc. In some embodiments, the network 105 may include point-to-point communications between devices, for example, using a point-to-point wireless protocol (e.g., Wi-Fi Direct, Ultrawideband, etc.) etc. An example of point-to-point communication between two client devices 115a and 115b is shown by arrow 132.
[0025] For ease of explanation, Figure 1 One block is shown for server system 101, server device 104, and database 199, and two blocks are shown for client devices 115a and 115n. Server blocks 101, 104, and 199 can represent multiple systems, server devices, and network databases, and the blocks can be provided in configurations different from those shown. For example, server system 101 can represent multiple server systems that can communicate with other server systems via network 105. In some embodiments, for example, server system 101 can include cloud hosting servers. In some examples, database 199 and / or other storage devices can be provided in a server system block that is separate from server device 104 and communicates with server device 104 and other server systems via network 105.
[0026] There may be any number of client devices 115. Each client device can be any type of electronic device, such as a desktop computer, a laptop computer, a portable or mobile device, a cellular phone, a smartphone, a tablet computer, a camera, a smart display, a television, a television set-top box or entertainment device, a wearable device (e.g., display glasses or goggles, a watch, headphones, an armband, jewelry, etc.), a personal digital assistant (PDA), a media player, a gaming device, etc. Some client devices may also include a local database or other storage device similar to database 199. In some embodiments, network environment 100 may not have all of the components shown and / or may have other elements, including other types of elements, instead of or in addition to the elements described herein.
[0027] In various embodiments, users 125 may communicate with server system 101 and / or with each other using respective client devices 115a, 115n. In some examples, users 125 may interact with each other via applications running on respective client devices and / or with server system 101 via web services implemented on server system 101, such as social networking services, image hosting services, or other types of web services. For example, respective client devices 115a, 115n may transfer data to and from one or more server systems (e.g., server system 101).
[0028] In some embodiments, the server system 101 can provide appropriate data to the client devices 115a, 115n so that each client device 115 can receive content uploaded to the server system 101 and / or the network service. In some examples, users 125 can interact through audio or video conferencing, audio, video or text chat, or other communication modes or applications. The network services implemented by the server system 101 may include systems that allow users 125 to perform various communications, form links and associations, upload and publish shared content (such as images, text, video, audio, and other types of content), and / or perform other functions. For example, the client device 115 can display received data, such as content published by the server and / or network service that is sent or streamed to the client device 115 and originates from a different client device 115 (or directly from a different client device 115) or originates from the server system 101 and / or the network service. In some embodiments, the client devices 115a, 115n can communicate directly with each other, for example, using the peer-to-peer communication between the client devices 115, 115n described above. In some embodiments, a “user” may include one or more programs or virtual entities as well as humans interfacing with the system or network 105 .
[0029] In some embodiments, any client device 115a, 115n can provide one or more applications. Figure 1 As shown, client device 115a may provide a camera application 152 and an image management application 103b. Client device 115n may also provide similar applications. Camera application 152 may provide a user 125a of a corresponding client device 115a with the ability to capture images using the camera of their corresponding client device 115a. For example, camera application 152 may be a software application executed on client device 115a.
[0030] In some embodiments, the camera application 152 may provide a user interface. For example, the user interface may enable a user of the client device 115a to select an image capture mode, such as a still image (or photo) mode, a burst mode (e.g., capturing a continuous number of images over a short period of time), a motion image mode, a video mode, a high dynamic range (HDR) mode, a resolution setting, and the like. For example, a video mode may correspond to the capture of a video comprising multiple frames and may be of any length. Further, the video mode may support different frame rates, such as 25 frames per second (fps), 30 fps, 50 fps, 60 fps, and the like. During the capture of an image or video, one or more parameters of the image capture may change. For example, when capturing a video, a user may use the client device 115a to zoom in or out of a scene.
[0031] In some embodiments, the camera application 152 may implement (eg, in part or in whole) the Figure 3 and Figure 4 In some embodiments, the image management application 103a and / or the image management application 103b may implement (eg, in part or in whole) the methods described herein. Figure 3 and Figure 4 The method described.
[0032] The camera application 152 and the image management application 103b can be implemented using the hardware and / or software of the client device 115a. In different embodiments, the image management application 103b can be a standalone application, for example, executed on any client device 115a, 115n, or can work in conjunction with the image management application 103a provided on the server system 101.
[0033] With user permission, the image management application 103 can perform one or more automatic functions, such as storing (e.g., backing up) the image or video (e.g., to the database 199 of the server system 101), enhancing the image or video, stabilizing the image or video, identifying one or more features in the image (e.g., a face, a body, type of object, type of movement), compressing the image, etc. In some examples, image or video stabilization can be performed based on input from an accelerometer, gyroscope, or other sensor of the client device 115a and / or based on comparison of multiple frames of a moving image or video.
[0034] The image management application 103 may also provide image management functionality, such as displaying images and / or videos in a user interface (e.g., in an exclusive view including a single image, in a grid view including multiple images, etc.), editing images or videos (e.g., adjusting image settings, applying filters, changing image focus, removing one or more frames of a motion image or video), sharing images with other users (e.g., of client devices 115a, 115n), archiving images (e.g., storing images so that they do not appear in a primary user interface), generating image-based creations (e.g., collages, albums, motion-based artifacts such as animations, stories, video loops, etc.), etc. In some embodiments, to generate image-based creations, the image management application 103 may utilize one or more tags associated with an image or video.
[0035] In some embodiments, the image management application 103 can determine one or more characteristics of the image and determine a compression setting for the image based on the one or more characteristics in the image. In some embodiments, the image management application 103 can store the compression setting associated with the image or video and the compressed image or video in the database 199 and / or in a local database (not shown) on the client device 115. In some embodiments, the image management application 103 deletes the original image immediately, saves the original image and asks the user to confirm the deletion, or saves the original image for a certain number of days before deleting the original image.
[0036] The user interface on the client device 115 can enable the display of user content and other content, including images, videos, data, and other content, as well as communications, privacy settings, notifications, and other data. Such a user interface can be displayed using software on the client device 115, software on the server device 104, and / or a combination of client software and server software (e.g., application software or client software that communicates with the server system 101) executed on the server device 104. The user interface can be displayed by a display device (e.g., a touch screen or other display screen, a projector, etc.) of the client device 115 or the server device 104. In some embodiments, an application running on the server system 101 can communicate with the client device 115 to receive user input at the client device 115 and output data, such as visual data, audio data, etc., at the client device 115.
[0037] In some embodiments, the server system 101 and / or any of the one or more client devices 115a, 115n may provide a communication application. The communication program may allow a system (e.g., client device 115 or server system 101) to provide options for communicating with other devices. The communication program may provide one or more associated user interfaces that are displayed on a display device associated with the server system 101 or client device 115. The user interface may provide a user with various options to select a communication mode, a user or device to communicate with, and the like. In some examples, the communication program may provide an option for sending or broadcasting a content release (e.g., to a broadcast area), and / or may output a notification indicating that the content release has been received by the device and, for example, that the device is in a defined broadcast area for the release. The communication program may display or otherwise output the transmitted content release and the received content release, for example, in any of a variety of formats. For example, a content release may include an image shared with other users.
[0038] Other embodiments of the features described herein may use any type of system and / or service. For example, other networked services (e.g., connected to the Internet) may be used instead of or in addition to social networking services. Any type of electronic device may utilize the features described herein. Some embodiments may provide one or more features described herein on one or more client or server devices that are disconnected or intermittently connected to a computer network. In some examples, a client device 115 including or connected to a display device may display data (e.g., content) stored on a storage device to the client device 115, for example, previously received over a communication network.
[0039] Example device 200
[0040] Figure 2 is a block diagram of an example device 200 that may be used to implement one or more features described herein. In one example, the device 200 may be used to implement a client device 115, e.g., Figure 1 Any of the client devices 115a, 115n shown. Alternatively, the device 200 may implement a server device, e.g. Figure 1 The server device 104 is shown. In some embodiments, the device 200 can be used to implement a client device, a server device, or both a client device and a server device. The device 200 can be any suitable computer system, server, or other electronic or hardware device described above.
[0041] One or more of the methods described herein can be run on a standalone program executing on any type of computing device, a program running on a web browser, a mobile application ("app") running on a mobile computing device (e.g., a cellular phone, a smart phone, a smart display, a tablet computer, a wearable device (a watch, an armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, a head-mounted display, etc.), a laptop computer, etc.). In one example, a client / server architecture can be used, for example, where a mobile computing device (as a client device) sends user input data to a server device and receives final output data from the server for output (e.g., for display). In another example, all computations can be performed within the mobile application (and / or other applications) on the mobile computing device. In another example, computations can be split between the mobile computing device and one or more server devices.
[0042] In some embodiments, the device 200 includes a processor 202, a memory 204, an input / output (I / O) interface 206, a camera 208, and a display device 210. The processor 202 can be one or more processors and / or processing circuits for executing program code and controlling the basic operations of the device 200. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. The processor can include a system having a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multi-processor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for implementing functionality, a dedicated processor for implementing processing based on a neural network model, a neural network, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, the processor 202 can include one or more coprocessors that implement neural network processing. In some embodiments, the processor 202 may be a processor that processes data to generate a probabilistic output. For example, the output generated by the processor 202 may be imprecise or may be accurate within a certain range of the expected output. The processing need not be limited to a specific geographic location or have time constraints. For example, the processor may perform its functions in real time, offline, in batch mode, etc. Portions of the processing may be performed at different times and in different locations by different (or the same) processing systems. The computer may be any processor that communicates with a memory.
[0043] Memory 204 is typically provided in device 200 for access by processor 202 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., suitable for storing instructions for execution by a processor or processors, and located separate from and / or integrated with processor 202. Memory 204 may store software operated by processor 202 on server device 200, including operating system 212, other applications 214, application data 216, and image management application 103.
[0044] Other applications 214 may include applications such as a camera application, an image gallery or library application, an image management application, a data display engine, a web hosting engine or application, an image display engine or application, a media display application, a communication engine, a notification engine, a social networking engine, a media sharing application, a mapping application, etc. One or more of the methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application with a web page, as a mobile application ("app") running on a mobile computing device, etc. In some embodiments, other applications 214 may each include instructions that enable processor 202 to perform the functions described herein, for example, Figure 3 and Figure 4 some or all of the methods.
[0045] Application data 216 may be data generated by other applications 214 or hardware of device 200. For example, application data 216 may include images captured by camera 208, user actions identified by other applications 214 (e.g., social networking applications), and so on.
[0046] The I / O interface 206 can provide functionality so that the device 200 can interface with other systems and devices. The interfaced devices can be included as part of the device 200 or can be separated from the device 200 and communicate with it. For example, a network communication device, a storage device (e.g., a memory and / or database 199) and an input / output device can communicate through the I / O interface 206. In some embodiments, the I / O interface can be connected to an interface device, such as an input device (keyboard, pointing device, touch screen, microphone, camera, scanner, sensor, etc.) and / or an output device (display device, speaker device, printer, motor, etc.).
[0047] Some examples of the devices that can be connected to the interface connection of I / O interface 206 can include one or more display devices 210, and this display device 210 can be used for display content, for example, the user interface of image, video and / or output application described herein.Display device 210 can be connected to device 200 by local connection (for example, display bus) and / or by network connection and can be any suitable display device.Display device 210 can include any suitable display device, such as, liquid crystal display (LCD), light emitting diode (LED) or plasma display screen, cathode ray tube (CRT), television, monitor, touch screen, three-dimensional display screen or other visual display device.For example, display device 210 can be arranged on the flat display screen on mobile device, the multiple display screens embedded in glasses form factor or earphone device or the monitoring screen of computer equipment.
[0048] The I / O interface 206 can interface to other input and output devices. Some examples include one or more cameras, such as camera 208, which can capture images. Some embodiments may provide a microphone for capturing sound (e.g., as part of a captured image, voice command, etc.), an audio speaker device for outputting sound, or other input and output devices.
[0049] The camera 208 can be any type of camera that can capture a video comprising a plurality of frames. As used herein, a camera can include any image capture device. In some embodiments, the camera 208 can include a plurality of lenses having different capabilities, e.g., front-facing versus rear-facing, different zoom levels, image resolution of the captured images, etc. In some embodiments, the device 200 can include one or more sensors, such as a depth sensor, an accelerometer, a position sensor (e.g., a global positioning system (GPS)), a gyroscope, etc. In some embodiments, the one or more sensors can operate in conjunction with the camera 208 to obtain sensor readings corresponding to different frames of video captured using the camera 208.
[0050] Sample Image Management Application 103
[0051] The image management application 103 may include a feature detection machine learning module 218 , a user-specific machine learning module 220 , a compression module 222 , and a user interface module 224 .
[0052] In some embodiments, the feature detection machine learning module 218 generates a feature detection machine learning model to identify features from an image. For example, a feature can be a vector (embedding) in a multidimensional feature space. Images with similar features may have similar feature vectors, for example, the vector distance between feature vectors of such images may be smaller than the vector distance between different images. The feature space can be a function of various factors of the image, for example, the subject depicted (the object detected in the image), the composition of the image, color information, image orientation, image metadata, specific objects recognized in the image (e.g., known faces, with user permission), etc.
[0053] The user-specific machine learning module 220 can generate a user-specific machine learning model that is personalized for a user account associated with the user. The user-specific machine learning module 220 can use the user-specific machine learning model to determine a compression setting for an image based on features in the image. This is beneficial in preserving storage space while keeping the image as high quality as possible based on the content that the user is interested in. For example, the user-specific machine learning model can output an indication that an image of a sunset is to be compressed with a high compression ratio. For example, such an indication can be determined based on the user-specific machine learning model analyzing an image that includes features detected by the feature detection machine learning module 218. If the image is not associated with features that determine that the image is not important to the user because the image includes a sunset, the compression setting should be the highest compression level for the image.
[0054] Example Feature Detection Machine Learning Module 218
[0055] The feature detection machine learning module 218 generates a feature detection machine learning model that determines one or more features of an input image. In some embodiments, the feature detection machine learning module 218 includes an instruction set that is executable by the processor 202 to generate the feature detection machine learning model. In some embodiments, the feature detection machine learning module 218 is stored in the memory 204 of the device 200 and is accessible and executable by the processor 202.
[0056] In some embodiments, the feature detection machine learning module 218 can use training data to generate a training model, specifically a feature detection machine learning model. For example, the training data can include any type of data, such as images (e.g., still images, dynamic photos / motion images, image frames from a video, etc.) and, optionally, corresponding features (e.g., labels or tags associated with each image that identify objects in the image).
[0057] For example, the training data may include a training set that includes a plurality of digital images and corresponding features. In some embodiments, the training data may include images with enhancements (such as rotation, light shift, and color shift) to provide invariance in the model when it is provided with user photos that may be rotated or have unusual characteristics (e.g., artifacts of the camera used to capture the image). The training data may be obtained from any source, such as a data repository specifically labeled for training, data that has been provided with permission to be used as training data for machine learning, etc. In embodiments where one or more users have permitted the use of their respective user data to train the machine learning model, the training data may include such user data. In embodiments where the user has permitted the use of their respective user data, the data may include permitted data such as images / videos or image / video metadata (e.g., images, corresponding features that may be derived from users who provided manual labels or tags), communications (e.g., messages on social networks; emails; chat data such as text messages, voice, video, etc.), documents (e.g., spreadsheets, text documents, presentations, etc.), and the like.
[0058] In some embodiments, the training data may include synthetic data generated for the purpose of training, such as data that is not based on user input or activity in the context being trained, for example, data generated by simulations or computer-generated images / video, etc. In some embodiments, the feature detection machine learning module 218 uses weights that were taken from another application and not edited / migrated. For example, in these embodiments, a training model may be generated, for example, on a different device, and provided as part of the image management application 103. In various embodiments, the training model may be provided as a data file that includes a model structure or form (e.g., which defines the number and type of neural network nodes, the connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. The feature detection machine learning module 218 may read the data file for the training model and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the training model.
[0059] The feature detection machine learning module 218 generates a trained model, referred to herein as a feature detection machine learning model. In some embodiments, the feature detection machine learning module 218 is configured to apply the feature detection machine learning model to data, such as application data 216 (e.g., an input image), to identify one or more features in the input image and generate a feature vector (embedding) representing the image. In some embodiments, the feature detection machine learning module 218 may include software code to be executed by the processor 202. In some embodiments, the feature detection machine learning module 218 may specify a circuit configuration (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables the processor 202 to apply the feature detection machine learning model. In some embodiments, the feature detection machine learning module 218 may include software instructions, hardware instructions, or a combination. In some embodiments, the feature detection machine learning module 218 may provide an application programming interface (API) that can be used by the operating system 212 and / or other applications 214 to call the feature detection machine learning module 218, for example, to apply the feature detection machine learning model to the application data 216 to determine one or more features of the input image.
[0060] In some embodiments, the feature detection machine learning model may include one or more model forms or structures. In some embodiments, the feature detection machine learning model may use a support vector machine, however, in some embodiments, a convolutional neural network (CNN) is more preferable. For example, the model form or structure may include any type of neural network, such as a linear network, a deep neural network that implements multiple layers (e.g., a "hidden layer" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (CNN) (e.g., a network that splits or divides input data into multiple parts or tiles, uses one or more neural network layers to process each tile separately, and aggregates the processing results for each block), a sequence-to-sequence neural network (e.g., a network that receives sequence data (such as words in a sentence, frames in a video, etc.) as input and produces a sequence of results as output), etc.
[0061] The model form or structure can specify the connectivity between the various nodes and the organization of the nodes into layers. For example, the nodes of a first layer (e.g., an input layer) can receive data or application data 216 as input data. For example, such data can include one or more pixels per node, such as when a feature detection machine learning model is used to analyze an input image, such as a first image associated with a user account. Subsequent intermediate layers can receive the outputs of the nodes of the previous layer as input according to the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. The last layer (e.g., an output layer) produces the output of the machine learning application. For example, the output can be image features associated with the input image. In some embodiments, the model form or structure also specifies the number and / or type of nodes in each layer.
[0062] Features output by the feature detection machine learning module 218 may include subject matter (e.g., a sunset versus a particular person); colors present in the image (e.g., a green mountain versus a blue lake); color balance; light source; angle and intensity; location of objects in the image (e.g., following the rule of thirds); location of objects relative to each other and the position of the lens (e.g., depth of field); focus (foreground versus background); or shadows. While the foregoing features are human-understood, it will be understood that the feature output may be an embedding or other mathematical value that represents the image and is not human-parsable (e.g., no individual feature value may correspond to a specific feature, such as the color present, object location, etc.); however, the trained model is robust to images such that similar features are output for similar images, and images with significant differences have correspondingly different features.
[0063] In some embodiments, the model form is a CNN having network layers, wherein each network layer extracts image features at a different level of abstraction. A CNN for identifying features in an image can be used for image classification. In some embodiments, a CNN can be used to identify features of an image and then apply transfer learning by replacing the classification layer, or more specifically a fully connected feedforward neural network output layer, with a user-specific machine learning model described below. In some embodiments, the CNN is a VGGnet, ResNet, AlexNet, Inception network, or any other advanced neural network considered for image processing applications, and is trained using a training set of digital images, such as ImageNet. The model architecture can include a combination and ordering of layers consisting of multi-dimensional convolutions, average pooling, maximum pooling, activation functions, normalization, regularization, and other layers and modules actually used to apply deep neural networks.
[0064] In various embodiments, the feature detection machine learning model may include one or more models. The one or more models may include multiple nodes, or in the case of a CNN - a filter bank, may be arranged into layers according to a model structure or form. In some embodiments, the node may be a computational node without memory, for example, configured to process one unit of input to produce one unit of output. For example, the computation performed by the node may include multiplying each of the multiple node inputs by a weight to obtain a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output.
[0065] In some embodiments, the computations performed by the nodes may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, separate processing units using a GPU, or dedicated neural circuitry. In some embodiments, the nodes may include memory, for example, to enable storage and use of one or more earlier inputs when processing subsequent inputs. For example, a node with memory may include a long short-term memory (LSTM) node. An LSTM node may use memory to maintain state that allows the node to operate like a finite state machine (FSM). Models with such nodes may be useful when processing sequential data, such as words in a sentence or paragraph, a series of images, frames in a video, speech or other audio, and the like. For example, a heuristic-based model used in a gating model may store one or more previously generated features corresponding to previous images.
[0066] In some embodiments, the feature detection machine learning model can include embeddings or weights for individual nodes. For example, the feature detection machine learning model can be initialized as a plurality of nodes organized into layers specified by the model form or structure. Upon initialization, corresponding weights can be applied to the connections between each pair of nodes connected according to the model form (e.g., nodes in successive layers of a neural network). For example, the corresponding weights can be randomly assigned or initialized to default values. The feature detection machine learning model can then be trained, for example, using a training set of digital images, to produce results. In some embodiments, subsets of the entire architecture can be reused as transfer learning methods from other machine learning applications to leverage pre-trained weights.
[0067] For example, training can include applying supervised learning techniques. In supervised learning, the training data can include multiple inputs (e.g., a set of digital images) and corresponding expected outputs for each input (e.g., one or more features for each image). Based on a comparison of the output of the feature detection machine learning model with the expected output, the values of the weights are automatically adjusted, for example, in a manner that increases the probability that the feature detection machine learning model will produce the expected output when provided with similar inputs.
[0068] In some embodiments, training can include applying unsupervised learning techniques. In unsupervised learning, only input data (e.g., images with labeled features) can be provided, and a feature detection machine learning model can be trained to distinguish the data, for example, clustering the features of the images into groups, where each group includes images with features that are similar in some way.
[0069] In various embodiments, the training model includes a set of weights corresponding to the model structure. In embodiments where a training set of digital images is omitted, the feature detection machine learning module 218 can generate a feature detection machine learning model based on previous training, for example, by the developer of the feature detection machine learning module 218, by a third party, etc. In some embodiments, the feature detection machine learning model can include a fixed set of weights (e.g., downloaded from a server that provides weights).
[0070] In some embodiments, the feature detection machine learning module 218 can be implemented in an offline manner. In these embodiments, the feature detection machine learning module can be generated in a first phase and provided as part of the feature detection machine learning module 218. In some embodiments, small updates to the feature detection machine learning model can be implemented in an online manner. In such embodiments, an application that calls the feature detection machine learning module 218 (e.g., the operating system 212, one or more of the other applications 214, etc.) can utilize the feature detections generated by the feature detection machine learning module 218, for example, provide the feature detections to a user-specific machine learning module 220, and can generate a system log (e.g., if the user permits, the actions taken by the user based on the feature detections; or if used as input for further processing, the results of the further processing). The system log can be generated periodically, for example, hourly, monthly, quarterly, etc., and can be used to update the feature detection machine learning model with the user's permission, for example, to update the embeddings of the feature detection machine learning model.
[0071] In some embodiments, the feature detection machine learning module 218 can be implemented in a manner that can be adapted to the specific configuration of the device 200 on which the feature detection machine learning module 218 is executed. For example, the feature detection machine learning module 218 can determine a computation graph that utilizes available computational resources, e.g., the processor 202. For example, if the feature detection machine learning module 218 is implemented as a distributed application across multiple devices, the feature detection machine learning module 218 can determine computations to be performed on separate devices in a manner that optimizes the computations. In another example, the feature detection machine learning module 218 can determine that the processor 202 includes a GPU with a specific number of GPU cores (e.g., 1000) and implement the feature detection machine learning module 218 accordingly (e.g., as 1000 separate processes or threads).
[0072] In some embodiments, the feature detection machine learning module 218 can implement an ensemble of training models. For example, the feature detection machine learning model can include multiple training models, each model being applied to the same input data. In these embodiments, the feature detection machine learning module 218 can select a particular training model based on, for example, available computing resources, the success rate of previous inferences, etc.
[0073] In some embodiments, the feature detection machine learning module 218 can execute multiple training models. In these embodiments, the feature detection machine learning module 218 can combine the outputs from applying separate models, for example, using a voting technique that scores the separate outputs from applying each training model or by selecting one or more specific outputs. In some embodiments, such a selector is part of the model itself and acts as a connection layer between the training models. Further, in these embodiments, the feature detection machine learning module 218 can apply a time threshold to apply the separate training models (e.g., 0.5ms) and only utilize those separate outputs that are available within the time threshold. Outputs that are not received within the time threshold may not be utilized, for example, and may be discarded. For example, these methods may be appropriate when there is a time limit specified when calling the feature detection machine learning module 218, for example, by the operating system 212 or one or more applications 612.
[0074] Example User-Specific Machine Learning Module 220
[0075] The user-specific machine learning module 220 generates a user-specific machine learning model that determines compression settings for the same input image analyzed by the feature detection machine learning module 218. In some embodiments, the user-specific machine learning module 220 includes an instruction set that is executable by the processor 202 to generate the user-specific machine learning model. In some embodiments, the user-specific machine learning module 220 is stored in the memory 204 of the device 200 and can be accessed and executed by the processor 202.
[0076] In some embodiments, the user-specific machine learning module 220 can use training data to generate a training model, specifically a user-specific machine learning model. The training data can include any type of data, such as user-specific features that indicate user actions with respect to one or more prior images of a feature detection machine learning model. For example, a user-specific feature can indicate the user's level of interest in an image. An image indicating that it is a favorite (e.g., marked as a favorite through explicit user input) can be considered an important image that should not be compressed. Other examples of user-specific features can include user actions, such as any of the following: tagging other users in an image; sharing an image; creating an album or other image-based creation; commenting on an image; metadata important to the user (e.g., location data indicating that the image was captured at a significant location); downloading an image, editing an image, ordering a print of an image; editing an image; data from an explicit query asking the user whether they liked an image, etc. In some embodiments, user actions can include actions directed at another user's image, such as using natural language processing to determine the sentiment of a comment on another user's image as a signal of interest in the image, indicating approval (e.g., liking) of another user's image, saving another user's image, downloading another user's image, etc. These user actions signal the value of the image and can be used as input to the model.
[0077] The training data may be obtained from any source, such as a data repository specifically labeled for training, permitted data provided for use as training data for machine learning, and the like. In embodiments where one or more users have permitted the use of their respective user data to train the machine learning model, the training data may include such user data. In embodiments where the user has permitted the use of their respective user data, the data may include permitted data such as images / videos or image / video metadata (e.g., images, corresponding features of the images, user-specific features associated with the user, how the user-specific features indicate descriptions of one or more prior images, and the like), communications (e.g., messages on social networks; emails; chat data such as text messages, voice, video, and the like), documents (e.g., spreadsheets, text documents, presentations, and the like), and the like. In some embodiments, the prior images that were used by the feature detection machine learning module 218 to generate a feature detection machine learning model for identifying one or more features are used by the user-specific machine learning module 220 to generate a user-specific machine learning model to determine compression settings based on user-specific features indicative of user actions and corresponding image features of the prior images.
[0078] For example, a training model can be generated, for example, on a different device, and provided as part of the image management application 103. In various embodiments, the training model can be provided as a data file that includes a model structure or form (e.g., which defines the number and type of neural network nodes, the connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. The user-specific machine learning module 220 can read the data file for the training model and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the training model.
[0079] The user-specific machine learning module 220 generates a training model, referred to herein as a user-specific machine learning model. In some embodiments, the user-specific machine learning module 220 is configured to apply the user-specific machine learning model to data, such as data of the compression module 222, to identify compression settings for an input image. In some embodiments, the user-specific machine learning module 220 may include software code to be executed by the processor 202. In some embodiments, the user-specific machine learning module 220 may specify a circuit configuration (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables the processor 202 to apply the user-specific machine learning model. In some embodiments, the user-specific machine learning module 220 may include software instructions, hardware instructions, or a combination.
[0080] In some embodiments, the user-specific machine learning model may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep neural network, such as a multi-layer feed-forward fully connected neural network, a CNN, or a sequence-to-sequence neural network, as discussed in more detail above.
[0081] The model form or structure can specify the connectivity between the various nodes and the organization of the nodes into layers. For example, the nodes of a first layer (e.g., an input layer) can receive data or application data 216 as input data. For example, such data can include one or more user-specific features for each node, such as when a feature detection machine learning model is used to analyze user-specific features that indicate user actions associated with an image. Subsequent intermediate layers can receive the outputs of the nodes of the previous layer as input according to the connectivity specified in the model form or structure. These layers can also be referred to as hidden layers. The last layer (e.g., an output layer) produces the output of the machine learning application. For example, based on the user-specific features, the output can be a compression setting for the image. More specifically, the output can be a determination of the user's level of interest in the image, which corresponds to a rating of the image, and the user-specific machine learning module 220 maps the rating to a compression setting. In some embodiments, the model form or structure also specifies the number and / or type of nodes in each layer.
[0082] The user-specific machine learning model receives input from the feature detection machine learning module 218 to identify features from the input image fed into the user-specific machine learning model. The input allows the feature detection machine learning model to identify features from the input image and determine whether the user is interested in the input image. The user-specific machine learning model is trained on user-specific features that indicate user actions to identify which features in the image the user is interested in based on signals such as sharing a photo, viewing a photo, etc. In some embodiments, the signal is a label that is used to explicitly estimate the relative importance to the user (e.g., user rating, stars on a photo, shared at least once, x amount of views, etc.). In some embodiments, these can be placed in a stack rating to create clusters. Within the cluster, the user-specific machine learning model uses a feature detection algorithm to generate a similarity measure, which is then used to estimate the relative importance to the user.
[0083] In some embodiments where no user-specific features are available, a user-specific machine learning model can generate a baseline compression setting. For example, the baseline compression setting can be generated by user input from other users, where the input is image ratings based on blurriness or other types of indicators of less interesting images. In this example, the user-specific machine learning model can apply higher compression ratios to these and specific types of images (e.g., receipts).
[0084] The output of the user-specific machine learning model can be a rating that identifies the user's level of interest, such as using a scale of 1-5, 1-10, etc., or a regression output (analog value), such as 7.83790. The user-specific machine learning model can map the rating to a compression setting, such as a compression ratio. Below is an example of how to map the rating to a compression ratio. In this example, for a rating of 1, the compressed image will occupy 0.2 of the original image's original resolution.
[0085] grade Compress(original size:new size) 1 1:0.2 2 1:0.3 3 1:0.5 4 1:0.8 5 1:1
[0086] While the above examples describe compression settings that include levels mapped to compression ratios, other examples of compression settings are possible. For example, a user-specific machine learning model may determine the compression technique, the image format (e.g., using JPEG, WebP, HEIF, etc.), the parameters selected for optimization (e.g., dynamic range, image resolution, color, whether the compression is progressive), etc. In another example, the compression setting may be a determination to maintain one or more features in a high-resolution image and compress the rest of the image, such as by determining one or more regions of interest in the image. In another example, in addition to the user-specific machine learning model determining levels of different features in the image, the user-specific machine learning model may also determine that quality trade-offs are considered for certain types of images. For example, dynamic range may be critical for sunset images, resolution may be more important for close-ups, etc.
[0087] In yet another example, the user-specific machine learning model may also indicate how to apply rankings when different features are included in the same image. For example, if images typically have a ranking of 5 when including sunsets, but typically have a ranking of 2 when including food, the user-specific machine learning model may apply a ranking that indicates the most interest, i.e., a ranking of 5. In another embodiment, the user-specific machine learning model may determine that a particular user's reaction to an image suggests that the rankings of multiple features should be averaged, that certain features should be associated with higher weights, etc.
[0088] In some embodiments, the user-specific machine learning model can determine that the rating of the input image satisfies an importance threshold, and in response to the rating satisfying the importance threshold, the user-specific machine learning model provides a suggestion to the user to share the input image. In another embodiment, in response to the rating satisfying the importance threshold, the user-specific machine learning model prioritizes backup of the input image over other images associated with the user account that do not meet the importance threshold. For example, if client device 115 is located in an area with limited internet access, the user-specific machine learning model can instruct client device 115 to transfer the image with the highest rating (or in descending order based on the rating) to be transferred to server system 101 for storage. In another embodiment, in response to the rating satisfying the importance threshold, the user-specific machine learning model instructs user interface module 224 to provide a user interface including instructions for capturing a subsequent image if the scene depicted in the subsequent image has at least one of the one or more features of the input image. For example, if the user is interested in taking a photo of a plant, the user interface can inform the user that water droplets in the photo may obscure the petals. In another example, if a feature of interest is observed in the image, the camera can automatically lock focus on the feature of interest (and allow the user to change focus with a click).
[0089] In various embodiments, the user-specific machine learning model may include one or more models. The one or more models may include multiple nodes, arranged in layers according to the model structure or form. In some embodiments, the node may be a computational node without memory, for example, configured to process one unit of input to produce one unit of output. For example, the computation performed by the node may include multiplying each of the multiple node inputs by a weight to obtain a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function.
[0090] In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, separate processing units using a GPU, or dedicated neural circuitry. In some embodiments, a node may include memory, for example, to enable storage and use of one or more earlier inputs when processing subsequent inputs. For example, a node with memory may include an LSTM node. An LSTM node may use memory to maintain state that allows the node to operate like an FSM.
[0091] In some embodiments, the user-specific machine learning model may include embeddings or weights for individual nodes. For example, the user-specific machine learning model may be initialized as a plurality of nodes organized into layers specified by the model form or structure. Upon initialization, corresponding weights may be applied to the connections between each pair of nodes connected according to the model form (e.g., nodes in successive layers of a neural network). For example, the corresponding weights may be randomly assigned or initialized to default values. The user-specific machine learning model may then be trained, e.g., using a user-specific training set, to produce results.
[0092] Training can include applying supervised learning techniques. In supervised learning, the training data can include multiple inputs (e.g., ratings given for different types of images) and corresponding expected outputs (e.g., compression settings) for each input. Based on a comparison of the output of the feature detection machine learning model (e.g., predicted ratings) with the expected outputs (e.g., ratings provided by the user), the values of the weights are automatically adjusted, for example, in a manner that increases the probability that the user-specific machine learning model will produce the expected output when provided with similar inputs. Included below is an example of user-provided ratings for images of different types associated with the user. In this example, the lowest rating is associated with the least important, and the highest rating is associated with the most important.
[0093] category Given level Receipt 1 screenshot 1 food 2 car 3 cat 4 Sunset 5 landscape 5
[0094] In various embodiments, the training model includes a set of weights or embeddings corresponding to the model structure. In embodiments where a training set of digital images is omitted, the user-specific machine learning module 220 can generate a user-specific machine learning model based on previous training, for example, by the developer of the user-specific machine learning module 220, by a third party, etc. In some embodiments, the user-specific machine learning model can include a fixed set of weights (e.g., downloaded from a server that provides weights).
[0095] The user-specific machine learning module 220 can be implemented offline and / or as a whole of a training model and using different formats. As described above with respect to the feature detection machine learning module 218, it is understood that the same description can also be applied to the user-specific machine learning module 220. Therefore, this description will not be repeated.
[0096] Example Compression Module 222
[0097] The compression module 222 compresses the input image based on the compression settings determined by the user-specific machine learning module 220. In some embodiments, the compression module 222 may include an instruction set that is executable by the processor 202 to compress the input image. In some embodiments, the compression module 222 is stored in the memory 204 of the device 200 and is accessible and executable by the processor 202.
[0098] The compression module 222 can receive the input image from the feature detection machine learning module 218 and the compression settings from the user-specific machine learning module 220. The compression module 222 applies the compression settings to the input image. The compression module 222 can replace the original input image with the compressed input image to reduce the file size, thereby more efficiently utilizing the memory 204 and / or storage device on which the image is stored. In some embodiments, the compression module 222 can transfer the compressed input image to another location for storage. For example, if the image management application 103b is part of the client device 115a, the compression module 222 can transfer the compressed input image to the server system 101 for storage.
[0099] Example User Interface Module 224
[0100] User interface module 224 generates a user interface that receives input from a user. In some embodiments, user interface module 224 includes an instruction set that is executable by processor 202 to compress an input image. In some embodiments, user interface module 224 is stored in memory 204 of device 200 and is accessible and executable by processor 202.
[0101] In some embodiments, the user interface module 224 generates a user interface for changing various settings associated with the image management application 103. In some embodiments, the user interface module 224 generates the user interface and receives user input to determine a baseline compression setting. For example, the user interface module 224 can generate a user interface that is visible to a user associated with a user account.
[0102] The user interface may include two or more versions of a sample image, where each image is compressed using a different compression setting. The user interface may include a prompt asking the user to identify a particular version of the sample image as a baseline compression setting for the user account. For example, the baseline compression setting may represent the lowest compression setting accepted by the user for images associated with the user account. The user interface module 224 may request user input regarding the baseline compression setting multiple times to confirm the accuracy of the baseline compression setting. For example, the user interface module 224 may provide different features to the image, display the user interface periodically (once a week, once a month, whenever there is a software update for the image management application 103), etc.
[0103] In some embodiments, the user interface module 224 can generate a warning in response to a user's selection. For example, if the user selects a high-resolution compression setting or no compression setting at all, the user interface module 224 can warn the user that storage space will be exhausted within a certain number of days or after a certain number of additional images are captured. This estimate can be based on the average size of photos uploaded by the user over the past x days.
[0104] The user interface module 224 can transmit the results of the user's selection to the user-specific machine learning module 220 for use in determining the compression setting for the input image. The machine learning module 220 can then map the grades to compression settings based on the baseline compression setting.
[0105] Any software in memory 204 may alternatively be stored in any other suitable storage location or computer-readable medium. Additionally, memory 204 (and / or other connected storage devices) may store one or more messages, one or more taxonomies, electronic encyclopedias, dictionaries, thesauri, knowledge bases, message data, grammars, and / or other instructions and data used in the features described herein. Memory 204 and any other type of storage (disk, optical disk, tape, or other tangible medium) may be considered "storage" or "storage devices."
[0106] For ease of explanation, Figure 2 A block for each of processor 202, memory 204, I / O interface 206, camera 208, display device 210, and software blocks 103, 218, 220, 220, 222, and 224 is shown. These blocks can represent one or more processors or processing circuit systems, operating systems, memories, I / O interfaces, applications, and / or software modules. In other embodiments, device 200 may not have all of the components shown and / or may have other elements, including other types of elements, rather than or in addition to the elements described herein. Although some components are described as performing the blocks and operations described in some embodiments herein, any suitable component or combination of components of any suitable processor of environment 100, device 200, similar systems, or associated with such systems can perform the described blocks and operations.
[0107] The methods described herein can be implemented by computer program instructions or codes that can be executed on a computer. For example, the code can be implemented by one or more digital processors (e.g., microprocessors or other processing circuit systems) and can be stored on a computer program product, including non-transitory computer-readable media (e.g., storage media), such as magnetic, optical, electromagnetic or semiconductor storage media, including semiconductor or solid-state memory, tape, removable computer floppy disk, random access memory (RAM), read-only memory (ROM), flash memory, rigid disk, optical disk, solid-state memory drive, etc. The program instructions can also be contained in an electronic signal and provided as an electronic signal, for example, in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods can be implemented in hardware (logic gates, etc.) or in a combination of hardware and software. Example hardware can be a programmable processor (e.g., a field programmable gate array (FPGA), a complex programmable logic device), a general-purpose processor, a graphics processor, an application-specific integrated circuit (ASIC), etc. One or more methods can be executed as part or component of an application running on a system or as an application or software running in combination with other applications and an operating system.
[0108] Example Method
[0109] Figure 3 is a flow chart illustrating an example method 300 of using a feature detection machine learning model 304 to identify one or more features 306 in an input image 302 and using a user-specific machine learning model 308 to determine compression settings 310 from the one or more features of the image, in accordance with some embodiments.
[0110] The feature detection machine learning model 304 may include a deep neural network, such as a convolutional neural network (CNN) with a set of layers that construct more abstract objects from pixels. In some embodiments, the early layers of the CNN detect edges, and as the depth of the layers increases, the human-defined meaning of the features increases. For example, mid-stage layers may detect components of an object, and later-stage layers may detect the object (or face) itself.
[0111] An input image 302 is provided as input to a feature detection machine learning model 304. The input image 302 is from a user associated with a user account. The input image 302 may be received by an input layer in a layer set. The input layer may be connected to a second layer in the layer set. In some embodiments, each of one or more additional layers receives the output of the previous layer as input and provides input to the next layer. The feature detection machine learning model 304 generates one or more features 306 based on the input image 302. The last layer in the layer set may be an output layer. Continuing with this example, the output layer may output one or more features 306.
[0112] In some embodiments, the output may include a corresponding probability that each feature has been accurately identified in the image. The output of the feature detection machine learning model 304 may be a numeric vector, a probability value, or a set of probability values (e.g., each corresponding to a specific stack of video frames). The output of the feature detection machine learning model 304 is provided as input to the user-specific machine learning model 308.
[0113] The user-specific machine learning model 308 may also include a deep neural network, such as a CNN. In some embodiments, the user-specific machine learning model 308 may be generated using transfer learning by replacing the classification layer of the feature detection machine learning model 304 with a component trained using user-specific features associated with the user, where the user-specific features indicate user actions with respect to prior images. In some embodiments, the prior images are also used to train the feature detection machine learning model 304.
[0114] The user-specific machine learning model 308 can receive one or more features 306 via an input layer. The input layer can be connected to the second layer of the plurality of layers. In some embodiments, one or more additional layers can be included in the user-specific machine learning model 308, each receiving the output of the previous layer as input and providing input to the next layer. The last layer of the user-specific machine learning model 308 can be an output layer. In some embodiments, the model 308 can have a single input layer that directly outputs the compression setting 310.
[0115] The user-specific machine learning model 308 may generate as output a compression setting 310 (prediction 310) and, optionally, a probability associated with the compression setting 310. In some embodiments, the compression setting 310 may include one or more levels of one or more features in the input image 302. The probability may include a probability value, a set of probability values (e.g., each corresponding to a level of a particular feature in the input image 302), or a vector representation generated by an output layer of the user-specific machine learning model 308.
[0116] In some embodiments, the method 300 can be implemented on one or more of the client devices 115a, 115n, for example, as part of the image management application 103b. In some embodiments, the method 300 can be implemented on the server device 104, for example, as part of the image management application 103a. In some embodiments, the method 300 can be implemented on the server device 104 and on one or more of the client devices 115a, 115n.
[0117] In some embodiments, method 300 may be implemented as software executable on a general-purpose processor (e.g., a central processing unit (CPU) of a device). In some embodiments, method 300 may be implemented as software executable on a dedicated processor (e.g., a graphics processing unit (GPU), a field programmable gate array (FPGA), a machine learning processor, etc.). In some embodiments, method 300 may be implemented as dedicated hardware, for example, as an application-specific integrated circuit (ASIC).
[0118] Figure 4 is a flow chart illustrating an example method 400 of creating a training model according to some embodiments.
[0119] Method 400 may begin at block 402. In block 402, a determination is made as to whether user consent to use user data has been obtained. For example, user interface module 224 may generate a user interface requesting user consent to use user data when generating a feature detection machine learning model and / or a user-specific machine learning model. If user consent is not obtained, then in block 404, a baseline model is used instead of the feature detection machine learning model and / or the user-specific machine learning model.
[0120] If user consent is obtained, method 400 may proceed to block 406, where a training set of digital images and corresponding features are obtained. Block 406 may be followed by block 408. In block 408, a feature detection machine learning model is trained based on the training set and the corresponding features. After training, the feature detection machine learning model is capable of identifying image features in an input image provided to the feature detection machine learning model. Block 408 may be followed by block 410.
[0121] In block 410, a training set of user-specific features associated with a user is obtained, wherein the user-specific features indicate user actions with respect to one or more prior images. In some embodiments, the prior images are the same as the set of digital images. In block 410, block 412 may be followed. In block 412, a user-specific machine learning model is trained based on the user-specific features and the one or more prior images, wherein, after training, the user-specific machine learning model determines a ranking for an input image provided to the user-specific machine learning model.
[0122] Figure 5 is a flow chart illustrating an example method 500 of applying a model to an input image according to some embodiments.
[0123] In block 502, an input image is obtained, the input image being associated with a user account. Block 502 may be followed by block 504. In block 504, one or more features of the input image are determined using a feature detection machine learning model. Block 504 may be followed by block 506. In block 506, a compression setting for the input image is determined based on the one or more features in the input image using a user-specific machine learning model personalized for the user account. Block 506 may be followed by block 508. In block 508, the input image is compressed based on the compression setting.
[0124] Prior to training, each node can be assigned an initial weight, and the connections between nodes in different layers of the neural network can be initialized. Training can include adjusting the weights of one or more nodes and / or the connections between one or more pairs of nodes.
[0125] In some embodiments, a subset of the training set may be excluded during the initial training phase. This subset may be provided after the initial training phase, and the accuracy of the prediction (indication of whether to analyze the video) may be determined. If the accuracy is below a threshold, further training may be performed using additional digital images from the training set or user-specific features, respectively, to adjust the model parameters until the model correctly predicts its output.
[0126] Further training (second phase) can be repeated any number of times, for example, until the model reaches a satisfactory level of accuracy. In some embodiments, the trained model can be further modified, for example, compressed (using fewer nodes or layers), converted (for example, to be usable on different types of hardware), etc. In some embodiments, different versions of the model can be provided, for example, a client version of the model can be optimized for size and have reduced computational complexity, while a server version of the model can be optimized for accuracy.
[0127] Although methods 400 and 450 have been referenced Figure 4 The various blocks in FIG are described, but it is understood that the technology described in this disclosure can be used without executing Figure 4 In some embodiments, Figure 4 One or more of the blocks shown in may be combined.
[0128] Furthermore, while training has been described with reference to a training set, the feature detection machine learning model and the user-specific machine learning model can be trained during operation. For example, if a user requests that an image be compressed using specific compression settings, the feature detection machine learning model and the user-specific machine learning model can be updated to include the user information. In some embodiments, the user can manually provide annotations, for example, providing a list of features and their corresponding ratings. With the user's permission, some embodiments can utilize such annotations to train the feature detection machine learning model and the user-specific machine learning model.
[0129] Although the description has been given with respect to specific embodiments thereof, these specific embodiments are merely illustrative and not restrictive. The concepts shown in the examples can be applied to other examples and embodiments.
[0130] Where certain embodiments discussed herein may collect or use personal information about a user (e.g., user data, information about the user's social network, the user's location and time at that location, the user's biometric information, the user's activities, and demographic information), the user may be provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use user personal information, particularly upon receiving explicit authorization from the relevant user to do so. The user has the ability to permanently delete these models.
[0131] For example, a user is provided with control over whether a program or feature collects user information about that particular user or other users associated with the program or feature. Each user whose personal information is to be collected is presented with one or more options that allow control over the collection of information related to that user to provide permission or authorization regarding whether information is to be collected and which portions of the information are to be collected. For example, one or more such control options may be provided to the user via a communication network. Additionally, certain data may be processed in one or more ways before being stored or used to remove personally identifiable information. As an example, the identity of a user may be processed so that personally identifiable information cannot be determined. As another example, the geographic location of a client device may be generalized to a larger area so that the specific location of the user cannot be determined.
[0132] It should be noted that the functional blocks, operations, features, methods, devices and systems described in this disclosure can be integrated into or divided into different combinations of systems, devices and functional blocks known to those skilled in the art. Any suitable programming language and programming technique can be used to implement the routine of a specific embodiment. Different programming techniques can be adopted, for example, procedural or object-oriented. The routine can be executed on a single processing device or multiple processors. Although steps, operations or calculations can be presented in a specific order, in different specific embodiments, the order can be changed. In some embodiments, multiple steps or operations shown in order in this specification can be performed simultaneously.
Claims
1. A computer-implemented method of determining a compression setting, the method comprising: providing a first user interface comprising two or more versions of a sample image to a user associated with a user account, each of the two or more versions of the sample image being compressed with a different compression setting; obtaining user input from the user, the user input identifying a particular version of the sample image from the two or more versions; selecting a compression setting associated with the particular version of the sample image as a baseline compression setting for the user account; obtaining an input image, wherein the input image is associated with the user account; determining one or more features of the input image including one or more regions of interest using a feature detection machine learning model; determining, based on the baseline compression setting and the one or more features in the input image, a second compression setting for the input image using a user-specific machine learning model personalized for the user account, wherein the second compression setting corresponds to maintaining the one or more regions of interest in the input image at a high resolution and compressing a remainder of the input image using the baseline compression setting; and The input image is compressed based on the second compression setting.
2. The computer-implemented method of claim 1 , wherein: The feature detection machine learning model is generated by the following operations: Obtaining a training set of digital images and corresponding features; and The feature detection machine learning model is trained based on the training set and the corresponding features, wherein, after training, the feature detection machine learning model is able to identify image features in the input image provided to the feature detection machine learning model.
3. The computer-implemented method of claim 2, wherein: The user-specific machine learning model is generated by: obtaining a training set of user-specific features associated with the user, the user-specific features indicating user actions with respect to one or more prior images; and The user-specific machine learning model is trained based on the user-specific features and the one or more prior images, wherein, after training, the user-specific machine learning model determines a rank of the input image provided to the user-specific machine learning model.
4. The computer-implemented method of claim 3, wherein: The second compression setting includes mapping the level to a compression ratio of the remaining portion of the input image, the compression ratio comparing an original size of the input image to a new size of the input image.
5. The computer-implemented method of claim 3, wherein: Determining one or more characteristics of the input image further includes identifying a type of the input image, and the ranking is based on the type of the input image.
6. The computer-implemented method of claim 3, wherein: The corresponding image features of the one or more prior images are obtained by applying the feature detection machine learning model to the one or more prior images.
7. The computer-implemented method of claim 1 , further comprising: In response to obtaining the user input, a warning is issued predicting that storage space in the user account will be exhausted after a specified number of days or after a predicted number of additional images are captured.
8. The computer-implemented method of claim 1 , wherein: Determining the second compression setting includes: determining a rating of the input image by the user-specific machine learning model; and The level is mapped to the second compression setting, wherein the mapping is based on the baseline compression setting.
9. The computer-implemented method of claim 8, further comprising: determining that the level of the input image satisfies a significance threshold; as well as In response to determining that the rating satisfies the importance threshold, performing at least one action selected from the group consisting of: providing a suggestion to the user to share the input image, prioritizing backup of the input image over backups of other images associated with the user account that do not meet the importance threshold, providing a second user interface including instructions for capturing the subsequent image if the scene depicted in the subsequent image has at least one of the one or more features of the input image, and A combination of the above actions.
10. A computing device comprising: processor; as well as a memory having instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising: providing a first user interface comprising two or more versions of a sample image to a user associated with a user account, each of the two or more versions of the sample image being compressed with a different compression setting; obtaining user input from the user, the user input identifying a particular version of the sample image from the two or more versions; selecting a compression setting associated with the particular version of the sample image as a baseline compression setting for the user account; obtaining an input image, wherein the input image is associated with the user account; determining one or more features of the input image including one or more regions of interest using a feature detection machine learning model; determining, based on the baseline compression setting and the one or more features of the input image, a second compression setting for the input image using a user-specific machine learning model personalized for the user account, wherein the second compression setting corresponds to maintaining the one or more regions of interest in the input image at a high resolution and compressing a remainder of the input image using the baseline compression setting; and The input image is compressed based on the second compression setting.
11. The computing device of claim 10, wherein: The feature detection machine learning model is generated by the following operations: Obtaining a training set of digital images and corresponding features; and The feature detection machine learning model is trained based on the training set and the corresponding features, wherein, after training, the feature detection machine learning model is able to identify image features in the input image provided to the feature detection machine learning model.
12. The computing device of claim 11, wherein: The user-specific machine learning model is generated by: obtaining a training set of user-specific features associated with the user, the user-specific features indicating user actions with respect to one or more prior images; and The user-specific machine learning model is trained based on the user-specific features and the one or more prior images, wherein, after training, the user-specific machine learning model determines a rank of the input image provided to the user-specific machine learning model.
13. The computing device of claim 10, wherein: The operations further include: In response to obtaining the user input, a warning is issued predicting that storage space in the user account will be exhausted after a specified number of days or after a predicted number of additional images are captured.
14. The computing device of claim 10, wherein: Determining the second compression setting includes: determining a rating of the input image by the user-specific machine learning model; and The level is mapped to the second compression setting, wherein the mapping is based on the baseline compression setting.
15. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations comprising: providing a first user interface comprising two or more versions of a sample image to a user associated with a user account, each of the two or more versions of the sample image being compressed with a different compression setting; obtaining user input from the user, the user input identifying a particular version of the sample image from the two or more versions; selecting a compression setting associated with the particular version of the sample image as a baseline compression setting for the user account; obtaining an input image, wherein the input image is associated with the user account; determining one or more features of the input image including one or more regions of interest using a feature detection machine learning model; determining, based on the baseline compression setting and the one or more features of the input image, a second compression setting for the input image using a user-specific machine learning model personalized for the user account, wherein the second compression setting corresponds to maintaining the one or more regions of interest in the input image at a high resolution and compressing a remainder of the input image using the baseline compression setting; as well as The input image is compressed based on the second compression setting.
16. The computer-readable medium of claim 15, wherein: The feature detection machine learning model is generated by the following operations: Obtaining a training set of digital images and corresponding features; and The feature detection machine learning model is trained based on the training set and the corresponding features, wherein, after training, the feature detection machine learning model is able to identify image features in the input image provided to the feature detection machine learning model.
17. The computer-readable medium of claim 16, wherein: The user-specific machine learning model is generated by: obtaining a training set of user-specific features associated with the user, the user-specific features indicating user actions with respect to one or more prior images; and The user-specific machine learning model is trained based on the user-specific features and the one or more prior images, wherein, after training, the user-specific machine learning model determines a rank of the input image provided to the user-specific machine learning model.
18. The computer-readable medium of claim 17, wherein: The training set further includes corresponding image features of the one or more prior images.
19. The computer-readable medium of claim 15, wherein: The operations further include: In response to obtaining the user input, a warning is issued predicting that storage space in the user account will be exhausted after a specified number of days or after a predicted number of additional images are captured.
20. The computer-readable medium of claim 15, wherein: Determining the second compression setting includes: determining a rating of the input image by the user-specific machine learning model; and The level is mapped to the second compression setting, wherein the mapping is based on the baseline compression setting.
Citation Information
Patent Citations
Content reproduction system and content storage system
JP2008099012A
Learning user preferences for photo adjustments
US20150098646A1