Video Enhancement
Front-end processing on mobile devices converts video from RGB to YUV color space, addressing storage and bandwidth limitations, enabling efficient cloud-enhanced video quality on smartphones and other devices.
Patent Information
- Application Number
- JP2025572628
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-03
- Filing Date
- 2024-10-03
- Publication Date
- 2026-07-24
AI Technical Summary
The quality of videos captured by smartphones and other client devices is limited by sensor hardware and local image/video processing capabilities, leading to storage and bandwidth constraints, as well as irreversible compression issues that degrade video quality.
Perform front-end processing on the mobile device using an image signal processor to convert video from RGB to YUV color space, reducing file size, and send the processed video to a server for cloud enhancement, which includes noise reduction, stabilization, and other improvements.
Enables high-quality video enhancement on mobile devices with reduced storage and bandwidth demands, preserving useful information and allowing for smooth, detailed video playback.
Smart Images

Figure 2026524810000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a non - provisional application claiming priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 542,285, filed on October 3, 2023, entitled "Video Enhancement", the content of which is hereby incorporated by reference in its entirety.
Background Art
[0002] Smartphones and other client devices are commonly used for video capture. The quality of videos captured by such devices is limited by sensor hardware and local image / video processing capabilities.
[0003] The description of the background art provided herein is intended to present a general overview of the context of the present disclosure. Within the scope described in this background art section, the achievements of the inventors named herein, as well as aspects of this document that may not meet the requirements of prior art at the time of filing, are not to be recognized as prior art to the present disclosure, either explicitly or implicitly.
Summary of the Invention
Means for Solving the Problems
[0004] A method performed by a computer on a mobile device includes receiving a request for enhanced video from a user. The method further includes recording an input video of a scene, the input video having a first format. The method further includes converting the input video to a second format by performing front-end processing and conversion from the red, green, blue (RGB) color space to the YUV color space using the mobile device's image signal processor, the file size of the input video in the second format being smaller than that of the input video in the first format. The method further includes sending the input video in the second format to a server for cloud processing. The method further includes receiving the enhanced video from the server.
[0005] In some embodiments, front-end processing includes one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof. In some embodiments, conversion to the YUV color space includes one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix. In some embodiments, converting the input video to a second format further includes quantizing the input video from a first format encoding the input video in 12 bits to a second format encoding the input video in 10 bits, and interpolating the input video in the first format by adding frames to increase the frames per second (FPS) for the input video in the second format. In some embodiments, the first format is a Bayer image format, and acquiring the input video in the first format includes acquiring camera sensor data from the camera sensor of a mobile device, and converting the input video to a second format includes converting the camera sensor data in Bayer image format to a YUV420 layout. In some embodiments, converting the input video to a second format includes performing swizzling using Y-as-green, RGGB quadrant, RGGB track, or YUV conversion.
[0006] In some embodiments, the method further includes acquiring camera sensor data in Bayer image format from a mobile device's camera sensor, demosaicing the camera sensor data, and binning the camera sensor data. In some embodiments, the method further includes displaying enhanced video playback on the mobile device, receiving a user selection indicating a pause for the enhanced video, and displaying enhanced frames from the enhanced video in a user interface, the user interface including an option to download the enhanced frames. In some embodiments, the method further includes recording a preview video of the scene while recording the input video, and providing an option to view the preview video before receiving the enhanced video from a server, the preview video being associated with a lower quality than the enhanced video. In some embodiments, the method further includes using an image signal processor to perform front-end processing of the preview video, conversion from RGB color space to YUV color space, demosaicing, application of a color correction matrix, and merging long and short frames of the preview video to create a combined frame.
[0007] A non-temporary computer-readable medium stores instructions therein, which, when executed by one or more computers, cause those computers to perform an action. The action includes receiving an enhanced video request from a user and recording an input video of a scene, the input video having a first format, and the action further includes converting the input video to a second format by performing front-end processing and conversion from RGB color space to YUV color space using the image signal processor of a mobile device, the file size of the input video in the second format being smaller than that of the input video in the first format, and the action further includes sending the input video in the second format to a server for cloud processing and receiving the enhanced video from the server.
[0008] In some embodiments, front-end processing includes one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof. In some embodiments, conversion to the YUV color space includes one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix. In some embodiments, converting the input video to a second format further includes quantizing the input video from a first format encoding the input video in 12 bits to a second format encoding the input video in 10 bits, and interpolating the input video in the first format by adding frames to increase the frames per second (FPS) for the input video in the second format. In some embodiments, the first format is a Bayer image format, and acquiring the input video in the first format includes acquiring camera sensor data from the camera sensor of a mobile device, and converting the input video to a second format includes converting the camera sensor data in Bayer image format to a YUV420 layout.
[0009] The system comprises a processor and memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform an action. The action includes receiving an enhanced video request from a user and recording an input video of a scene, the input video having a first format, and further includes converting the input video to a second format by performing front-end processing and conversion from RGB color space to YUV color space using the image signal processor of a mobile device, the file size of the input video in the second format being smaller than that of the input video in the first format, and further includes sending the input video in the second format to a server for cloud processing and receiving the enhanced video from the server.
[0010] In some embodiments, front-end processing includes one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof. In some embodiments, conversion to the YUV color space includes one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix. In some embodiments, converting the input video to a second format further includes quantizing the input video from a first format encoding the input video in 12 bits to a second format encoding the input video in 10 bits, and interpolating the input video in the first format by adding frames to increase the frames per second (FPS) for the input video in the second format. In some embodiments, the first format is a Bayer image format, and acquiring the input video in the first format includes acquiring camera sensor data from the camera sensor of a mobile device, and converting the input video to a second format includes converting the camera sensor data in Bayer image format to a YUV420 layout. [Brief explanation of the drawing]
[0011] [Figure 1] This is a block diagram of an exemplary network environment according to some embodiments described herein. [Figure 2] This is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3] A to D show exemplary user interfaces for obtaining enhanced video according to some embodiments described herein. [Figure 4]This block diagram shows exemplary video stream processing and blocks of different processing stages in which the video can be sent to a server, according to some embodiments described herein. [Figure 5] This block diagram shows the processing of camera sensor data when it is transmitted to a server, according to some embodiments described herein. [Figure 6A] Examples of input and preview video files according to some embodiments described herein are shown below. [Figure 6B] The following are exemplary parameters for the input video file in Figure 6A, according to some embodiments described herein. [Figure 6C] The following are exemplary parameters for the preview video file in Figure 6A, according to some embodiments described herein. [Figure 7A] This is an example of remosaicing of pixels in a Bayer array according to some embodiments described herein. [Figure 7B] This is an example of Bayer array binning according to some embodiments described herein. [Figure 7C] This specification shows combinations of binning and remosaicing according to several embodiments described herein. [Figure 8A] This specification describes exemplary camera image sensors with phase difference capabilities according to several embodiments described herein. [Figure 8B] The following are examples of phase difference (PD) layouts according to some embodiments described herein. [Figure 9A] This specification shows different pixel patterns for Bayer arrays and YUV image formats according to several embodiments described herein. [Figure 9B] This specification shows different pixel patterns for Bayer arrays and YUV image formats according to several embodiments described herein. [Figure 10] This flowchart shows an exemplary method for obtaining enhanced video according to some embodiments described herein.
Mode for Carrying Out the Invention
[0012] Overview The quality of video captured by a mobile device is limited by the sensor hardware and the local image / video processing capabilities. The video can be processed on a server with more computing resources to enhance aspects such as video resolution, color, dynamic range, etc. However, storing the raw video captured by the image sensor of a mobile device on the mobile device can be restrictive due to high storage capacity requirements, energy usage during video capture, and / or storage bandwidth limitations. For example, on some mobile devices, storing the raw sensor data of a video indicates a storage load of 0.5 gigabytes per second on the storage device, which can overwhelm the mobile device. Further, even if the raw video file is stored on the mobile device, sending the raw video from the mobile device to the server can require significant cost and / or time due to the bandwidth requirements for sending the large-sized raw video file. Such transmission can also drain the battery of the mobile device. There are also problems with compressing the raw video file because irreversible compression irreversibly changes the video and causes loss of information essential for improving video quality in post-processing.
[0013] The techniques described herein enhance video favorably by performing lossless processing on input video captured by a camera sensor of a mobile device, such as a smartphone, tablet, wearable device, portable camera, or any other device with a camera. This processing makes it possible to send the processed video to a remote server by providing a video file that is smaller in size than the raw format. For example, in some embodiments, a media application converts the input video from a first format to a second format by performing front-end processing and red, green, and blue (RGB) processing using an image signal processor, which is a dedicated processor (e.g., separate from the device's main processor) that is part of the image processing pipeline, before the video is written to the mobile device's storage device. The remote server receives the input video in the second format (which has a smaller file size than the raw file and retains useful information captured by the image sensor), enhances the video, and sends the enhanced video file back to the mobile device. In some embodiments, video enhancement by the server may include correcting shaky video, grainy video, poorly lit video, and other defective video (e.g., by performing video stabilization). The server provides smooth, detailed, and well-lit enhanced versions of videos for viewing or storing on mobile devices, storing in user accounts hosted by video hosting services associated with the mobile device user, and sharing with other users, all of which involve specific user permissions to access, enhance, and store and / or transmit videos.
[0014] For example, a media application receives a request for an enhanced video from a user. The media application records an input video of a scene, and the input video has a first format. The media application uses an image signal processor of a mobile device to perform front-end processing and conversion from an RGB color space to a YUV color space, thereby converting the input video to a second format, and the file size of the input video in the second format is smaller than that of the input video in the first format. The media application transmits the input video in the second format to a server for cloud processing. The media application receives the enhanced video from the server.
[0015] Exemplary Environment FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101 coupled to a network 105, a mobile device 115a, and a mobile device 115n. Users 125a, 125n may be associated with their respective mobile devices 115a, <115n>. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, the letter after a reference number, e.g., "115a", represents a reference to an element having that particular reference number. A reference number in the text without a subsequent letter, e.g., "115", represents a general reference to embodiments of an element bearing that reference number.
[0016] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicably coupled to the network 105 via signal lines 102. The signal lines 102 may be a wired connection such as Ethernet®, coaxial cable, or fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more mobile devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.
[0017] Database 199 may store machine learning models, training datasets, images, etc. Database 199 may also store social network data associated with user 125, user preferences of user 125, etc.
[0018] The mobile device 115 may be a computing device that includes memory coupled to a hardware processor. For example, the mobile device 115 may include a tablet computer, a mobile phone, a smart device, a wearable device, a head-mounted display, a portable game player, a portable music player, a reader device, or other electronic device that can access the network 105.
[0019] A mobile device may include a camera that includes an image sensor, such as a CMOS / CCD sensor. In some embodiments, the mobile device may include an image signal processor (ISP), such as an application-specific integrated circuit (ASIC) or other type of dedicated processor, coupled to the image sensor. In these embodiments, as further described below, raw image data (e.g., one or more frames of video) captured by the image sensor is provided directly to the ISP (without the involvement of the mobile device's main processor or CPU) for various operations. In some embodiments, the ISP may be dedicated hardware including image / video processing circuitry capable of performing various operations. In some embodiments, the ISP may include a processing unit coupled to memory that stores a set of instructions for the various operations the ISP performs. In some embodiments, the mobile device may implement an image processing pipeline including an image sensor (capturing raw data) and an ISP that performs different processing operations on the captured video frames. In some embodiments, video capture by the mobile device may support multiple modes by different combinations of parameters such as video frame rate, video resolution, and dynamic range. In some embodiments, the ISP may perform specific processing corresponding to a user-selected mode.
[0020] In the illustrated embodiment, mobile device 115a is connected to network 105 via signal line 108, and mobile device 115n is connected to network 105 via signal line 110. Media application 103 may be stored as media application 103b on mobile device 115a and / or as media application 103c on mobile device 115n. Signal lines 108 and 110 may be wired connections such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections such as Wi-Fi®, Bluetooth®, or other wireless technologies. Mobile devices 115a and 115n are accessed by users 125a and 125n, respectively. The mobile devices 115a and 115n in Figure 1 are used as examples. Although Figure 1 shows two mobile devices 115a and 115n, this disclosure applies to system architectures having one or more mobile devices 115.
[0021] The media application 103 may be stored on the media server 101 or the mobile device 115. In some embodiments, the operations described herein are performed on the media server 101 or the mobile device 115. In some embodiments, some operations may be performed on the media server 101 and some on the mobile device 115. The execution of operations is subject to user settings. For example, user 125a may specify that operations are performed on each device 115a and not on the media server 101. With such a setting, the operations described herein are performed entirely on the mobile device 115a and not on the media server 101. Furthermore, user 125a may specify that user images and / or other data are stored locally only on the mobile device 115a and not on the media server 101. With such a setting, user data is not transmitted to or stored on the media server 101. The transmission of user data to the media server 101, the media server 101's arbitrary temporary or permanent storage of such data, and the media server 101's execution of actions on such data will only occur if the user consents to the transmission, storage, and execution of actions by the media server 101. The user will be given the option to change settings at any time, for example, to enable or disable the use of the media server 101.
[0022] The media application 103b of the mobile device 115a receives a request for enhanced video from the user. The media application 103b instructs the camera of the mobile device 115a to record a preview video of the scene and an input video of the scene. The input video is recorded in a first format. The media application 103b converts the input video to a second format, and the file size of the input video in the second format is smaller than that of the input video in the first format. In some embodiments, the user may record the video first and then request the enhanced video after the initial recording. In some embodiments, the user may start recording by providing a command that the enhanced video will be provided to the user.
[0023] Media application 103b sends the input video in a second format to media server 101 for cloud processing. Media application 103a on media server 101 generates an enhanced video. In some embodiments, media application 103a enhances the input video by performing noise reduction, blur reduction, brightness enhancement, three-dimensional stabilization, and / or interpolation to correct blur, poor image quality, insufficient lighting, and other defects in the video.
[0024] While the media server 101 is processing the input video, the media application 103b provides an option to view a preview video. The media application 103a of the media server 101 enhances the video. For example, the media application 103a may perform one or more color corrections, sharpen the image (one or more frames of the video), improve the visibility of scenes where the video is captured under night or low-light conditions, remove or reduce blur, enhance the dynamic range, etc. The media application 103b receives the enhanced video from the media server 101. The media application 103b provides the enhanced video, for example, by adding the enhanced video to the camera roll of the mobile device 115a.
[0025] In some embodiments, the media application 103 may run using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may run using a combination of hardware and software.
[0026] Exemplary computing device Figure 2 is a block diagram of an exemplary computing device 200 that may be used to perform one or more features described herein. The computing device may be any suitable computer system, server, or other electronic or hardware device. In some embodiments, the computing device 200 is the mobile device 115 in Figure 1.
[0027] In some embodiments, the computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, display 241, camera 243, digital signal processor 245, image signal processor 247, and storage device 249, all connected via a bus 218. The processor 235 may be connected to the bus 218 via signal line 222, the memory 237 may be connected to the bus 218 via signal line 224, the I / O interface 239 may be connected to the bus 218 via signal line 226, the display 241 may be connected to the bus 218 via signal line 228, the camera 243 may be connected to the bus 218 via signal line 230, the digital signal processor 245 may be connected to the bus 218 via signal line 232, the image signal processor 247 may be connected to the bus 218 via signal line 234, and the storage device 249 may be connected to the bus 218 via signal line 236.
[0028] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling the basic operation of the computing device 200. “Processor” includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configurations), multiple processing units (e.g., multi-processor configurations), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), composite programmable logic devices (CPLDs), dedicated circuits for implementing functions, dedicated processors for performing processing based on neural network models, neural circuits, processors optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, the processor 235 may include one or more coprocessors for performing neural network processing. In some embodiments, the processor 235 may be a processor that processes data to produce a probabilistic output, for example, the output produced by the processor 235 may be inaccurate or accurate within a range from an expected output. The processing does not need to be limited to a specific geographical location or have temporal constraints. For example, the processor may perform its functions in real time, offline, or batch mode. Multiple parts of the processing may be performed at different times and in different locations by different (or the same) processing systems. The computer can be any processor that communicates with memory.
[0029] Memory 237 can be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erase-read-only memory (EEPROM), or flash memory, typically located separately from and / or integrated with the processor 235, provided in the computing device 200 for access by the processor 235 and suitable for storing instructions for execution by the processor or a set of processors. Memory 237 can store software that runs on the computing device 200 by the processor 235, including media applications 103.
[0030] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, and the like. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application ("App") that runs on a mobile computing device, and so on.
[0031] Application data 266 may be data generated by other applications 264 or the hardware of the computing device 200. For example, application data 266 may include images used by an image library application and user actions identified by other applications 264 (e.g., a social networking application).
[0032] The I / O interface 239 can provide functionality that allows the computing device 200 to interface with other systems and devices. Interface devices can be included as part of the computing device 200 or are separate and capable of communicating with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 249), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can be connected to interface devices such as input devices (keyboards, pointing devices, touchscreens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).
[0033] Some examples of interface devices that can be connected to the I / O interface 239 include a display 241 that can be used to display content, such as images, videos, and / or a user interface for an output application as described herein, and to receive touch (or gesture) input from the user. For example, display 241 may be used to display a user interface, including a graphical guide, in a viewfinder. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, three-dimensional display screen, or other visual display device. For example, display 241 may be a flat display screen provided in a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen in a computer device.
[0034] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 includes multiple lenses, such as a front lens, a main lens, and an ultra-wide-angle lens. Camera 243 includes a camera image sensor (e.g., a CMOS sensor, a CCD sensor, or any sensor that captures light as an image) that captures sensor data transmitted to the digital signal processor 245 and / or the image signal processor 247 via the I / O interface 239.
[0035] In some embodiments, the camera 243 includes a phase-difference (PD) sensor function, where every pixel on the camera image sensor is composed of two parallel diodes under a single lens. In some embodiments, other combinations of lenses and diodes can be used for different PD sensor confirmations.
[0036] The digital signal processor (DSP) 245 includes hardware for converting digital electrical signals into digital output signals. In some embodiments, the digital signal processor 245 measures, filters, or compresses signals from the camera sensor. The digital signal processor 245 receives analog signals from the camera sensor, converts analog signals into digital signals, manipulates digital signals, and converts digital manipulated signals back into analog manipulated signals.
[0037] In some embodiments, the camera 243 may be directly coupled to the ISP 247 and / or DSP 245, bypassing the system bus 218 and processor 235. In these embodiments, image / video frames (raw sensor data) captured by the camera are directly provided to the ISP 247 and / or DSP 245 for processing. The processed video can then be displayed on the display 241 (e.g., preview video) and / or stored in the storage device 249 (e.g., compressed video obtained after processing the input video). In some embodiments, the ISP 247 and / or DSP 245 may include dedicated circuitry for image / video processing of the raw data. In some embodiments, mode selection of image / video capture may cause the ISP 247 and / or DSP 245 to perform a specific set of operations corresponding to the selected mode.
[0038] The image signal processor (ISP) 247 receives camera image sensor data from the camera 243 and performs image processing on the camera image sensor data associated with the video captured by the camera 243. In some embodiments, the ISP 247 receives commands from the media application 103 via the I / O interface 239 and performs one or more of the following on the image data associated with the video: Bayer transformation, demosaicing, noise reduction, and image sharpening. In some embodiments, the ISP 247 may include a multi-camera and frame processor (MCFP) that combines long 12-bit frames and short 12-bit frames to create a single high-dynamic 12-bit frame.
[0039] The storage device 249 stores data related to the media application 103. For example, the storage device 249 may store images, preview videos, input videos in a first format, input videos in a second format, and enhanced videos received from the media server 101.
[0040] Figure 2 shows an exemplary media application 103 stored in memory 237. The media application 103 includes a user interface module 202 and a processing module 204.
[0041] The user interface module 202 generates graphic data for displaying the user interface associated with the camera 243. For example, the user interface may include options for capturing images, options for capturing video, and options for initiating settings to acquire enhanced video.
[0042] The user interface module 202 obtains permission from the user to modify the video, including uploading the video to the server, performing server-side video processing to generate enhanced video, and downloading the enhanced video from the server. The user may be provided with controls that allow them to choose whether and when the systems, programs, or functions described herein can collect user information (e.g., video captured by the user with a camera or video obtained by the user by other means, user preferences, etc.), and whether the user receives content or communications from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, user identity may be processed so that personally identifiable information about the user cannot be determined. Thus, the user can control what information is collected about them, how that information is used, and what information is provided to the user.
[0043] Figures 3A–D illustrate exemplary user interfaces for the process of acquiring enhanced video according to some embodiments described herein. Figure 3A includes a first user interface 300, which includes an image of a mobile device 302 and a hand 304 approaching a settings button 306 (indicating user touch input). The first user interface 300 includes a “Turn on Video Boost” button 308, which, when selected, displays an option to turn on the video boost setting for the video.
[0044] In some embodiments, after the video boost setting is enabled, the user can select video boost whenever they want to acquire enhanced video. In some embodiments, with user permission, video boost may be automatically turned on under certain conditions, such as when recording in low light or in situations where the video is shaky (e.g., due to user camera shake or recording while moving). In some embodiments, the user interface module 202 suggests to the user to turn on video boost in response to certain lighting conditions, such as low light.
[0045] Figure 3B shows a second user interface 325 with video boost enabled. In some embodiments, the second user interface 325 includes a message “Video Boost On” 327 when the user first enables the video boost setting. The second user interface 325 includes a video boost icon 329 that is displayed to signal to the user that an enhanced video will be created based on the video the user has captured. The user starts recording by selecting the record button 331. The second user interface 325 also includes a camera icon 333 and a video icon 335, allowing the user to capture images and videos, respectively. The video icon 335 is highlighted to indicate that the mobile device is in video capture mode. Various user interfaces may provide additional options to allow the user to, for example, select an image / video capture mode, set or adjust the zoom level, control camera settings, etc.
[0046] Once video recording is complete, the high-resolution (4K) video data of the input data is securely transmitted to the media server 101 for processing, with the permission of a specific user.
[0047] Figure 3C shows the third user interface 350 after the video has been captured. The third user interface 350 includes text 352 informing the user that video boost is being prepared and instructing the user to tap the enhanced video icon 354 for more details. Tapping the enhanced video icon 354 may display an estimated time (not shown) for processing the input video and providing the enhanced video. The enhanced video icon 354 is highlighted to indicate that the setting of the delete button 560 will be applied to the enhanced video. While the enhanced video is being prepared, the first frame 358 of the enhanced video is displayed in the third user interface 350. If the user selects the delete button 360, the user interface module 202 notifies the media server 101 to stop generating the enhanced video.
[0048] During recording, the mobile device captures the preview video and input video in a first format used to generate the enhanced video. After the video is recorded, the preview video can be viewed by pressing the preview video button 356.
[0049] Figure 3D shows the fourth user interface 375 after the enhanced video 379 has become available. The enhanced video icon 377 is highlighted, and the enhanced video 379 can be played by pressing the play button 381. In response to the user pressing the play button 381, the user interface module 202 displays the enhanced video playback. The user interface may receive a user selection indicating that the enhanced video is paused. The user interface displays the enhanced frame of the enhanced video and presents the option to download the enhanced frame.
[0050] The user can share the enhanced video by selecting the share button 383, edit the enhanced video by selecting the edit button 385, or delete the enhanced video by selecting the delete button 387. In some embodiments, when the delete button 387 is pressed, the user interface module 202 displays a question regarding whether the user wants to delete the enhanced video only from the mobile device or also from cloud storage. In some embodiments, the user may also be provided with the option to extract individual frames (or portions thereof) from the enhanced video as still images.
[0051] When video image data is captured by the camera image sensor of camera 243, ISP247 processes the image data. Some processing is advantageous for sending the input video to the media server 101 because it reduces the video file size compared to the input video captured by the camera image sensor. However, different processing steps in the image processing pipeline may result in corresponding irreversible changes made to the input video, for example, altering the data captured by the camera image sensor. Such changes may limit the video enhancements that can be performed by the media server 101. As a result, there may be different advantages and disadvantages to choosing when to select a particular processing step for acquiring the video for enhancement and sending it to the media server 101.
[0052] Figure 4 is a block diagram 400 showing exemplary video stream processing and blocks of different processing stages that can transmit the video stream to the media server 101, according to some embodiments described herein. Processing is performed by ISP247.
[0053] The initial video data 405 captured by the camera image sensor can be 8 megapixels (MP) at 30 frames per second (FPS) in Bayer image format and can be encoded in 10 bits. Bayer image format is a color image encoding format for capturing color information from a single sensor. 10 bits refers to the number of bits (bit depth) occupied by the image format. The initial video data has not been processed by ISP247 and is referred to as raw image data.
[0054] Front-end processing 410 includes one or more of the following: linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, and highlight recovery. Linearization occurs when sensor values describing red, blue, and green pixels that form a nonlinear plot are converted to a linear plot, for example, by using an inverse output curve. In some embodiments, linearization includes ISP247 which applies tone reduction, linearization of extreme highlights, and nonlinear compensation in shadows caused by flare. ISP247 may perform black level correction by subtracting a black level offset value from the pixel value. ISP247 may perform digital gain by scaling the pixel values of the red, blue, and green channels using scalar values to improve image exposure. ISP247 may perform green channel imbalance correction by adjusting the gain of green pixels present within the red and blue lines to align the lines more tightly. ISP247 may perform lens shading correction to compensate for distortion caused by the use of spherical lenses. The ISP247 can perform white balance adjustment by calibrating and adjusting color gain to achieve neutral whites and neutral grays in an image. The ISP247 can perform highlight recovery by applying positive luminance correction and moving the exposure slider of the RAW converter to reveal hidden information in the image. The ISP247 can perform highlight recovery by first optimizing the overall tone, then optimizing the highlights, and then combining the two processed versions.
[0055] During front-end processing 410, ISP247 may also create a small buffer for analysis, slightly processed, and a Gaussian pyramid for motion estimation used in staggered high dynamic range (sHDR) frame integration or temporal denoising. In some embodiments, sHDR may allow multiple exposure values to be read from the image sensor, for example, as the camera image sensor continues to be exposed (to light from the scene), short exposure values corresponding to short exposure images may be captured first. Later, second exposure values may be captured for the same scene to obtain longer exposure images.
[0056] In some embodiments, zigzag HDR may be used, where different sensor pixels of the camera image sensor are exposed at different times. This makes it possible to acquire multiple exposures in a single readout. However, in this technique, only a subset of sensor pixels is used for data capture, which can result in lower resolution image data for single exposures.
[0057] In some embodiments, the output video data 415 after front-end processing is 8MP at 30FPS in Bayer image format and encoded with 13 bits. Different frame rates and / or bit depths may be used in various embodiments.
[0058] Red, green, and blue processing (RGBP) 420 involves conversion from the RGB color space to the YUV color space. YUV stands for luminance (i.e., brightness) and color difference, represented by the blue component (U) and the red component (V).
[0059] In some embodiments, RGBP420 includes a first-stage spatial denoising, demosaicing, application of a color correction matrix (CCM), and RGB2YUV, which converts the RGB matrix to a YUV matrix. After RGBP420, the video data is 8MP at 30FPS in the YUV425 image format and encoded in 12 bits. The YUV422 image format is a YCbCr format that can describe any 4:2:2 chroma subsampling format with 8 bits per color sample. The YUV data format shares U and V values between two pixels.
[0060] The multi-camera and frame processor (MCFP) 430 performs motion estimation and expands the dynamic range of the image. After the third level of processing 430, the video data 435 is encoded in 12 bits at 8MP at 30 FPS in YUV image format.
[0061] The additional processing 440 may include one or more of the following: second-stage spatial noise reduction, local tone mapping (HDRnet), sharpening, and color enhancement. Typically, these operations are non-linear and difficult to reverse. Therefore, if these result in information (e.g., texture, detail), loss, or artifacts (e.g., blurring, aliasing), reversing / correcting the processing may require more effort, and the processing may be irreversible. The additional processing 440 may also include one or more of the following: obtaining motion estimation results, applying time filters, performing mesh-based warping on frames, cropping, and scaling frames to match the final target resolution. Mesh-based warping may be used for stabilization, lens distortion correction, focus breathing compensation, and combinations thereof. After the additional processing 440, the video data 445 is 8MP at 30FPS in YUV image format and encoded in 10 bits.
[0062] In some embodiments, video data 405 may be sent to the media server 101 before or after each processing block in Figure 4. If video data 405 is sent to the media server 101 before the front-end processing 410, the ISP247 may swizzle the video data 405 to YUV420 because many video codecs do not support the raw format of video data 405. The swizzling and other processes performed by the ISP247 are described in more detail below with reference to Figures 7-9. YUV420 is a YCbCr format that describes a buffer of any 4:2:0 chroma subsampling planar or semiplanar with 8 bits per color sample.
[0063] The advantages of sending video data 405 at this stage include that video data 405 is the least modified version of the sensor data, there is little processing required from ISP247, clipping to avoid artifacts, as performed by ISP247, is avoided, and testing of the sensor data becomes easier.
[0064] When the video data 415, which is generated after the front-end processing 410, is sent to the media server 101, some common distortions in the raw image are corrected in the video data 415, which may make the data more compressible. However, without denoising, the image may still be noisy under low-light conditions. As above, ISP247 swizzles the video data 415 to YUV420. The advantages of sending the video data 415 at this stage are that some corrections may help with compression efficiency, and that no clipping occurs due to RGBP, MCFP, or other processing blocks.
[0065] When the video data 425 after RGBP420 is sent to the media server 101, the advantages include that sHDR frame fusion has not yet been applied, most nonlinear processes have not yet been applied, the first stage of spatial noise reduction has been applied, and clipping by MCFP has not occurred.
[0066] When the video data 435 after MCFP processing 430 is sent to the media server 101, the advantages include the fact that the size of the video data is halved because long-exposure and short-exposure frames are integrated, and that the first stage of spatial noise reduction is applied.
[0067] When the video data 445 after the fourth level of processing 430 is sent to the media server 101, the advantages include that all ISP noise reduction is applied, testing is easier because the YUV10 stream is acquired via ISP247, and the image format is viewable and shareable without modification (e.g., as a preview video).
[0068] In various embodiments, as illustrated with reference to Figure 4, video data from a particular stage of processing may be sent to a media server 101 for enhancement. In some embodiments, the stage selection may be based on available local processing resources (e.g., ISP247 capabilities), power (device battery level), communication bandwidth to the media server 101, etc. In some embodiments, the stage selection may be further based on user-selected settings (e.g., video capture mode), scene attributes (e.g., low-light scene vs. normal light, scene with significant motion, or static scene with little or no motion), etc.
[0069] The processing module 204 obtains video data of the input video from ISP247 (e.g., video data 405, 415, 425, 435, or 445) and compresses the video data by applying a hardware and / or software codec. The codec compresses the video data by pruning coefficients from a discrete cosine transform (DCT) table of blocks. In some embodiments, this is done by quantizing the coefficients and removing any non-essential values (e.g., zero). The codec may control the compression based on a bitrate that specifies a minimum quantization value per frame (e.g., 0), a maximum quantization value per frame (e.g., 22), and a desired number of bits / bytes to write per second (e.g., 240 megabits per second (Mbps)).
[0070] Applying a compression format reduces the file size of the input video. After the video data is processed and compressed by ISP247, the input video is associated with a second format. The processing module 204 sends the input video associated with the second format to the media server 101.
[0071] Figure 5 is a block diagram of an exemplary flowchart 500 illustrating the image processing of camera sensor data when the camera sensor data is transmitted to the media server 101. During video recording, the camera image sensor 505 captures camera sensor data. The camera sensor data is sent to the ISP 247 for front-end processing 510 and then for RGBP processing 515. In some embodiments, the camera sensor data is split at a tap-out point 517, and the camera sensor data received at the tap-out point 517 is prepared for transmission to the media server 101. The camera sensor data also undergoes MCFP processing 520 to obtain a preview video, which can be accessed locally on a mobile device, e.g., the smartphone or other device that captured the video, while the camera sensor data is used at the media server 101 to generate an enhanced video.
[0072] Camera sensor data may be processed before it is sent to the media server 101 (525). Processing may include converting the camera sensor data from a 12-bit image to a 10-bit image (YUV420-10b image format 530) using a quantization method that rounds values to the nearest corresponding value. In some embodiments, the camera sensor data is converted to a 10-bit image because the encoder used by the media server 101 supports 10-bit images rather than 12-bit images, and 10-bit images have a smaller file size. The source YUV422 image is subsampled to YUV420 using an interpolation / sampling method (530). In some embodiments, the interpolation / sampling method converts the camera sensor data from 30 FPS to 60 FPS by using adjacent frames during interpolation and adding frames to the camera sensor data to transition the camera sensor data to 60 FPS. The 10-bit image is sent to the image reader 535.
[0073] In some embodiments, images are read from a camera using an image reader 535. The images are further processed and used by retrieving a hardware buffer that stores the image data. The images may have the YCBCR_P010 image format, and the hardware buffer format may be YCBCR_P010. The images read from the image reader 535 may be compressed and sent to a media server 101.
[0074] The determination of the stage with a tap-out point 517, where camera sensor data is extracted, compressed, and stored on the mobile device and sent to the server as input video, is based on the time required for ISP247 to process the camera sensor data. The longer the camera sensor data is stored on the mobile device during the video recording process, the more the camera sensor data is processed locally, which may result in irreversible changes being made to the image data that prevent reconstruction of the original sensor values (as illustrated with reference to Figure 4). These changes may manifest as a reduction in detail due to processes such as denoising, clipping of highlights and shadows due to adjustments such as white balance and lens shading, and quantization due to a reduction in bit depth. Placing the tap-out point 517 between RGBP processing 515 and MCFP processing 520 represents a compromise between file size and avoidance of potential irreversible processing.
[0075] Figure 6A shows an example 600 of an input video file 602 and a preview video file 614. The input video file 602 uses the Moving Pictures Expert Group 4 (MP4) container format 604. The input video file 602 includes a RAWish image stream 606, an audio stream 608, frame-by-frame metadata 610, and static metadata 612. In some embodiments, the RAWish image stream 606 is camera sensor data read from the image reader 535 in Figure 5.
[0076] The camera sensor data is referred to as "RAWish" because, although it undergoes some processing by ISP247 (e.g., minimal modification of the raw sensor data), it is similar to the RAW image format. The per-frame metadata 610 may include the frame metadata version, the serialized frame metadata length, the serialized frame metadata, the serialized spatial gain map length, and the serialized spatial gain map. The static metadata 612 may include the version, the serialized static metadata length, and the serialized static metadata.
[0077] The preview video file 614 is unenhanced and is therefore referred to as a 0.8x video. The preview video file 614 also uses an MP4 container 615 and includes a video stream 616 and an audio stream 618. In some embodiments, the order of tracks in the MP4 container 615 (or other video container) may be undefined. Other container types may be used for the input video file 602 and the preview video file 614.
[0078] Figure 6B shows exemplary parameters for the input video file 650 in Figure 6A. The input video file 650 has a bitrate 652 of 240 megabytes per second (Mbps) (653), a quantization parameter (QP) range 654 of 0 to 20 (i.e., a maximum of 20 QP, with no minimum value set) (655), a frame rate 656 of 30 FPS (657), a keyframe rate 658 of 30 FPS (i.e., every frame is encoded to be a keyframe), an image layout 660 of semiplanar YUV420 (661), and a bit depth 662 of 10 bits (663). In some embodiments, the file may be generated at specific stages, such as the different processing stages described in Figure 4.
[0079] Figure 6C shows the preview video file parameters 675 of the preview video file in Figure 6A. The preview video file parameters 675 are: bitrate 676 of 20 Mbps, QP range 678 with no quantization upper limit or lower limit set (679), frame rate 680 of 30 FPS (681), key frame rate 682 of 1 FPS (683), image layout 684 of YUV420 (685), and bit depth 686 of 8 bits (687). In some embodiments, the preview video file may be generated in the MCFP processing step 520 in Figure 5.
[0080] Figure 7A shows an example of remosaic processing of image pixels according to some embodiments described herein. When an image sensor captures image data composed of a quad-Bayer structure or a tetracell structure, the image sensor captures red, blue, and green at each photosite. Because the human eye is more sensitive to green, twice as many green photosites are recorded as blue and green photosites.
[0081] The Bayer array 700 is arranged using four adjacent pixels clustered together by pixels of the same color. Pixels are denoted as R for red, Gr for green pixels adjacent to red pixels, Gb for green pixels adjacent to blue pixels, and B for blue pixels.
[0082] ISP247 can remosaic an image by further subdividing all color pixels (R, Gr, Gb, B) into four subpixels and rearranging the pattern into a higher-resolution Bayer array 725 using a layout in which R, Gr, Gb, and B are arranged alternately. Remosaicing can enhance resolution, reduce blurring, decrease artifacts, and potentially provide image data up to 50 megapixels (MP).
[0083] Figure 7B shows an example of binning of a Bayer array 750 according to several embodiments described herein. In some embodiments, the ISP247 performs binning by combining each quadrant into a single channel to obtain a lower-resolution image. Binning is advantageous for capturing images in low-light conditions and improving quality by combining pixels to create larger pixels. In some embodiments, the size is changed from 50MP to 12MP (because 4 pixels are combined into 1 pixel during binning).
[0084] Using binning alone on an encoded video stream may result in insufficient video zoom resolution. Using remosaicing alone may result in video file sizes that are too large for mobile devices, exceeding the processing capabilities of the codec. In some embodiments, the ISP247 performs both remosaicing and binning. For example, the ISP247 may perform binning on an image sensor and crop to the central region to obtain a 12MP sensor crop with a resolution equivalent to that of remosaicing alone. This can result in a sharper digital zoom quality than using upscaling techniques. Other types of Bayer arrays, such as 5x5 tetracells that produce different remosaicing results, can also be used.
[0085] Figure 7C illustrates a combination of binning and remosaicing according to several embodiments described herein. Figure 7C shows an exemplary portion of image 775 cropped to 12MP and remosaiced in the central region, the original 50MP quad-bayer structure 785, and an exemplary image 795 reduced to 12MP as a result of binning. Image 795 is a lower-resolution image compared to image 785, while image 775 is a 12MP image and is a zoomed-in region of image 785 (indicated by a dotted line in image 785).
[0086] In some embodiments, the ISP247 can achieve high dynamic range (HDR) by combining multiple exposures of a scene into a single shot. For example, the camera may capture long and short shots. However, this increases the exposure time of the image. In some embodiments, the sensor utilizes zigzag HDR, where different sensor pixels are exposed at different times. This makes it possible to acquire multiple exposures in a single readout, although the resolution of individual exposures may be lower.
[0087] In some embodiments, the ISP247 uses staggered HDR (sHDR) to read out multiple exposures simultaneously. The sensor continues to be exposed while sensor data is being read out for a single exposure. The ISP247 can perform other simultaneous readouts for longer exposure images.
[0088] In some embodiments, a multicamera and frame processor (MCFP) combine long 12-bit frames and short 12-bit frames to create a single high-dynamic 12-bit frame.
[0089] In some embodiments, the camera 243 includes a phase-difference (PD) sensor function, where every pixel on the sensor consists of two side-by-side diodes located beneath a single lens. The sensor acquires two values per pixel, each measuring a different phase (or directionality) of incident light. Figure 8A shows an exemplary camera image sensor 800 with a phase-difference function according to some embodiments described herein. The camera image sensor 800 includes a lens 805, two diodes 807a and 807b, and two diodes 809a and 809b. The PD signal is useful for autofocus and, more generally, provides data about the distance from an object to the sensor. The PD signal is useful in techniques where depth-of-field (or bokeh) effects are applied to images or videos.
[0090] Figure 8B shows the types of PD layouts according to some embodiments described herein. In some embodiments, the camera utilizes a sparse PD layout 825 in which only a portion of the pixels on the sensor measure phase difference. A front camera (e.g., on the same side as the user facing the primary display of a smartphone or other device) may use a dual PD layout 835 in which all pixels on the sensor have two diodes to measure phase. An ultra-wide-angle camera (e.g., a second camera in a smartphone on the opposite side of the device's primary display) may use a quad PD layout 845 in which all pixels have four diodes to measure phase difference in both horizontal and vertical directions. A main camera on the same side as the ultra-wide-angle camera may use an octa PD layout 855 (e.g., a 4x2 pattern) in which all subpixels of a quad Bayer array have two aligned diodes.
[0091] In some embodiments, the input video image format is a single 10-bit image format that can be compressed by a hardware codec called YUVP010. YUVP010 may be a YUV420 semiplanar layout, where the U and V chroma channels are subsampled 4:1 with reference to luminance (Y). Different tap-out points during ISP processing are in different formats, except for the final tap-out point. As a result, image data acquired from ISP247 can be converted to this image format.
[0092] Figures 9A and 9B show different pixel patterns for Bayer arrays and YUV image formats according to some embodiments described herein. The YUV420 image format has a larger capacity than the RAW image format for the same dimensions and bit depth. The Bayer array 900 of RAW data contains all color data in a single width × height (W × H) plane of alternately arranged pixels, whereas the YUV pattern 905 includes a grayscale W × H plane (Y) followed by a chroma plane (U, V) with half the width and height (W / 2, H / 2).
[0093] Swizzling is used to reinterpret raw data (e.g., with four channels in RGGB format) as a three-channel YUV image. Swizzling can take several forms. For example, swizzling from Bayer array 910 to YUV915 uses the Y-as-green technique, where the Y channel of the YUV is the pixel value G R and G B The U and V channels are used to store the pixel values red and blue (R, B) from the Bayer array. Y-as-green techniques may require additional calculations for interpolation, but the colors are more natural and compression is easier.
[0094] In other examples, swizzling from a Bayer array 920 to YUV925 using the RGGB quadrant may be used. In this example, the Y channel is obtained from the pixel values G of the Bayer array. R , G B It has four quadrants, one for each of R, B, and R. In this example, the U and V channels are set to zero.
[0095] In another example, swizzling from a Bayer array 930 to YUV935 using RGGB tracks may be used, as shown in Figure 9B. In this example, the pixel values from the Bayer array are divided into four tracks (T1-T4), with each track containing the pixel values G from the Bayer array. R , G B It corresponds to R, and B. The Y channel of each track stores the pixel value, while the U and V channels are set to zero.
[0096] Finally, another example shows the conversion from YUV940 to YU'V'945.
[0097] Exemplary flowchart Figure 10 shows an exemplary flowchart for acquiring enhanced video. Method 1000 can be performed by the computing device 200 in Figure 2. In some embodiments, Method 1000 is performed by the mobile device 115 in Figure 1.
[0098] Method 1000 in Figure 10 may begin at block 1002. In block 1002, it is determined whether user permission has been obtained from the user to generate enhanced video. If permission has not been obtained, the method may end at block 1004, and no processing for generating enhanced video is performed. In this case, the captured video is stored locally on the user's device but is not sent to a server or other device for video enhancement. If user permission is obtained, block 1006 may follow block 1002.
[0099] In block 1006, a request for enhanced video is received from the user. The user can request enhanced video during recording or select preferences for the input video to be automatically converted to enhanced video. Block 1008 may follow block 1006.
[0100] In block 1008, the input video of the scene is recorded, and the input video has a first format. Block 1010 may follow block 1008.
[0101] In block 1010, the input video is converted to a second format by performing front-end processing and conversion from the RGB color space to the YUV color space using the image signal processor of the mobile device 115, the size of the input video in the second format being smaller than that of the input video in the first format. The front-end processing may include one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof. The conversion to the YUV color space may include one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix.
[0102] In some embodiments, converting an input video to a second format further includes quantizing the input video from a first format encoding the input video in 12 bits to a second format encoding the input video in 10 bits, and interpolating the input video in the first format by adding frames to increase the frames per second (FPS) for the input video in the second format. In some embodiments, the first format is a Bayer image format, and acquiring an input video in the first format includes acquiring camera sensor data from a mobile device's camera sensor, and converting the input video to a second format includes converting the camera sensor data in Bayer image format to a YUV420 layout. In some embodiments, converting an input video to a second format includes swizzling using Y-as-green, RGGB quadrant, RGGB track, or YUV conversion. In some embodiments, method 1000 further includes acquiring camera sensor data in Bayer image format from a mobile device's camera sensor, remosaicing the camera sensor data, and binning the camera sensor data. Block 1010 may be followed by Block 1012.
[0103] In block 1012, the input video in the second format is sent to a server (e.g., media server 101) for cloud processing. Block 1014 may follow block 1012.
[0104] In block 1014, the enhanced video is received from the server. In some embodiments, the method further includes displaying the enhanced video on a mobile device, receiving a pause of the enhanced video, and displaying enhanced frames from the enhanced video via a user interface, the user interface including an option to download the enhanced frames. In some embodiments, the method further includes providing a preview video of lower quality than the enhanced video before the enhanced video is received, in response to the termination of recording of the input video.
[0105] In some embodiments, Method 1000 further includes displaying enhanced video playback on a mobile device, receiving a user selection indicating a pause for the enhanced video, and displaying enhanced frames from the enhanced video in a user interface, the user interface including an option to download the enhanced frames. In some embodiments, Method further includes recording a preview video of a scene while recording an input video, and providing an option to view the preview video before receiving the enhanced video from a server, the preview video being associated with a lower quality than the enhanced video. In some embodiments, Method further includes using an image signal processor to perform front-end processing of the preview video, conversion from RGB color space to YUV color space, demosaicing, application of a color correction matrix, and merging long and short frames of the preview video to create a combined frame.
[0106] In some embodiments, blocks 1002 and 1004 / 1006 may be performed during the initial setup of the media application 103, where the user indicates whether video enhancements should be enabled, as illustrated with reference to Figure 3. The user may change their preferences at any time, which may be supported by additional execution of blocks 1002 and 1004 / 1006.
[0107] In some embodiments, the user may record the video and later select enhancement options. In these embodiments, block 1006 may occur after blocks 1008-1014.
[0108] The above description includes many specific details for illustrative purposes to provide a complete understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In some cases, structures and devices are shown in block diagram form to avoid obscuring this specification. For example, embodiments may be described above with reference primarily to user interfaces and specific hardware. However, embodiments can be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.
[0109] Any reference in this specification to “some embodiments” or “some examples” means that certain features, structures, or characteristics described in relation to the embodiments or examples may be included in at least one embodiment of this specification. The phrase “in some embodiments” appearing in various places in this specification does not necessarily refer to the same embodiments.
[0110] Some parts of the detailed explanation above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These descriptions and representations of algorithms are means used by those skilled in the art to most effectively communicate the content of their work to others skilled in the art. Here, and also generally, an algorithm is considered to be a self-consistent set of steps that lead to a desired result. These steps are steps that require the physical manipulation of physical quantities. Usually, though not essential, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, or otherwise manipulated. For reasons of general use, it is sometimes convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.
[0111] However, it should be recognized that all these terms and similar terms should correspond to appropriate physical quantities and are merely convenient labels applied to those quantities. As will become clear from the following discussion, unless otherwise specifically stated, throughout this specification, discussions using terms such as “process,” “calculate,” “compute,” “determine,” or “display” refer to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memory of the computer system and convert it into other data similarly represented as physical quantities in the memory or registers of the computer system, or other such information storage devices, transmission devices, or display devices.
[0112] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-temporary computer-readable storage medium, including but not limited to any type of disk including an optical disk, ROM, CD-ROM, magnetic disk, RAM, EPROM, EEPROM, magnetic card or optical card, flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0113] The specification may take the form of several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, the specification is implemented with software including, but not limited to, firmware, resident software, and microcode.
[0114] Furthermore, the specification may take the form of a computer program product accessible from a computer-enabled medium or computer-readable medium, which provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this specification, the computer-enabled medium or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, instruction execution unit, or instruction execution device.
[0115] A data processing system suitable for storing or executing program code would include at least one processor directly or indirectly connected to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory providing temporary storage for at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.
Claims
1. A method performed on a mobile device and executed by a computer, Receiving requests for enhanced videos from users, The method includes recording an input video of a scene, wherein the input video has a first format, and the method further includes The method includes converting the input video to a second format by performing front-end processing and conversion from the red, green, and blue (RGB) color space to the YUV color space using the image signal processor of the mobile device, wherein the file size of the input video in the second format is smaller than that of the input video in the first format, and the method further includes Sending the input video in the second format to a server for cloud processing, A method comprising receiving the enhanced video from the server.
2. The method according to claim 1, wherein the front-end processing includes one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof.
3. The method according to claim 1, wherein the conversion to the YUV color space includes one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix.
4. Converting the aforementioned input video to the second format is The input video is quantized from a first format that encodes the input video in 12 bits to a second format that encodes the input video in 10 bits, The method according to claim 1, further comprising interpolating the input video in the first format by adding frames in order to increase the frames per second (FPS) for the input video in the second format.
5. The first format mentioned above is the Bayer image format, Acquiring the input video in the first format includes acquiring camera sensor data from the camera sensor of the mobile device, The method according to claim 1, wherein converting the input video to the second format includes converting the camera sensor data in the Bayer image format to a YUV420 layout.
6. The method according to claim 5, wherein converting the input video to the second format includes swizzling using Y-as-green, RGB quadrant, RGB track, or YUV conversion.
7. The acquisition of camera sensor data in Bayer image format from the camera sensor of the aforementioned mobile device, The camera sensor data is remosaiced, The method according to claim 1, further comprising performing binning of the camera sensor data.
8. Displaying the enhanced video playback on the mobile device, Receiving a user selection indicating the pause of the enhanced video, The method according to claim 1, further comprising displaying enhanced frames from the enhanced video in a user interface, wherein the user interface includes an option to download the enhanced frames.
9. While recording the aforementioned input video, a preview video of the aforementioned scene is recorded, The method according to claim 1, further comprising providing an option to view the preview video before receiving the enhanced video from the server, wherein the preview video is associated with a lower quality than the enhanced video.
10. The method according to claim 9, further comprising using the image signal processor to perform front-end processing of the preview video, conversion from the RGB color space to the YUV color space, demosaicing, application of a color correction matrix, and merging long and short frames of the preview video to create a merged frame.
11. A non-temporary computer-readable medium that, when executed by one or more computers, stores instructions that cause the one or more computers to perform an action, wherein the action is: Receiving requests for enhanced videos from users, The operation includes recording an input video of a scene, wherein the input video has a first format, and the operation further includes, The process includes converting the input video to a second format by performing front-end processing and conversion from the red, green, and blue (RGB) color space to the YUV color space using the image signal processor of a mobile device, wherein the file size of the input video in the second format is smaller than that of the input video in the first format, and the process further includes: Sending the input video in the second format to a server for cloud processing, A non-temporary computer-readable medium, including receiving the enhanced video from the server.
12. The non-temporary computer-readable medium according to claim 11, wherein the front-end processing includes one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof.
13. The non-temporary computer-readable medium according to claim 11, wherein the conversion to the YUV color space includes one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix.
14. Converting the aforementioned input video to the second format is The input video is quantized from a first format that encodes the input video in 12 bits to a second format that encodes the input video in 10 bits, The non-temporal computer-readable medium according to claim 11, further comprising interpolating the input video of the first format by adding frames in order to increase the number of frames per second (FPS) for the input video of the second format.
15. The first format mentioned above is the Bayer image format, Acquiring the input video in the first format includes acquiring camera sensor data from the camera sensor of the mobile device, The non-temporary computer-readable medium according to claim 11, wherein converting the input video to the second format includes converting the camera sensor data in the Bayer image format to a YUV420 layout.
16. It is a system, Processor and The system includes a memory coupled to the processor, the memory storing instructions that cause the processor to perform an action when executed by the processor, and the action is Receiving requests for enhanced videos from users, The operation includes recording an input video of a scene, wherein the input video has a first format, and the operation further includes, The process includes converting the input video to a second format by performing front-end processing and conversion from the red, green, and blue (RGB) color space to the YUV color space using the image signal processor of a mobile device, wherein the file size of the input video in the second format is smaller than that of the input video in the first format, and the process further includes: Sending the input video in the second format to a server for cloud processing, A system including receiving the enhanced video from the server.
17. The system according to claim 16, wherein the front-end processing includes one or more actions selected from the group consisting of linearization, black level correction, digital gain, green channel imbalance correction, lens shading correction, white balance adjustment, highlight recovery, and combinations thereof.
18. The system according to claim 16, wherein the conversion to the YUV color space includes one or more actions selected from the group consisting of spatial noise reduction, demosaicing, application of a color correction matrix, and conversion from an RGB matrix to a YUV format matrix.
19. Converting the aforementioned input video to the second format is The input video is quantized from a first format that encodes the input video in 12 bits to a second format that encodes the input video in 10 bits, The system according to claim 16, further comprising interpolating the input video in the first format by adding frames in order to increase the number of frames per second (FPS) for the input video in the second format.
20. The first format mentioned above is the Bayer image format, Acquiring the input video in the first format includes acquiring camera sensor data from the camera sensor of the mobile device, The system according to claim 16, wherein converting the input video to the second format includes converting the camera sensor data in the Bayer image format to a YUV420 layout.