Image frame preprocessing based on camera statistics

Through a video preprocessing system based on camera statistics, selectively identifying and providing image frame subsets, the problem of inefficiency in high-quality images and video processing in the prior art is solved, and more efficient and accurate image processing output is achieved.

CN113557522BActive Publication Date: 2025-08-29MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080020602.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-11
Filing Date
2020-03-02
Publication Date
2025-08-29
Estimated Expiration
2040-03-02

AI Technical Summary

Technical Problem

Existing media processing applications are inefficient and consume a lot of computing resources when processing high-quality images and videos, resulting in exhaustion of client computing devices and increased cloud computing spending.

Method used

Through a video preprocessing system based on camera statistics, input video is received from the video capture device, image frames are identified and provided to the image processing model, and the image frame subset is selectively identified by the camera statistics to reduce the consumption of processing resources and improve processing efficiency.

Benefits of technology

The image processing model is realized to produce more efficient and accurate output when processing images and videos, while reducing the consumption of processing resources and enhancing the effectiveness of the image processing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113557522B_ABST
    Figure CN113557522B_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems, methods, and computer-readable media for selectively identifying image frames from an input video to provide to an image processing model based on camera statistics. For example, the system disclosed herein includes receiving an input video and associated camera statistics from a video capture device. The system disclosed herein also includes identifying, based on the camera statistics and based on the application of an image processing model, select image frames to provide to the image processing model. The system disclosed herein also includes selectively identifying and providing the camera statistics to the image processing model. By selectively providing data to the image processing model based on the camera statistics, the system disclosed herein can leverage the capabilities of the video capture device to significantly reduce the expenditure of processing resources when utilizing various image processing models.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In recent years, the use of computing devices (e.g., mobile devices, personal computers, server devices) to capture, store, and edit digital media has increased dramatically. Indeed, it is now common for electronic devices to capture and process digital media in various ways. For example, conventional media systems typically include various applications or tools for processing digital media. These media processing applications provide a wide range of uses in processing images and videos.

[0002] Nevertheless, while media processing applications provide useful tools for analyzing digital media and generating useful output, these applications and tools include various problems and shortcomings. For example, many media processing applications are inefficient and / or consume a large amount of processing resources to operate effectively. Indeed, with video capture devices capturing and storing higher-quality images than ever before, conventional applications require a large amount of computing resources and processing time to successfully execute the application. For example, media processing applications that utilize machine learning techniques may exhaust the processing power of client computing devices and result in significant cloud computing expenditures. Furthermore, conventional media processing applications may take a significant amount of time to produce the desired results.

[0003] These and other issues exist regarding the use of various applications and software tools to analyze and process digital media. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 An example environment including a statistics-based video pre-processing system is shown in accordance with one or more implementations.

[0005] Figure 2A An example process for identifying an image frame and providing the image frame to an image processing model is shown in accordance with one or more implementations.

[0006] Figure 2B

[0014] An example process for identifying image frames and camera statistics and providing the image frames and camera statistics to an image processing model is shown in accordance with one or more implementations.

[0007] Figure 3

[0014] Example processes for selectively identifying and providing image frames from multiple video feeds are shown in accordance with one or more implementations.

[0008] Figure 4 An example workflow for transforming video content and providing a subset of image frames to an image processing model is shown in accordance with one or more implementations.

[0009] Figure 5 A schematic diagram illustrating an example computing device including a statistics-based video pre-processing system in accordance with one or more implementations is shown.

[0010] Figure 6

[0014] Example methods of selectively providing image frames to an image processing model in accordance with one or more implementations are shown.

[0011] Figure 7 Another example method of identifying image frames and camera statistics and providing the image frames and camera statistics to an image processing model in accordance with one or more implementations is shown.

[0012] Figure 8 Some components that may be included in a computer system are shown. DETAILED DESCRIPTION

[0013] The present disclosure relates to a statistics-based video preprocessing system (or simply, "preprocessing system") implemented in conjunction with an image processing model. Specifically, as discussed in further detail below, the preprocessing system can receive an input video from one or more video capture devices. The preprocessing system can also identify camera statistics for the content of the input video (or simply, "input video content") to determine one or more operations to be performed on the input video content before providing one or more image frames to the image processing model, which can include a deep learning model.

[0014] For example, the pre-processing system can analyze input video content received from the camera device in light of camera statistics also received from the camera device. Specifically, the pre-processing system can utilize camera statistics (such as measurements of camera focus, white balance, lighting conditions, detected objects, or other statistics obtained by the video capture device) to identify a subset of image frames from the input video to feed to the image processing model. For example, the pre-processing system can selectively identify a subset of image frames based on the camera statistics to produce more accurate or more useful output and utilize fewer processing resources when applying the image processing model to the subset of images.

[0015] In addition to identifying images from the input video and providing the images to the image processing model, the preprocessing system can also provide the image processing model with camera statistics obtained by the video capture device. When the image processing model is processed based on the images and associated camera statistics training, the image processing model can generate useful output more accurately and / or efficiently based on the specific application or function of the image processing model. In addition to enabling the image processing model to accurately and efficiently generate useful output, preprocessing the input video according to one or more implementations described herein enables various types of image processing models to process data and generate useful output while utilizing fewer resources than conventional media processing systems.

[0016] The present disclosure includes numerous practical applications that provide benefits and / or address issues associated with analyzing and processing images. For example, as described above, by intelligently selecting a subset of image frames from input video content, the preprocessing system enables an image processing model to more efficiently generate output based on the specific application or function of the image processing model while using fewer processing resources. Furthermore, by preprocessing input video content in various ways, the preprocessing system similarly enhances the utility of the image processing system while reducing the consumption of processing resources.

[0017] In addition, by identifying image frames and preprocessing video content based on camera statistics, the preprocessing system can take advantage of information readily available using the built-in functionality of the video capture device. Specifically, where the video capture device has already used camera statistics to refine video footage and generate input video, the preprocessing system can leverage these statistics to strategically feed image data to the image processing model in a more efficient manner. Additionally, the preprocessing system can enable the camera statistics obtained from the video capture device to be provided as input to the image processing model to enhance the functionality and accuracy of the image processing model for many different applications.

[0018] As indicated in the foregoing discussion, the present disclosure utilizes various terms to describe the features and advantages of the pre-processing system. Additional details regarding the meaning of these terms will now be provided. For example, as used herein, a "video capture device" refers to an electronic device capable of capturing video clip footage and generating input video content. A video capture device may refer to a standalone device that communicates with a computing device. Alternatively, a video capture device may refer to a camera or other video capture device that is integrated within a computing device. In one or more implementations, a video capture device captures video clip footage (e.g., unrefined or raw video content) and refines the video clip footage using any number of camera statistics to generate an input video comprising a plurality of image frames (e.g., a plurality of images representing a video), which may include a refined video according to the camera statistics.

[0019] As used herein, "camera statistics" refers to various metrics and measurements associated with capturing video footage and generating input video content. Camera statistics may refer to any features or measurements acquired or generated by a video capture device and / or an application running on the video capture device. For example, camera statistics may refer to specifications and characteristics of a camera capture device, such as the resolution of images captured by the video capture device, detected device movement, and / or the configuration (e.g., zoom, orientation) of one or more lenses used in conjunction with capturing video footage. In addition to device-related characteristics, camera statistics may refer to heuristic data associated with operations performed by one or more applications running on the video capture device to transform, refine, or otherwise modify images. For example, camera statistics may refer to focus measurements, white balance measurements, lighting conditions, scene detection statistics, object detection, or other metrics that one or more applications on the video capture device can identify and provide to the pre-processing system. In addition, in one or more embodiments, camera statistics may include features or characteristics of the video content. For example, camera statistics may include content characteristics such as the resolution of individual image frames comprising the video content, the frame rate of the video content (e.g., frames per second), and / or the display ratio of the image frames.

[0020] As described above, the pre-processing system can provide video frames (e.g., a subset of video frames, transformed video frames) to the image processing model. As used herein, "image processing model" refers to any model that is trained to generate an output based on one or more input image frames. An image processing model can refer to one or more of a computer algorithm, a classification model, a regression model, an image transformation model, or any type of model with a corresponding application or defined functionality. Additionally, an image processing model can refer to a deep learning model, such as a neural network (e.g., a convolutional neural network, a recurrent neural network), or other machine learning architecture that is trained to perform various applications based on an input image. In one or more implementations, the image processing model is trained to generate an output based on both the input image and associated camera statistics.

[0021] As used herein, the "output" of an image processing model can refer to any type of output based on the type of image processing model or the application implemented by the image processing model. For example, where the image processing model refers to a classification model, the output of the image processing model can include a classification of one or more images, such as whether a face is detected, the identification of an individual associated with a face, the identification of an object within (multiple) images, a rating of an image or video, or any other classification of one or more image frames. As another example, where the image processing model includes a quick read (QR) code or barcode reading application, the output can refer to an output image that includes a clear representation of the QR code, barcode, and / or a decoded value extracted from (multiple) images with the displayed code. As a further example, where the image processing model includes an optical character recognition (OCR) application, the output can include text data, character data, or other data generated based on analysis of the image content provided to the image processing model. In fact, it should be understood that the output can refer to any desired output corresponding to multiple applications and processing models generated based on one or more input parameters (e.g., images, camera statistics).

[0022] Additional details regarding the pre-processing system will now be provided in conjunction with illustrative drawings depicting example implementations. For example, Figure 1 An example environment 100 is shown for pre-processing input video to identify image frames and providing the image frames to an image processing model. Figure 1 As shown, environment 100 includes one or more server devices 102, which include image processing models 104. In addition, environment 100 includes computing device 106, which includes a statistics-based video pre-processing system 108 (or simply "pre-processing system 108"). Environment 100 also includes video capture device 110.

[0023] like Figure 1 As shown, (multiple) server devices 102 and computing devices 106 can communicate with each other directly or indirectly via a network 112. Network 112 can include one or more networks and can use one or more communication platforms or technologies suitable for transmitting data. Network 112 can refer to any data link that enables the transmission of electronic data between devices and / or modules of environment 100. Network 112 can refer to a hardwired network, a wireless network, or a combination of hardwired and wireless networks. In one or more embodiments, network 112 includes the Internet.

[0024] Computing device 106 may refer to various types of computing devices. For example, computing device 106 may include a mobile device, such as a mobile phone, a smartphone, a personal digital assistant (PDA), a tablet computer, or a laptop computer. Additionally or alternatively, computing device 106 may include a non-mobile device, such as a desktop computer, a server device, or other non-portable device. (Multiple) server devices 102 may similarly refer to various types of computing devices. Each of computing device 106 and (multiple) server devices 102 may include the following in combination: Figure 8 Describe the features and functionality.

[0025] Furthermore, video capture device 110 may refer to any type of camera or other electronic device capable of capturing video footage and providing the generated input video and associated camera statistics to pre-processing system 108. In one or more embodiments, video capture device 110 is a standalone digital camera or other video capture device that includes video capture capabilities. Alternatively, in one or more embodiments, video capture device 110 is integrated within computing device 106.

[0026] As will be discussed in further detail below, the video capture device 110 can capture a video clip shot and apply a plurality of camera statistics to the video clip shot to generate an input video having a plurality of image frames. For example, while capturing the video clip shot, the video capture device 110 can focus the image, adjust the white balance, and apply one or more settings to compensate for lighting or other ambient conditions. Additionally, the video capture device 110 can use various tools to analyze the captured content to detect scenes, detect objects, or otherwise identify specific types of content within the video clip shot. Furthermore, the video capture device 110 can perform one or more operations on the captured content, including, for example, fusing multiple video frames together or enhancing one or more captured frames of the video content.

[0027] The video capture device 110 can provide input video content to the computing device 106 for pre-processing based on the camera statistics. For example, after using various camera statistics to refine the captured video clip footage, the video capture device 110 can provide the video content to the computing device for further processing. In one or more embodiments, the video capture device 110 provides a video stream (e.g., a live video stream) as the video capture device 110 captures and refines the video clip footage. Alternatively, the video capture device 110 can provide a video file to the computing device 106.

[0028] In addition to providing input video content to computing device 106, video capture device 110 may also provide any number of camera statistics to computing device 106. For example, video capture device 110 may provide a collection of all camera statistics acquired while capturing and generating the input video content. The camera statistics may include a file of camera statistics corresponding to the video content. Alternatively, video capture device 110 may provide the camera statistics as part of a digital video file (e.g., as metadata for the video file). In one or more embodiments, when video capture device 110 provides input video to computing device 106, video capture device 110 provides camera statistics associated with corresponding frames of the input video content.

[0029] As described above, and as will be discussed in further detail through the examples below, the camera statistics may include any number of different statistics depending on the characteristics of the input video and the features and capabilities of the video capture device 110. For example, the video capture device 110 may obtain and provide camera statistics including an indication of image frames that are focused (or a measure of focus with respect to various image frames), an indication of the white balance of one or more image frames, a measure of lighting conditions detected by the video capture device 110, an identification of which image frames correspond to scene changes, the resolution of the image frames comprising the video content, the frame rate of the video content (e.g., the frames per second of the input video content), identification of objects and associated frames in which one or more objects appear, and information regarding the merging of multiple image frames together in generating the input video.

[0030] In one or more embodiments, the pre-processing system 108 identifies one or more camera statistics for use in identifying image frames to provide to the image processing model 104. For example, the pre-processing system 108 may identify a portion or subset of camera statistics from the collection of all camera statistics provided by the video capture device 110. The pre-processing system 108 may identify relevant camera statistics based on application of the image processing model 104. As another example, the pre-processing system 108 may identify camera statistics based on which image frames are provided to the image processing model 104 (discussed below).

[0031] As will be discussed in further detail below, the pre-processing system 108 can identify image frames and camera statistics to provide to the image processing model 104. For example, the pre-processing system 108 can identify a subset of image frames from a plurality of image frames representing all frames of the input video to provide as input to the image processing model 104. As another example, the pre-processing system 108 can perform additional processing on the input video to generate one or more transformed or otherwise modified image frames to provide as input to the image processing model 104. As a further example, the pre-processing system 108 can identify any camera statistics to provide to the image processing model 104 in conjunction with providing the corresponding image frames to the image processing model 104.

[0032] After receiving model input data (e.g., image frames, camera statistics), the image processing model 104 may apply one or more applications and / or algorithms of the image processing model 104 to the input image frames (and / or camera statistics) to generate an output. For example, the image processing model 104 may generate one or more classifications, output images, decoded data, extracted text, or other outputs based on the training of the image processing model 104 to generate the desired output.

[0033] Although Figure 1 An example environment 100 is shown that includes a particular number and arrangement of server device(s) 102, computing device 106, and video capture device 110, but it should be understood that the environment 100 can include any number of devices, including the image processing model 104 and pre-processing system 108 implemented on the same device network and / or across multiple devices, such as Figure 1 For example, in one or more embodiments, the image processing model 104 is implemented on a cloud computing system including (a plurality of) server devices 102. Alternatively, in the case of communication between modules or internal components of a single computing device, the image processing model 104 can be implemented on an edge device and / or on the same device as the pre-processing system 108 and / or the video capture device 110.

[0034] Go to Figure 2A , Figure 2A An example framework for selectively identifying image frames to provide to the image processing model 104 according to one or more embodiments is shown. Figure 2AAs shown, video capture device 110 can capture video clip footage 202. Video clip footage 202 can include visual data captured in the visible light spectrum, such as red, green, and blue (RGB) data, or visual data captured in the invisible light spectrum, such as infrared data. Video clip footage 202 can also include depth data. Video capture device 110 can capture the footage and generate video content at a specific frame rate and resolution based on the specifications and capabilities of video capture device 110. Additionally, video capture device 110 can capture the footage and generate video content at a specific display ratio based on the specifications and capabilities of the video capture device.

[0035] The video capture device 110 can generate an input video 204 based on an input video clip shot 202. Specifically, the video capture device 110 can generate the input video 204 by transforming, refining, or otherwise modifying the incoming video clip shot 202. While processing the video clip shot to generate the input video 204, the video capture device 110 can track or otherwise collect a number of camera statistics, such as specifications and characteristics of the video capture device 110, heuristic data regarding the transformations or other modifications performed on the video clip shot 202 in generating the input video 204, and identification of content within image frames comprising the input video 204 (e.g., detected objects, scene changes). Additionally, the video capture device 110 can identify camera statistics such as depth data, lens type (e.g., fisheye lens), or various simple scalars (e.g., exposure, ISO measurement, center focus quality). The video capture device 110 can further identify metrics such as vectors (e.g., saturation of primary colors, focus measurements in multiple regions of the viewport). The video capture device 110 can further identify a spatial map (e.g., depth quality).

[0036] like Figure 2A As shown, the video capture device 110 can provide both the input video 204 and the associated camera statistics 206 to the pre-processing system 108. The pre-processing system 108 can perform a number of actions associated with the input video 204 and the camera statistics 206. For example, the pre-processing system 108 can identify a subset of the camera statistics 206 that are relevant to a particular application of the image processing model 104. For example, where the image processing model includes a deep learning model trained to identify or classify various types of objects shown within digital images, the pre-processing system 108 can selectively identify camera statistics associated with image sharpness (e.g., focus statistics) and camera statistics associated with identification of frames that include one or more detected objects to further analyze or process image frames of the input video 204.

[0037] The pre-processing system 108 can utilize the identified statistics provided from the video capture device 110 to identify a subset of image frames 208 from the plurality of image frames of the input video 204. Specifically, using the camera statistics, the pre-processing system 108 can selectively identify a subset of image frames 208 that includes images that are in focus and / or images that include detected objects based on the associated camera statistics, which images will provide more useful data to the image processing model 104 when classifying the image frames and / or content shown in the image frames.

[0038] The pre-processing system 108 may select any number of image frames to provide to the image processing model 104. In one or more embodiments, the pre-processing system 108 identifies the number of image frames based on the ability of the computing device (e.g., the server device 102) to apply the image processing model 104 at a particular frame rate. For example, if the image processing model 104 is capable of analyzing two image frames per second, the pre-processing system 108 may provide fewer image frames to the image processing model 104 than if the server device(s) 102 are capable of applying the image processing model 104 to ten image frames per second.

[0039] As another example, the pre-processing system 108 may identify a number of image frames or a rate of image frames based on the complexity and application of the image processing model 104. For example, the pre-processing system 108 may provide a greater number or rate of image frames for analysis to a less complex image processing model 104 (e.g., a simple algorithm) than if the image processing model 104 were more complex (e.g., a complex neural network or deep learning model). In one or more embodiments, the pre-processing system 108 determines the rate or number of image frames to provide to the image processing model 104 based on the processing power of the computing device that includes the image processing model 104 and based on the complexity of the image processing model 104 and / or the application of the image processing model 104.

[0040] When identifying a subset of image frames 208 to provide to the image processing model 104, the pre-processing system 108 may identify video frames at corresponding frame rates for different portions or durations of the input video 204 based on camera statistics corresponding to the respective portions of the input video. For example, if a first portion of the input video 204 is associated with camera statistics 206 indicating that image frames from the first portion are out of focus or do not include an object or motion detected therein, the pre-processing system 108 may identify one or more video frames from the first portion at a low frame rate. In practice, because the content of blurry images and / or images including redundant content may provide less useful or redundant data for analysis using the image processing model 104, the pre-processing system 108 may provide fewer image frames (e.g., one frame for every five seconds of the input video 204) as input to the image processing model 104 to avoid wasting processing resources of a computing device implementing the image processing model 104.

[0041] As a further example, where a second portion of the input video 204 is associated with camera statistics 206 indicating that image frames from the second portion are in focus and / or include objects or motion detected therein, the pre-processing system 108 may identify image frames from the second portion at a higher frame rate than the first portion (wherein the images are out of focus and / or do not include detected objects). Because the content of the images that are in focus and / or include objects detected therein may provide more useful results and / or non-redundant data for analysis using the image processing model 104, the pre-processing system 108 may provide a higher number or rate of image frames (e.g., 2-10 frames per second of the input video 204) as input to the image processing model 104 because these frames are more likely to include useful data than image frames from other portions of the input video 204.

[0042] like Figure 2A As shown, the image processing model 104 can generate output 210 including various values ​​and / or images. For example, as described above, depending on the training or application of the image processing model 104, the output 210 can include a classification of an image, a value associated with an image, information about an image, a transformed image, or any other output associated with the input video 204 captured and generated by the video capture device 110. The output 210 can be provided to the computing device 106 for storage, display, or further processing.

[0043] Figure 2B Another example framework is shown that includes identifying and providing image frames and related statistics as input to the image processing model 104 according to one or more embodiments described herein. Specifically, similar to Figure 2A, the video capture device 110 can capture a video clip shot 212. The video capture device 110 can similarly generate an input video 214 (e.g., input video content) comprising a plurality of image frames and camera statistics 216 associated with the input video 214. The video capture device 110 can provide both the input video 214 and the associated camera statistics 216 to the pre-processing system 108, as described above in conjunction with Figure 2A discussed.

[0044] The pre-processing system 108 can utilize the camera statistics 216 to generate transformed video content 218. For example, the pre-processing system 108 can modify content from the input video 214 by further refining the image frames, modifying color or brightness, downsampling the resolution of the image frames, or otherwise modifying the input video 214. In one or more embodiments, the pre-processing system 108 transforms the input video 214 based on the camera statistics 216 received in conjunction with the input video 214. The pre-processing system 108 can also transform the input video 214 based on the application of the image processing model 104. For example, where the image processing model 104 implements a QR code reading algorithm, the pre-processing system 108 can transform the video content 204 by further enhancing, removing color, cropping extraneous content, or otherwise modifying the image frames that include the detected QR code image, particularly where the modification enables the image processing model 104 to more accurately or efficiently decode or decipher the QR code included within the image frame(s).

[0045] In addition to providing the transformed video content 218 to the image processing model 104, the pre-processing system 108 may also provide one or more identified camera statistics 220 as input to the image processing model 104. For example, as described above, the pre-processing system 108 may identify camera statistics 220 that include a subset of the camera statistics 216 provided by the video capture device 110 in conjunction with the input video 214. Specifically, the pre-processing system 108 may identify those camera statistics that are relevant to a selected or transformed image frame (e.g., an image frame of the transformed video content 218) and / or based on application of the image processing model 104 itself.

[0046] By providing the identified statistics 220 to the image processing model 104, in addition to the video content 218 (e.g., the transformed images or video content), the pre-processing system 108 can provide additional input information that enables the image processing model 104 to more efficiently or accurately generate the desired output 222. The image processing model 104 can select a particular algorithm that is best suited for processing the transformed video content 218 based on the statistics 220 provided to the image processing model 104. Additionally, the image processing model 104 can modify one or more algorithms applied to the transformed video content 218 to more efficiently or effectively analyze selected video content 218 based on the identified statistics 220. In this manner, even in situations where the transformed video content 218 includes repeated images, or the pre-processing system 108 has not yet incorporated the statistics 220 as described above, the pre-processing system 108 can provide additional input information that enables the image processing model 104 to more efficiently or accurately generate the desired output 222. Figure 2A With the image frames selectively identified as discussed, the image processing model 104 may still determine or identify the most relevant image frames for applying one or more algorithms of the image processing model 104 to the transformed video content 218 .

[0047] Although Figure 2A and Figure 2B Different inputs are shown that are selected and provided to the image processing model 104, but it should be understood that in combination Figure 2A The features and functionalities discussed can be combined with Figure 2B The features and functionalities discussed may be applied in combination (and vice versa). Figure 2A In addition to the selected subset of image frames 208, the pre-processing system 108 may also provide the image processing model 104 with identified camera statistics associated with the subset of image frames 208. Figure 2B As another example, the pre-processing system 108 can identify a subset of the transformed image frames 208 to provide to the image processing model 104 in addition to the identified camera statistics 220, thereby further enhancing the functionality of the image processing model 104 while conserving processing resources by having image frames at a lower frame rate than the input video 214 provided from the video capture device 110.

[0048] Figure 3 Another example implementation of a pre-processing system 108 for pre-processing video data and associated camera statistics from multiple video capture devices is shown. Specifically, Figure 3An example framework is shown in which multiple video capture devices 302a-302c use different hardware to capture video clip shots 304a-304c. In addition, the video capture devices 302a-302c can generate input videos 306a-306c and associated camera statistics 308a-308c for generating corresponding input videos 306a-306c for the respective video capture devices 302a-302c. Capturing the video clip shots 304a-304c and providing the input videos 306a-306c and the associated camera statistics 308a-308c can include combining the above with the following examples: Figure 2A The captured video footage 202 is shown as well as generating and providing similar features to those discussed above for the input video 204 and associated camera statistics 206 .

[0049] The pre-processing system 108 can similarly pre-process the input videos 306a-306c from the video capture devices 302a-302c based on the associated camera statistics 308a-308c to identify a subset of image frames 310 to be provided to the image processing model 104. For example, the pre-processing system 108 can selectively identify image frames from each of the input videos 306a-306c to enable the image processing model 104 to efficiently generate the output 312 for the multiple videos 306a-306c. As another example, where the input videos 306a-306c refer to input video streams provided simultaneously from the video capture devices 302a-306c, the pre-processing system 108 can selectively identify image frames from a single input video from the multiple input videos 306a-306c that are determined (e.g., based on the camera statistics for the single input video) to include content for generating the output 222 that is more useful than the other input videos.

[0050] In one or more embodiments, the pre-processing system 108 does not provide image frames of the input videos 306a-306c until the camera statistics 308a-308c indicate that the input videos 306a-306c may include interesting content that the image processing model 104 can use to generate the output 312. For example, where the camera statistics 308a-308c indicate that there is no motion or detected objects within the input videos 306a-306c, the pre-processing system 108 may determine to send zero image frames from any of the input videos 306a-306c until motion or other objects are detected within the input videos 306a-306c.

[0051] As an illustrative example, where the video capture devices 302a-302c refer to a network of secure video capture devices 302a-302c that simultaneously capture input video streams and provide them to the pre-processing system 108, the pre-processing system 108 may use the camera statistics 308a-308c provided by each of the video capture devices 302a-302c to identify which of the input videos 306a-306c include content of interest (e.g., an identified individual, animal, or other object) based on application of the image processing model 104. Where the pre-processing system 108 identifies that the first video 306a includes a detected individual or motion during a time period, the pre-processing system 108 may select a subset of image frames 310 from the first video 306a during the time period to provide to the image processing model 104, while discarding image frames from the second and third input videos 306b-306c for the same time period. The pre-processing system 108 may similarly switch between identifying subsets of image frames 310 from different videos 306a-306c based on camera statistics 308a-308c that change over time and as content of interest is detected at different time periods within the respective input videos 306a-306c.

[0052] As a further example, where the video capture devices 302a-302c selectively provide one input video at a time to the pre-processing system 108, the pre-processing system 108 can identify image frames in response to a detected scene change (e.g., a switch between input video streams) and provide them to the image processing model 104. For example, the pre-processing system 108 can detect a scene change based on changes in camera statistics for different input videos (e.g., changes in white balance, changes in focus). The pre-processing system 108 can respond to a detected scene change by quickly providing a few image frames to the image processing model 104 to classify the new scene, after which the pre-processing system 108 can wait until a new scene is detected before sending additional image frames to the image processing model 104.

[0053] like Figure 3 As shown, the image processing model 104 can generate output 312 based on the application of the image processing model 104. For example, in an example implementation including multiple security cameras, the image processing model 104 can include a face counter, a face identifier, or an application for identifying key images that include useful representations of individuals or other objects detected therein. In practice, similar to one or more implementations discussed herein, the image processing model 104 can generate various types of outputs depending on the training of the image processing model 104 to achieve a specific application.

[0054] Figure 4An example process is shown for identifying an image frame to provide as input to the image processing model 104, and generating an output based on application of the image processing model 104. For example, Figure 4 As shown, the pre-processing system 108 can perform action 402 of receiving video content (e.g., input video) and associated camera statistics from a video capture device (e.g., video capture device 110). The video content can include one or more digital video files, including the input video and associated camera statistics. In one or more embodiments, the video content includes one or more incoming streams of video content and associated camera statistics provided in real time from the video capture device(s).

[0055] In one or more embodiments, the pre-processing system 108 performs act 404 of identifying content of interest within the video content based on camera statistics. For example, the pre-processing system 108 can identify a set of image frames from the input video content that have been identified as being of high quality (e.g., focused, good lighting conditions) by the camera statistics. As another example, the pre-processing system 108 can identify image frames that include one or more detected objects depicted therein.

[0056] While identifying the content of interest may include selectively identifying image frames that include the content of interest, identifying the content of interest may include identifying portions of the image frames having the content of interest from the input video content. Figure 4 As shown, the pre-processing system 108 can identify regions within individual image frames that correspond to more or less related content (e.g., regions A and B). In one or more implementations, the pre-processing system 108 identifies content of interest based on the application of the image processing model 104. As an illustrative example, the pre-processing system 108 can identify regions of an image frame that include detected faces or other content of interest associated with the application or desired output of the image processing model 104. As another example, the pre-processing system 108 can identify foreground and background portions of an image frame and determine that the foreground portion corresponds to content of interest within the image.

[0057] like Figure 4 As further shown, the pre-processing system 108 can perform an act of transforming the video content 406. For example, as discussed in one or more embodiments above, transforming the video content can include enhancing pixels, removing colors, adjusting brightness between two input video streams, or otherwise modifying the image frame based on camera statistics and / or application of the image processing model 104. In one or more embodiments, the pre-processing system 108 transforms the video content by performing a cropping operation on the image frame by removing portions of the image frame that do not include the content of interest.

[0058] For example, Figure 4As shown, the pre-processing system 108 can transform the image frame by removing the area that does not include the content of interest to generate a cropped image that includes only the portion of the image frame that includes the content of interest. The pre-processing system 108 can similarly transform multiple image frames to generate a set of transformed image frames that include cropped portions associated with the area of ​​interest within the image frames from the input video content.

[0059] The pre-processing system 108 may additionally perform an act of identifying a subset of the video content 408. For example, the pre-processing system 108 may identify a subset of image frames from a plurality of image frames representing the input video content. Figure 4 As shown, the pre-processing system 108 can identify a subset of image frames that includes an identified portion of the image frames corresponding to the content of interest. For example, where multiple image frames have been transformed to include only cropped portions of the input image frames, the pre-processing system 108 can identify a subset of the cropped image frames that have been identified as including the content of interest.

[0060] Although Figure 4 An example is shown in which the pre-processing system 108 first identifies content of interest and transforms the image frames before selecting a subset of image frames to provide to the image processing model 104, but in one or more embodiments, the pre-processing system 108 first identifies a subset of image frames and subsequently transforms the subset of image frames based on camera statistics and application of the image processing model 104. For example, in one or more implementations, the pre-processing system 108 first identifies a subset of image frames, analyzes the content of the subset of image frames to identify regions of interest, and crops the subset of image frames (or otherwise modifies the image frames) based on the identified regions of interest.

[0061] like Figure 4 As further shown, the pre-processing system 108 can perform act 410 of providing the subset of video content as input to the image processing model 104. In one or more embodiments, the pre-processing system 108 provides the camera statistics in conjunction with the subset of video content to the image processing model 104. For example, the pre-processing system 108 can identify camera statistics relevant to the identified subset of image frames and application of the image processing model 104.

[0062] As further shown, the pre-processing system 108 can perform an action 412 of generating an output for the video content subset. The pre-processing system 108 can generate the output based on the application of the image processing model 104. As described above, the output can include various outputs (e.g., output images, image or video classifications, decoded values) based on various potential applications of the image processing model 104. Additionally, in one or more embodiments, the output can be based on a combination of the video content subset and selected camera statistics provided as input to the image processing model 104.

[0063] Now go to Figure 5 , additional details regarding the components and capabilities of an example architecture of the pre-processing system 108 will be provided. Figure 5 As shown and combined as above Figure 1 As described above, the pre-processing system 108 can be implemented by the computing device 106, which can refer to various devices, such as mobile devices (e.g., smartphones, laptops), non-mobile consumer electronic devices (e.g., desktop computers), edge computing devices, server devices, or other computing devices. According to one or more implementations described above, the pre-processing system 108 can selectively provide image frames to the image processing model 104 based on camera statistics received in conjunction with input video generated by a video capture device. In addition, in one or more embodiments, the pre-processing system 108 identifies camera statistics and provides them to the image processing model 104 for generating output based on the selected image frames.

[0064] like Figure 5 As shown, the pre-processing system 108 includes a camera statistics identifier 502, a video content analyzer 504, a content transformation manager 506, a frame selection manager 508, and a data storage device 510. The data storage device 510 may store camera data 512 and model data 514.

[0065] like Figure 5 As further shown in FIG, in addition to the pre-processing system 108, the image processing model 104 may also be included on the computing device 106. Specifically, as an alternative to the image processing model 104 being implemented on one or more server devices 102 (e.g., on a cloud computing system), the image processing model 104 is implemented on the computing device 106 to collaboratively generate output based on the image frames identified and provided by the pre-processing system 108.

[0066] Additionally, in one or more embodiments, computing device 106 includes a video capture device 110 implemented thereon. For example, where computing device 106 refers to a mobile device, video capture device 110 may refer to a front-facing camera, a rear-facing camera, or a combination of cameras implemented thereon. As another example, where computing device 106 includes a desktop computer, video capture device 110 may refer to an auxiliary device that is plugged into the desktop computer and operates in conjunction with a camera application operating on computing device 106.

[0067] like Figure 5As shown, the pre-processing system 108 includes a camera statistics identifier 502. The camera statistics identifier 502 can receive camera statistics received from the video capture device 110, the camera statistics provided in conjunction with the input video generated by the video capture device 110. In addition, the camera statistics identifier 502 can selectively identify some or all of the camera statistics based on the application of the image processing model 104. For example, the camera statistics identifier 502 can identify a set of relevant statistics for pre-processing the input video. In addition, the camera statistics identifier 502 can identify a set of relevant statistics to be provided to the image processing model 104 based on the corresponding application of the image processing model 104.

[0068] The pre-processing system 108 may also include a video content analyzer 504. For example, the video content analyzer 504 may analyze image frames of the input video to identify content of interest depicted within the image frames of the input video. The video content analyzer 504 may analyze the image frames of the input video based on camera statistics received in conjunction with the input video. For example, if the camera statistics indicate the selection of image frames in which one or more detected objects or motion are depicted, the video content analyzer 504 may analyze those image frames to identify regions or portions of the image frames having detected objects or motion.

[0069] The pre-processing system 108 may also include a content transformation manager 506. The content transformation manager 506 may perform a number of operations on the input video before providing any number of image frames to the image processing model 104. For example, the content transformation manager 506 may perform operations such as cropping image frames, smoothing image frames, combining multiple images (e.g., from subsequently captured image frames or received from different video capture devices), correcting brightness or focus issues of different image frames, enhancing different portions of image frames, or modifying the input video in a variety of other ways. The content transformation manager 506 may modify the image based on camera statistics and / or based on application of the image processing model 104.

[0070] As further shown, the pre-processing system 108 includes a frame selection manager 508. The frame selection manager 508 can identify selected portions of video content to provide to the image processing model 104. For example, the frame selection manager 508 can select a subset of frames from a plurality of frames identifying an input video received from a video capture device. As another example, the frame selection manager 508 can select portions of frames (e.g., cropped portions) that include identified content of interest to provide to the image processing model 104. Furthermore, the frame selection manager 508 can selectively identify transformed image frames to provide to the image processing model 104.

[0071] The pre-processing system 108 may also include a data storage device 510. The data storage device 510 may include camera data 512. The camera data 512 may include any information about one or more video capture devices in communication with the pre-processing system 108. For example, the camera data 512 may include device-related camera statistics, including information about camera specifications, image resolution, display ratio, brightness settings, frame rate, and other camera statistics about the one or more video capture devices stored on the data storage device 510. The camera data 512 may also include information about the orientation of one or more cameras to enable the pre-processing system 108 to merge images from multiple cameras or more accurately analyze video content captured by multiple video capture devices.

[0072] The data storage device 510 may also include model data 514. The model data 514 may include any information about one or more image processing models and / or applications executed by the image processing models. For example, the model data 514 may include an identification of one or more relevant statistics corresponding to a particular application or image processing model. The model data may also include content types (e.g., QR codes, faces) that the pre-processing system 108 may consider when determining which image frames to provide to the one or more image processing models.

[0073] Each of the components of computing device 106 may communicate with each other using any suitable communication technology. Figure 5 Although shown as being separated in the drawings, any of these components or subcomponents may be combined into fewer components, such as a single component, or divided into more components, depending on the particular implementation.

[0074] The components of computing device 106 may include software, hardware, or both. For example, Figure 5 The components of the computing device 106 shown may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the computing device. When executed by one or more processors, the computer-executable instructions of the computing device 106 may perform one or more methods described herein. Alternatively, the components of the pre-processing system 108 may include hardware, such as a dedicated processing device that performs a certain function or group of functions. Additionally or alternatively, the components of the computing device 106 may include a combination of computer-executable instructions and hardware.

[0075] Now go to Figure 6-Figure 7 , these figures illustrate example flow diagrams comprising a series of actions for selectively identifying image frames and / or camera statistics to provide to an image processing model. Figure 6-Figure 7Actions according to one or more embodiments are shown, but alternative embodiments may omit, add, reorder, and / or modify Figure 6-Figure 7 Any of the actions shown. Figure 6-Figure 7 Alternatively, the non-transitory computer readable medium may include instructions that, when executed by one or more processors, cause the computing device to perform Figure 6-Figure 7 In another embodiment, the system may perform Figure 6-Figure 7 action.

[0076] like Figure 6 As shown, a series of actions 600 may include an action 610 of receiving input video content from a video capture device. In one or more embodiments, action 610 includes receiving input video content comprising a plurality of image frames from one or more video capture devices. The input video content may include captured video footage that has been locally refined by the one or more video capture devices based on camera statistics.

[0077] like Figure 6 As further shown in FIG. 6 , the series of acts 600 may include an act 620 of identifying camera statistics for the input video content. In one or more embodiments, act 620 includes identifying camera statistics for the video content, where the camera statistics include data acquired by one or more video capture devices in conjunction with generating the video content. Identifying the camera statistics may include receiving a set of camera statistics from the one or more video capture devices and identifying one or more camera statistics based on application of an image processing model.

[0078] As further shown, the series of acts 600 may include an act 630 of determining a subset of image frames from the input video content. In one or more embodiments, act 630 includes determining the subset of image frames from the plurality of image frames based on camera statistics. In one or more embodiments, determining the subset of image frames includes selecting the image frames from the plurality of image frames based on a rate at which the image processing model is configured to generate output based on the input images.

[0079] In one or more embodiments, determining the subset of image frames includes identifying image frames that include content of interest based on camera statistics. Identifying the image frames may include identifying a first set of image frames to provide as input to an image processing model at a first frame rate, wherein the first set of image frames corresponds to a first duration of the video content that includes the content of interest. Additionally, identifying the image frames may include identifying a second set of image frames to provide as input to the image processing model at a second frame rate, wherein the second set of image frames corresponds to a second duration of the video content that does not include the content of interest. Based on the first set of image frames including the content of interest and the second set of image frames not including the content of interest, the second frame rate may be higher than the first frame rate.

[0080] In addition, the series of actions 600 may include an action 640 of providing the subset of image frames as input to an image processing model. In one or more embodiments, action 640 includes providing the subset of image frames as input to an image processing model that is trained to generate an output based on one or more input images. In one or more embodiments, the image processing model is a deep learning model. The deep learning model (or other type of image processing model) can be implemented on a cloud computing system. Additionally or alternatively, the deep learning model (or other type of image processing model) can be implemented on a computing device that receives input video content from a video capture device.

[0081] In one or more implementations, receiving input video content includes receiving a plurality of input video streams from a plurality of video capture devices, wherein the plurality of input video streams include image frames from the plurality of input video streams. Additionally, camera statistics may include data obtained by the plurality of video capture devices in combination to generate the plurality of input video streams. Additionally, determining the subset of image frames may include selectively identifying a subset of image frames from a first input video stream of the plurality of input video streams based on identified content of interest detected within video content from the first input video stream. Additionally, determining the subset of image frames may include selectively identifying image frames from the first input video stream based on camera statistics obtained by the video capture device that generated the first input video stream.

[0082] like Figure 7 As shown, another series of actions 700 may include an action 710 of receiving input video content and associated camera statistics from a video capture device. In one or more embodiments, action 710 includes receiving input video content and camera statistics from one or more video capture devices, wherein the camera statistics include data obtained by the one or more video capture devices in conjunction with generating the video content. In one or more embodiments, the input video content includes captured video footage that has been locally refined by the one or more video capture devices based on the camera statistics.

[0083] like Figure 7 As further shown in FIG, the series of acts 700 may include an act 720 of identifying camera statistics for the input video content based on the application of the image processing model. In one or more embodiments, act 720 includes identifying a subset of camera statistics from a set of camera statistics associated with a subset of the video content based on the application of the image processing model.

[0084] As further shown, the series of actions 700 may include providing the identified camera statistics and associated video content as input to an image processing model. In one or more embodiments, action 730 includes providing the identified subset of camera statistics and the associated subset of video content as input to a deep learning model (or other image processing model) that is trained to generate output based on the video content and the camera statistics. Providing the identified subset of camera statistics and the associated subset of video content as input to the deep learning model may include providing a cropped portion of the video content to the deep learning model.

[0085] In one or more embodiments, a series of actions 700 includes transforming the input video content based on the identified camera statistics subset. Additionally, providing the identified camera statistics subset and the associated video content subset to the deep learning model may include providing the transformed video content to the deep learning model.

[0086] The deep learning model can be trained based on both training data including video content and related camera statistics. Additionally, the deep learning model can be implemented on one or more of a cloud computing system or computing devices that receives input video content from one or more video capture devices. Additionally, in one or more embodiments, the one or more video capture devices and the deep learning model are implemented on a computing device and coupled to one or more processors of the system.

[0087] Figure 8 Illustrated are certain components that may be included within computer system 800. One or more computer systems 800 may be used to implement the various devices, components, and systems described herein.

[0088] The computer system 800 includes a processor 801. The processor 801 may be a general-purpose single-chip or multi-chip microprocessor (e.g., Advanced RISC (Reduced Instruction Set Computer) Machine (ARM)), a dedicated microprocessor (e.g., a digital signal processor (DSP)), a microcontroller, a programmable gate array, etc. The processor 801 may be referred to as a central processing unit (CPU). Although Figure 8 Just a single processor 801 is shown in the computer system 800 , but in an alternative configuration, a combination of processors (eg, an ARM and DSP) could be used.

[0089] Computer system 800 also includes memory 803 in electronic communication with processor 801. Memory 803 can be any electronic component capable of storing electronic information. For example, memory 803 can be embodied as random access memory (RAM), read-only memory (ROM), magnetic disk storage media, optical storage media, flash memory devices in RAM, onboard memory included with a processor, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, etc., including combinations thereof.

[0090] Instructions 805 and data 807 may be stored in memory 803. Instructions 805 are executable by processor 801 to implement some or all of the functionality disclosed herein. Executing instructions 805 may involve using data 807 stored in memory 803. Any of the various examples of modules and components described herein may be implemented in part or in whole as instructions 805 stored in memory 803 and executed by processor 801. Any of the various examples of data described herein may be among the data 807 stored in memory 803 and used during execution of instructions 805 by processor 801.

[0091] The computer system 800 may also include one or more communication interfaces 809 for communicating with other electronic devices. The communication interface(s) 809 may be based on wired communication technology, wireless communication technology, or both. Some examples of the communication interface 809 include a Universal Serial Bus (USB), an Ethernet adapter, a wireless adapter operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 wireless communication protocol, Wireless communication adapter and infrared (IR) communication port.

[0092] The computer system 800 may also include one or more input devices 811 and one or more output devices 813. Some examples of input devices 811 include a keyboard, a mouse, a microphone, a remote control device, buttons, a joystick, a trackball, a touchpad, and a light pen. Some examples of output devices 813 include speakers and a printer. One specific type of output device typically included in the computer system 800 is a display device 815. The display device 815 used with the embodiments disclosed herein can utilize any suitable image projection technology, such as a liquid crystal display (LCD), a light emitting diode (LED), gas plasma, electroluminescence, etc. A display controller 817 may also be provided for converting data 807 stored in the memory 803 into text, graphics, and / or moving images (as appropriate) displayed on the display device 815.

[0093] The various components of the computer system 800 may be coupled together via one or more buses, which may include a power bus, a control signal bus, a status signal bus, a data bus, etc. For clarity, the various buses are shown in FIG. Figure 8 Shown in the figure is a bus system 819.

[0094] Unless specifically described as being implemented in a particular manner, the techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. Any features described as modules, components, etc. may also be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a non-transitory processor-readable storage medium comprising instructions that, when executed by at least one processor, perform one or more of the methods described herein. Instructions may be organized into routines, programs, objects, components, data structures, etc., which may perform specific tasks and / or implement specific data types, and may be combined or distributed as needed in various embodiments.

[0095] The steps and / or actions of the methods described herein may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is required for proper operation of the described method, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.

[0096] The term "determining" includes various actions, and thus, "determining" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database, or another data structure), ascertaining, etc. Furthermore, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determining" may include resolving, selecting, choosing, establishing, etc.

[0097] The terms "comprising," "including," and "having" are intended to be inclusive and mean that additional elements may be present in addition to the listed elements. Furthermore, it should be understood that reference to "one embodiment" or "an embodiment" in this disclosure is not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. For example, any element or feature described with respect to an embodiment herein can be combined with any element or feature of any other embodiment described herein, where compatible.

[0098] The present disclosure may be implemented in other specific forms without departing from its spirit or characteristics. The described embodiments should be considered as illustrative rather than restrictive. Therefore, the scope of the present disclosure is indicated by the appended claims rather than by the foregoing description. Changes within the equivalent meaning and range of the claims are intended to be included within their scope.

Claims

1. A method for preprocessing image frames based on camera statistics, comprising: receiving input video content comprising a plurality of image frames from one or more video capture devices; identifying camera statistics for the video content, the camera statistics comprising data acquired by the one or more video capture devices in conjunction with generating the video content; providing a first subset of image frames from the plurality of image frames as input to an image processing model at a first frame rate, the image processing model being trained to generate an output based on one or more input images; detecting content of interest in a second subset of image frames of the plurality of image frames based on the camera statistics; as well as Based on detecting the content of interest in the second subset of image frames of the plurality of image frames, the second subset of image frames is provided as input to the image processing model at a second frame rate higher than the first frame rate. 2 . The method of claim 1 , wherein the input video content comprises captured video footage that has been locally refined by the one or more video capture devices based on the camera statistics.

3. The method of claim 1 , wherein identifying the camera statistics comprises: receiving a set of camera statistics from the one or more video capture devices; as well as One or more camera statistics are identified based on application of the image processing model.

4. The method of claim 1 , further comprising selecting an image frame from the plurality of image frames at a higher frame rate based on a rate at which the image processing model is configured to generate an output based on an input image. 5 . The method of claim 1 , wherein detecting content of interest in the second subset of image frames comprises identifying the second subset of image frames including content of interest based on the camera statistics.

6. The method according to claim 5, wherein the first subset of image frames corresponds to a first duration of the video content that does not include content of interest; and The second subset of image frames corresponds to a second duration of the video content including content of interest.

7. The method according to claim 1, wherein receiving the input video content comprises receiving a plurality of input video streams from a plurality of video capture devices, the plurality of input video streams comprising image frames from the plurality of input video streams; and The camera statistics include data obtained by combining the plurality of video capture devices to generate the plurality of input video streams.

8. The method according to claim 7, further comprising: identifying the first subset of image frames from a first input video stream of the plurality of input video streams, and determining the first frame rate based on determining that no content of interest is present in video content from the first input video stream; as well as The second subset of image frames is identified from a second input video stream of the plurality of input video streams, and the first frame rate is determined based on determining the presence of interesting content in video content from the second input video stream.

9. The method of claim 8, wherein identifying the second subset of image frames further comprises: Image frames are selectively identified from the second input video stream based on camera statistics acquired by a video capture device generating the second input video stream.

10. The method of claim 1 , wherein the image processing model comprises a deep learning model, and wherein providing the first subset of image frames and providing the second subset of image frames comprises providing the first subset of image frames and the second subset of image frames to the deep learning model, the deep learning model being implemented on a cloud computing system.

11. The method of claim 1 , wherein the image processing model comprises a deep learning model, and wherein providing the first subset of image frames and providing the second subset of image frames comprises providing the first subset of image frames and the second subset of image frames as input to the deep learning model, the deep learning model being implemented on a computing device that receives the input video content from the video capture device.

12. The method of claim 1 , wherein providing the second subset of image frames as input to the image processing model at a second frame rate higher than the first frame rate is based on an ability of a computing device to apply the image processing model at a particular frame rate.

13. The method of claim 1, further comprising modifying the second subset of image frames based on the content of interest before providing the second subset of image frames as input to the image processing model.

14. The method of claim 13, wherein modifying the second subset of image frames comprises: Content is removed from the second subset of image frames based on the content of interest.

15. The method of claim 13, wherein modifying the second subset of image frames comprises one or more of: refining, modifying color, modifying brightness, downsampling resolution, enhancing, removing color, enhancing pixels, or cropping extraneous content of one or more image frames in the second subset of image frames.

16. The method of claim 1, wherein the content of interest is at least one of: an identified individual, animal, face, object, or motion data detected in the image frames from the second subset of image frames.

17. A system for preprocessing image frames based on camera statistics, comprising: one or more processors; a memory in electronic communication with the one or more processors; as well as Instructions stored in the memory, the instructions being executable by the one or more processors to cause the computing device to: receiving input video content comprising a plurality of image frames from a video capture device; identifying camera statistics for the video content, the camera statistics comprising data acquired by the video capture device in conjunction with capturing the video content; providing a first subset of image frames from the plurality of image frames as input to a deep learning model at a first frame rate, wherein the deep learning model is trained to generate an output based on one or more input images; detecting content of interest in a second subset of image frames from the plurality of image frames based on the camera statistics; as well as Based on the content of interest detected from the second subset of the plurality of image frames, the second subset of image frames is provided as input to the deep learning model at a second frame rate higher than the first frame rate.

18. The system of claim 17, wherein the video capture device and the deep learning model trained to generate the output based on the one or more input images are both implemented on the computing device and coupled to one or more processors of the system.

19. The system of claim 17, wherein providing the second subset of image frames as input to the deep learning model at a second frame rate higher than the first frame rate comprises: The second subset of image frames is provided at a frame rate higher than the first frame rate based on the ability of the computing device to apply the deep learning model at a particular frame rate.

20. The system of claim 17, further comprising modifying the second subset of image frames based on the content of interest before providing the second subset of image frames as input to the deep learning model.

Citation Information

Patent Citations

  • Processing method and device with video temporal up-conversion

    CN101223786A

  • Method and system for indexing and searching objects of interest across a plurality of video streams

    CN101663676A

  • Selecting key frames from video frames

    US20070147504A1

  • Compact video representation for video event retrieval and recognition

    US20180041765A1