Smart cropping of images
Through intelligent cropping technology, the image cropping area is automatically determined, which solves the problem of poor display of images on multi-functional mobile devices on different devices and orientations, and achieves high-quality image cropping and cropping quality evaluation.
Patent Information
- Application Number
- CN202110568735.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-19
- Filing Date
- 2021-05-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-05-25
AI Technical Summary
On multi-functional mobile devices, users want to automatically crop images into screen savers or background images suitable for different device displays and applications, however, single cropping is difficult to provide visually pleasing images on individual devices and orientations due to differences in device display properties and application content area constraints.
Through intelligent cropping technology based on image content, the cropping area is automatically determined, taking into account the significance, aspect ratio, resolution and orientation of the image, the cropping area is optimized to maximize overlap with the region of interest, and the cropping score is calculated to evaluate the cropping quality.
Automatically determine visually pleasing image cropping areas on different devices and orientations, ensuring that important parts of the image are displayed on each device and of high quality, providing a method of quantifying cropping quality for users to choose the best cropping.
Smart Images

Figure CN113822898B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to the field of digital image processing. More specifically, but not by way of limitation, the present disclosure relates to techniques for automatically cropping images in an intelligent manner, e.g., based on image content, and the aspect ratio, resolution, orientation, etc. of various display screens and / or display areas on which such images may be displayed. Background Art
[0002] The advent of mobile, multi-function devices, such as smartphones and tablet devices, has created a desire for high-quality display screens and small form factor cameras capable of generating high levels of image quality in near real time to be integrated into such mobile, multi-function devices. As users increasingly rely on these multi-function devices as their primary displays and cameras for daily use, users are able to capture and view images with image quality levels approaching (or exceeding) those to which they have become accustomed by using dedicated display monitors and camera devices.
[0003] Thus, users may typically want to use such captured images (or images obtained from other sources), for example, as part of a screen saver and / or as a "wallpaper" or "background image" on any of their devices with a display. However, many users have a variety of devices with different display screen sizes, orientations, aspect ratios, resolutions, etc., and may want to use one or more of their images as background images on any of their devices. Additionally, in some cases, one or more applications installed on a user's device may also wish to display such images within a designated content area (e.g., within a predetermined area on a display) as part of a user interface (UI) or other multimedia presentation application. In such cases, each designated content area of each application may also have its own constraints on the size, orientation, aspect ratio, resolution, etc. of the image content that may be used within the designated content area of the application, i.e., independent of the screen size, orientation, aspect ratio, resolution, etc. of the overall device display.
[0004] Due to the differences in the aforementioned device display properties and application-specific content area constraints (such as display screen size, orientation, aspect ratio, specified content area size, and resolution), a single crop obtained from one of a user's images is unlikely to provide a visually pleasing image on each of the user's devices and applications in every possible orientation of such devices. For example, it may be advantageous and visually pleasing for a crop of an image to be used on a user's device to include as many portions of the image as possible that are considered important, significant, and / or otherwise relevant (such portions of the image are also collectively referred to herein as "important"). It may also be advantageous and visually pleasing for a crop of an image to be used on a user's device to be able to take into account areas on the device display such that it would be preferred that important portions of the image do not overlap therewith (e.g., it may not be visually pleasing if a determined crop of a background image to be used on a device display results in important portions of the cropped image being covered by text, a title, a clock, a battery indicator, or other display elements present on the device's display screen during normal operation of the device's operating system).
[0005] Therefore, it would be advantageous to have methods, computer executable instructions, and systems that provide automatic and intelligent cropping of images, for example based on image content, as well as aspect ratios, resolutions, orientations, etc. of various display screens and designated content areas on which such images may be displayed. It would also be desirable to be able to automatically calculate a score for such an intelligent crop so that an entity (e.g., an end user or application) requesting the crop can quantify the likely quality of the crop for use on a particular device display screen at a particular orientation or within a particular designated content area. Summary of the invention
[0006] The present invention discloses an apparatus, method and non-transitory program storage device for providing automatic and intelligent cropping of an image given a requested target size of a crop region, from which an aspect ratio and / or orientation can be determined. In some embodiments, the location of the requested crop region within the image can be determined, for example, by using a saliency map or other object detection and / or classifier system to identify the portion of the image containing the most important or relevant content - and, if possible, ensure that such content is included in the determined crop region from the image (such determined crop region may also be referred to herein as a "crop box" or simply "crop").
[0007] Specifically, various devices, methods, and non-transitory program storage devices disclosed herein may be capable of: defining a first region of interest (ROI) in a given image that is most essential to include in an automatically determined crop region; defining a second (e.g., larger) ROI in the given image that is preferred to be included in the automatically determined crop region; and then determining a crop region from the given image based on a requested aspect ratio in an attempt to maximize the amount of overlap between the determined crop region and the first ROI and / or second ROI.
[0008] In a preferred embodiment, a crop score for the determined crop is determined based at least in part on how much of the first ROI and the second ROI are surrounded by the determined crop. In some cases, an interpolation operation such as linear interpolation can be used to determine the crop score for a given crop, for example, an interpolation between two predetermined crop scores assigned to crops that surround certain defined areas of the image (e.g., a defined area such as the first ROI, the second ROI, or the entire image extent). The crop score can be used to help an end user or application evaluate whether the determined crop is actually a good candidate to be used, for example, as part of a screen saver, as a wallpaper or background image, or for display in a designated content area on a display of a particular device.
[0009] According to other embodiments, additional crops of a given image may be determined using the techniques disclosed herein, for example, multiple crops of a given image having different target sizes, aspect ratios, different orientations, different resolution requirements, etc. may each be returned (along with corresponding crop scores) to a requesting end user or application.
[0010] According to other embodiments, the first ROI may be determined to surround all parts of the image having a significance score greater than a first threshold, and the second ROI may be determined to cover all parts of the image having a significance score greater than a second threshold, where, for example, the second threshold significance score is lower than the first threshold significance score. Due to having a lower threshold significance score, the second ROI will necessarily be larger than (and possibly cover) the first ROI. Each ROI may be continuous or discontinuous within the image. As described above, the first ROI may represent content that is considered "essential" to be included in the determined cropping, while the second ROI may represent content that is considered "preferred" to be included in the determined cropping. According to some embodiments, the more original images are included in the determined cropping, the higher the cropping score of the determined cropping will be, and the cropping score will reach a maximum value if the entire original image (or at least the entire horizontal range or the entire vertical range of the image) can be included in the determined cropping.
[0011] According to some crop scoring schemes, if the first ROI is completely enclosed in the determined crop region, the crop score for a given determined crop region is set to at least a first minimum score, and if the second ROI is completely enclosed in the determined crop region, the crop score is set to at least a second minimum score, wherein the second minimum score is greater than the first minimum score. In other words, if the determined crop region includes an "essential" portion of the image (i.e., the first ROI), it will be assigned a score of at least X, and if the determined crop region includes both an "essential" portion and a "preferred" portion of the image (i.e., the second ROI), it will be assigned a score of at least Y, wherein Y is greater than X. In other crop scoring schemes, the image may be divided into a plurality of ranked regions, wherein each ranked region is assigned a particular weighted score, and wherein the assigned crop score may include a weighted sum of the portions of each ranked region encompassed by the determined crop region. In some crop scoring schemes, a determined crop region may be assigned a maximum crop score, such as a score of 100%, if the determined crop region is coextensive with the original image (i.e., includes all content from the original image) in at least one dimension, or if the determined crop region encompasses all identified ROIs. In some cases, cropping may not be used (or recommended for an end user or requesting application) unless the crop score is greater than a minimum score threshold, such as a score of 50%.
[0012] In other embodiments, the requested crop may also include a specification of a "focus area", e.g., in addition to the requested aspect ratio, the requested crop may also specify a portion of the determined crop area (e.g., the bottom 75% of the crop area, the bottom 50% of the crop area, etc.), referred to herein as the focus area, wherein the crop score for the determined area is further determined based at least in part on the amount of the first ROI and / or the second ROI enclosed by the focus area. In other words, if the portion of the first ROI and / or the second ROI included in the determined crop area extends beyond the specified focus area, it may adversely affect the crop score of the determined crop area. For example, in some crop scoring schemes, if any portion of the first ROI (or some other ROI) in the determined crop area extends beyond the specified boundaries of the focus area, the determined crop area may be given a crop score below a minimum threshold score (and therefore may not be recommended for use by an end user or application).
[0013] In some embodiments, in addition to (or instead of) a saliency map, one or more of an object detection box, a face detection box, or a face recognition box generated based on the image may be used to determine the first ROI or the second ROI.
[0014] In other embodiments, when determining the size of the determined crop region, at least one of the width or height of the crop region may be selected to match a corresponding size of the image.
[0015] Various embodiments of non-transitory program storage devices are disclosed herein. Such program storage devices can be read by one or more processors. Instructions can be stored on the program storage device to cause one or more processors to perform any of the techniques disclosed herein.
[0016] According to the embodiments of the program storage device listed above, various programmable electronic devices are also disclosed herein. Such electronic devices may include: one or more image capture devices, such as optical image sensors / camera units; displays; user interfaces; one or more processors; and memory coupled to the one or more processors. Instructions may be stored in the memory, which instructions cause the one or more processors to execute instructions according to the various techniques disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A and 1B Exemplary images, saliency maps, and regions of interest (ROIs) are shown according to one or more embodiments.
[0018] Figure 2A and 2B An exemplary determined cropping region according to one or more embodiments is shown.
[0019] Figure 3A A graph illustrating exemplary cropping scores according to one or more embodiments is shown.
[0020] Figure 3B Exemplary interpolation techniques for determining clipping fractions are shown in accordance with one or more embodiments.
[0021] Figure 4 is a flow chart illustrating a method of performing an automatic image cropping technique according to one or more embodiments.
[0022] Figure 5 is a flow chart illustrating a method of performing an automatic image cropping technique according to one or more embodiments.
[0023] Figure 6 is a block diagram illustrating a programmable electronic computing device in which one or more of the techniques disclosed herein may be implemented. DETAILED DESCRIPTION
[0024] In the following description, for the purpose of explanation, a lot of specific details are set forth in order to provide a thorough understanding of the invention disclosed herein. However, it is obvious to those skilled in the art that the present invention can be practiced in the absence of these specific details. In other cases, structures and devices are shown in the form of block diagrams to avoid blurring the present invention. References to numbers without subscripts or suffixes should be understood as references to all subscripts and suffixes corresponding to the figure marks. In addition, the language used in this disclosure has been selected primarily for readability and instructive purposes, and may not be selected to delimit or limit the subject matter of the present invention, and therefore may need to resort to the claims to determine such subject matter of the invention. References to "one embodiment" or "embodiment" (or similar expressions) in the specification refer to specific features, structures or characteristics described in conjunction with the embodiment included in at least one embodiment of one of the invention, and multiple references to "one embodiment" or "embodiment" should not be understood as all necessarily referring to the same embodiment.
[0025] Introduction and Problem Background
[0026] Now go to Figure 1A , shows an exemplary image, saliency map, and region of interest (ROI) according to one or more embodiments. The first image 100 will be used as a sample image to discuss the various techniques presented herein. As can be seen, the image 100 is a rectangular, landscape-oriented image that includes various human subjects 102 / 104 / 106 positioned from left to right within the entire image. The image 100 also reflects an outdoor scene, in which the background of the human subjects includes various objects, such as walls, trees, the moon, etc.
[0027] Assuming that the user wants to use the first image 100 as a background image on the display of one of their electronic devices, and therefore provides a target size for a crop region that they wish to determine based on the first image 100, a first determination may be made as to whether the aspect ratio of the target size of the first image 100 matches the aspect ratio of the display of the target electronic device on which the user is interested in using the image 100 as a background image. If the aspect ratio of the target size of the image 100 and the aspect ratio of the display of the target device match, then (assuming the image has sufficient resolution) the image 100 may simply be used as a background image on the display of the target device without further modification.
[0028] However, more often, there will be a mismatch between the aspect ratio and / or target size requested for cropping of a given image and the aspect ratio and / or target size of the target display (or area of the display) on which the user wishes to use the image. In addition, many electronic device displays can be used in multiple orientations (e.g., portrait and landscape), which means that even for a single image intended for a single display device, there may be multiple different cropping regions that need to be determined. For example, using an unaltered landscape image 100 as a background image on a device (e.g., a smart phone) operating in a portrait orientation will not be visually pleasing because, for example, the sky will appear on the right side of the device display, and three human subjects appear to appear from the left side of the device display and are stacked vertically on each other. Instead, it is desirable to automatically determine a visually pleasing vertical cropping region that will fit the device display when in a portrait orientation, while still displaying an important portion of the image (and in the correct orientation).
[0029] As another example, a user may want to use image 100 as a background image on two (or more) different devices with different display properties, such as a smartphone with a 16:9 screen aspect ratio in a portrait orientation, a desktop monitor with a 16:9 screen aspect ratio in a landscape orientation, and a tablet device with two possible orientations, portrait and landscape (each with a 4:3 screen aspect ratio). Thus, in summary, a user may desire to automatically determine four different smart crop regions for image 100, such that each determined crop region has the correct target size and aspect ratio and includes important content when used as a background image on its corresponding device (and at its corresponding orientation). (It should be understood that all references to the desired use of image 100 as a background image on a display device also apply to the desired use of image 100 within a specified content area of a given aspect ratio and / or size within an application UI.)
[0030] As described above, one aspect of automatically determining a smart crop region for a given image is the ability to understand which portions of the image contain content that may be important, relevant, or otherwise significant to the user. Once such a determination is made, it may be desirable to include as much of such important content as possible in the determined crop region (while also optionally further aiming to keep as much of the important content as possible within the focus region within the determined crop region, as will be discussed below with respect to Figure 2B described in more detail).
[0031] In some embodiments, a saliency heatmap (e.g., the exemplary saliency heatmap 110 in FIG. 1 ) may be utilized to generate bounding boxes around salient objects (i.e., saliency-O) and / or salient regions in an image to which a user's attention may be directed when viewing the image (i.e., saliency-A). For purposes of this specification, a salient object or salient region refers to a portion of an image that is potentially of interest, and a saliency value refers to the likelihood that a particular pixel belongs to a salient object or region within the image.
[0032] A saliency heatmap may provide a binary determination for each pixel in the image (e.g., a value of "0" for non-salient pixels, and a value of "1" for salient pixels). In other cases, as shown in the exemplary saliency heatmap 110 of FIG. 1 , there may be a continuous saliency score assigned to each pixel, covering a range of possible scores, such as scores from 0% up to 100%. For example, the smallest dark square centered over the faces of the human subjects in image 110 may represent an area of pixels with a saliency score of 60% or greater. The next larger square with a slightly lighter shading on each human subject's face may represent an area of pixels with a saliency score of 50% or greater. Finally, the outermost largest square with the lightest shading on each human subject's face may represent an area of pixels with a saliency score of 15% or greater. Areas of image 110 not covered by the box in this heatmap example may simply represent areas of pixels with saliency scores below 15%, i.e., areas of the image that are unlikely to have interesting or important content that a user would find essential or important to include in a determined cropped area to be used for a background image or in a designated content area on one of their devices. It will be appreciated that, if desired for a given implementation, a saliency heatmap may alternatively be generated on the downsampled image, such that each portion of pixels is assigned an estimated saliency value in the heatmap.
[0033] According to some embodiments, the saliency model used to generate the saliency heat map 110 may include a trained saliency network, through which the saliency of an object can be predicted for an image. In one or more embodiments, the saliency model can be trained with still image data or video data, and can be trained to predict the saliency of various objects in an image. The saliency model can be trained in a category-independent manner. That is, the type of object may be irrelevant in the saliency network, which may only involve whether a particular object is salient. In addition, the saliency network can be trained based on RGB image data and / or RGB+depth image data. According to one or more embodiments, by incorporating depth into the training data, a more accurate saliency heat map may be generated. For example, depth can be used to identify object boundaries, scene layouts, and the like.
[0034] In one or more embodiments, a trained saliency network may take as input an image, such as image 100, and output a saliency heatmap, such as saliency heatmap 110, that indicates the likelihood that a particular portion of the image is associated with a salient object or region. In addition, in one or more embodiments, the trained saliency network may also output one or more bounding boxes indicating regions of interest within the saliency heatmap. In one or more embodiments, such as those described in commonly assigned, co-pending U.S. patent application Ser. No. 16 / 848,315 (hereinafter, the "'315 patent application," the entirety of which is hereby incorporated by reference), a saliency model may be combined with or fed into a bounding box neural network that may be used to predict an optimal size and / or location for a bounding box.
[0035] In other embodiments, such as those to be shown herein, a simple threshold operation may be used to determine the bounding box. For example, as shown in image 120, a first ROI 122 (which may also be referred to herein as an "inner region," "inner crop," or "tight crop") may be determined as the smallest rectangle that can encompass all portions of the image having a significance score greater than a first threshold (e.g., a 60% score associated with the darkest square region in the significance heatmap, as described above). Similarly, as shown in image 130, a second ROI 132 (which may also be referred to herein as an "outer region," "outer crop," or "loose crop") may be determined as the smallest rectangle that can encompass all portions of the image having a significance score greater than a second threshold (e.g., a 15% score associated with the lightest square region in the significance heatmap, as described above), wherein the second threshold significance score is lower than the first threshold significance score. As described above, if possible, the first ROI may be used as a proxy for portions of the image that are considered "essential" in the cropped image, and the second ROI may be used as a proxy for portions of the image that are considered "preferred" in the cropped image. In some cases, the determined ROI itself may simply be used as the determined cropping region for a given image, e.g., assuming it has a target size that meets the requirements of the end user or application. It should be understood that in a given implementation, a different threshold significance score may be used for each ROI, and any desired number of ROIs may be identified in a given smart cropping scheme, which may be continuous or non-contiguous within the image, and may be non-overlapping or at least partially overlapping.
[0036] Now go to Figure 1B, which shows an exemplary ROI and an expanded ROI according to one or more embodiments. As shown in image 140, one or more object detection classifiers or algorithms may also be run on image 100 to identify various objects, such as tree 142 and / or moon 144. In some cases, for example, depending on the type of object identified (or in the case of a facial recognition algorithm, the identity of the person identified), the boundaries of the ROI (e.g., determined by a saliency heat map) may be expanded (or otherwise modified) to incorporate (or exclude) one or more of the identified objects. For example, as shown in image 150, if it is determined that tree 142 and moon 144 are types of objects that a user would typically consider salient (and, therefore, would like to include in any determined cropped area to be used as a background image or within a designated content area), then the above reference may be used to define the cropped area. Figure 1A The original rectangular region defining the second ROI 132 discussed above may be expanded to include the tree 142 and the moon 144, as shown in the expanded ROI 152. It will now be appreciated that the exact specifications of how the boundaries of the ROIs are defined, how large they are made, and / or what objects are considered to be included (or excluded) within the ROIs (as well as how many different levels or "tiers" of ROIs are used on a given image) may all be customized based on the needs of a given implementation.
[0037] Example cropping area
[0038] Now go to Figure 2A , which shows an exemplary determined cropping region 202 / 212 according to one or more embodiments. Turning first to image 200 (which includes the same content as image 100, and which shows the same overlapping first ROI 122 and second ROI 132, as described above with respect to Figure 1A In some embodiments, it may be preferred to attempt to match at least one dimension (e.g., width or height) of the determined crop region to a corresponding dimension of the first image. In the illustrated example, the method is able to match the width of crop region 202 to the width of image 200. (It should be understood that in some cases, it may not be possible to match one of the dimensions of the determined crop region to a corresponding dimension of the first image for various reasons, including the aspect ratio and / or resolution of the first image.)
[0039] Once the width of the cropping region 202 has been determined, the height of the cropping region 202 may be determined based on the specific aspect ratio of the target size requested by the end user or application for the potential background image or the specified content region cropping. After determining the size of the cropping region 202, the method may then attempt to determine where the cropping region should be located within the original image 200 so as to produce the most visually pleasing background image or specified content region cropping from the first image. In some embodiments, this may include setting at least one of the first width, the first height, and the first position of the determined cropping region based at least in part on an effort to maximize the amount of overlap between the first cropping region and the first ROI. In other embodiments, the effort to determine the size and position of the cropping region may be configured to give priority to covering the entire first ROI, and then, assuming that the first ROI is completely covered, it is further configured to try to overlap as much of the second ROI as possible given the image constraints and the target size requested for the cropping region. As shown in image 200, given the target size requested for cropping, the location of a crop region 202 can be determined that encompasses the entire first ROI 122 and the second ROI 132. Therefore, based on the manner in which the first and second ROIs are specified using the saliency heatmap, the determined crop region 202 will most likely encompass all essential and preferred subject matter of the original first image.
[0040] Further consideration may be given to the position of the determined crop region 202 vertically placed within the scope of the image 200. For example, the determined crop region 202 may be placed vertically at various positions within the scope of the image 200 and still encompass all of both the first ROI 122 and the second ROI 132. Thus, the exact position of the crop region must still be determined. According to some embodiments, it may be preferred to center the crop region 202 relative to one or more of the ROIs because there may be an implicit assumption that the importance of a given ROI derives from the center of the ROI. For example, as shown in the image 200, the position of the determined crop region 202 has been centered so that the top of the crop region 202 is located midway between the top of the second ROI 132 and the top boundary of the image 200, while the bottom of the crop region 202 is simultaneously located midway between the bottom of the second ROI 132 and the bottom boundary of the image 200. It should be understood that different criteria may be used when determining the placement of the crop region (e.g., where a user has defined a "focus area" within the crop region, as will be described below with reference to Figure 2B ), and centering the cropped region within the image relative to the largest ROI is only one exemplary approach that may be followed.
[0041] As will be referred to below Figure 3AExplaining in more detail, according to some embodiments, a crop score may be determined for each determined crop region. The crop score may include a score designed to quantify the likely quality of a crop region for use on a display screen of a particular device. In some cases, there may be a minimum score threshold defined as the minimum score that a determined crop region must achieve before the crop region is recommended to an end user or application for use as a background image or for use within a specified content area. For example, in one embodiment, a simple minimum score threshold may be set based on the score achieved when the crop region encompasses the entire first ROI (i.e., the portion of the image deemed most essential by the saliency network). In other words, in Figure 2A In the example of , a given determined crop may be rejected if it does not encompass at least the entire first ROI 122. Since the determined crop region 202 does encompass the entire first ROI 122, it is shown with a check mark below the image 200, indicating that a successful lateral crop region 202 has been automatically and intelligently determined by the method. It should be understood that other minimum score thresholds may also be employed, such as threshold scores based on the crop region having to encompass all identified ROIs, having to encompass a certain percentage of the total pixels in the image, having to have a certain minimum size, etc.
[0042] Turning now to image 210, in contrast, the end user (or application) has requested a crop region 212 having a similar aspect ratio as crop region 202, but having a different orientation (i.e., a portrait orientation rather than a landscape orientation). Following the same process outlined above for image 200, the method may attempt to match the height dimension of crop region 212 to the height dimension of image 210, and then seek a location within the extent of image 200 where crop region 212 may overlap the maximum amount with the first ROI and / or the second ROI. As shown, no matter where crop region 212 is located on the horizontal extent of image 210, it will not encompass the entire first ROI 122 (let alone the entire larger second ROI 132). Therefore, assuming a similar minimum score threshold as described above with respect to image 200 is applied, the determined crop region 212 will be rejected (indicated by the "X" mark below image 210) because there is no location within the extent of image 200 where it can be placed that will encompass the entire first ROI 122. It appears that the optimal placement of the determined crop region 212 may be as shown in image 210, i.e., encompassing the faces of the two leftmost human subjects in images 104 and 106, but not encompassing the face of the human subject located on the right in image 102. As described above, if the minimum score threshold is relaxed in a given implementation (e.g., requiring that only 50% of the first ROI 122 need be encompassed in the determined crop region), then the determined crop 212 may be considered successful or acceptable.
[0043] According to other embodiments, for example, as described above with reference to Figure 1B As described, the size and / or range of the determined ROI may be modified (e.g., expanded or reduced) based on one or more classifiers or object detection systems. For example, if a facial recognition system is used in conjunction with a saliency network, any unrecognized faces in the image may be excluded from the ROI, even if their saliency scores would otherwise cause them to be included in the ROI. Thus, in such an example, if two human subjects of images 104 and 106 located on the left are recognized by the user's device (e.g., via a stored database of facial models of individuals known to the user), while the human subject of image 102 located on the right is not recognized, the human subject of image 102 located on the right may be excluded from the ROI, which may reduce the size of ROI 122 / 132 such that the determined cropping region 212 may be able to successfully or acceptably crop image 210 in a portrait orientation (i.e., cropping of the entire reduced-size ROI 122 / 132 that excludes the human subject located on the right side of the image). In other cases, other heuristics may be employed, such as modifying the crop region based on the most visually salient person, e.g., the person with the largest face in the image (rather than the most important person or the most closely related identified person).
[0044] Now go to Figure 2B , which illustrates an exemplary crop region with a focus region according to one or more embodiments. As described above, the focus region may include further specifications of a portion of the determined crop region (e.g., the bottom 75% of the crop region, the bottom 50% of the crop region, etc.), wherein the crop score of the determined region is further determined based at least in part on the amount of the first ROI and / or the second ROI enclosed by the first focus region. In other words, if the portion of the first ROI and / or the second ROI included in the determined crop region extends beyond the specified focus region, it may adversely affect the crop score of the determined crop region.
[0045] Figure 2B The image 250 in FIG. 2 shows a successful lateral crop region 256 that adopts the bottom 50% focus region based on the first ROI, i.e., it is expected that the first ROI 122 does not extend beyond the bottom 50% of the determined crop region 256. Figure 2ACompared to crop region 202 shown above, the width dimension of crop region 256 has been slightly reduced from the entire extent of image 250 in order to define crop region 256 in which first ROI 122 (i.e., primarily containing the faces of the three human subjects in the image) is completely contained within the bottom 50% of defined crop region 256, as defined by horizontal line 252. Shaded region 254 above horizontal line 252, i.e., the upper 50% of defined crop region 256, can now be safely reserved for covering text, titles, clocks, battery indicators, or other display elements that may be present on the display screen of the device during normal operation, without obscuring the essential subject matter of the image appearing in the crop region (i.e., the content of the image within first ROI 122).
[0046] In contrast, Figure 2B 260 shows a failed vertical crop region 266, which is based on the same bottom 50% focus area constraint of the first ROI 122, i.e., it is expected that the first ROI 122 does not extend beyond the bottom 50% of the determined crop region 266. It can be seen that in order to ensure that the content of the first ROI 122 only appears in the bottom 50% of the determined crop region 266 and does not appear within the shadow region 264 (as defined by the horizontal line 262), the size of the determined crop region 266 must become very small. In fact, the determined crop region 266 is so small that it again fails to meet the exemplary minimum score threshold based on covering the entire first ROI 122. Therefore, with Figure 2A As with image 210 , the attempted vertical crop of image 250 having the bottom 50% of the focus area fails.
[0047] Some implementations may also set minimum resolution requirements for the determined crop regions in order for them to also be considered successful. For example, if the determined crop region must be sized to a 600 pixel by 400 pixel region on the first image in order to satisfy various ROI and / or focus region crop criteria in place in a given crop request, then the method may not suggest or recommend the determined crop to a device display or designated content region having a resolution greater than a predetermined multiple of one or more dimensions of the determined crop. For example, if the device display (or designated content area) requesting the crop has a target size of 1200 pixels by 800 pixels (or larger), i.e., a horizontal rectangular crop area with a 3:2 aspect ratio, then a determined crop area of 600 pixels by 400 pixels may be considered too small to be used as a background image (or for use within the designated content area) even if it otherwise meets all other cropping criteria, because enlarging the crop area too much to fit on the device's display (or within the designated content area) as a background image may also result in visually unpleasant results, i.e., even if significant content from the image is included in the crop, the crop area may be too blurry or jagged due to enlargement to be used well as a background image (or for use within the designated content area). As can now be appreciated, the requested target size, aspect ratio, orientation, image resolution, and minimum score threshold, as well as the actual size and location of significant content in the image, may all have a significant impact on whether a determined crop area for a given image can be considered successful and / or recommendable for use by a requesting end user or application.
[0048] Clipping fraction
[0049] As described above, the crop score for each crop region may be determined based on any number of desired criteria, such as whether an identified ROI is encompassed by the crop region, the relative importance of the ROI (e.g., based on the type of object or person present), the total number of image pixels encompassed by the crop region, the percentage of total image pixels encompassed by the crop region, the size of the crop region, the user's likely familiarity with the location where the image was taken, and the like.
[0050] Now go to Figure 3A , which shows a graph 300 of exemplary cropping scores according to one or more embodiments. Figure 3AIn the example of , the crop score of the determined crop region is based at least in part on whether the determined crop region covers the defined first ROI and / or second ROI (and to what extent). As shown on the left side of the horizontal axis of the graph 300, if no pixels from the image are covered in the determined crop region, this will be equivalent to a crop score of 0% on the vertical axis of the graph 300. At the other extreme, as shown on the right side of the horizontal axis of the graph 300, if all pixels from the image are covered in the determined crop region, this will be equivalent to a perfect crop score of 100% on the vertical axis of the graph 300. Between these end values on the horizontal axis, various threshold scores can be specified. For example, as shown in the graph 300, if the determined crop region covers all of the first ROI (i.e., the inner region or tighter crop, including all the parts of the image that are considered "essential"), the crop region will be assigned a first minimum score, such as 50%. Moving to the right along the horizontal axis, if the determined crop region encompasses all of the second ROI (i.e., the outer region or looser crop, containing all portions of the image considered “essential” and all portions considered “preferred”), the crop region will be assigned a second minimum score greater than the first minimum score, e.g., 75%.
[0051] If the amount of the ROI encompassed by the determined crop region is somewhere between the range of the first ROI and the range of the second ROI, the crop score may be determined by applying an interpolation (e.g., linear interpolation) between the first minimum score (e.g., 50%) and the second minimum score (e.g., 75%), such as by referring to Figure 3B , shown in more detail. Similarly, if the amount of the ROI encompassed by the determined crop region is somewhere between no pixels and the extent of the first ROI, the crop score may be determined by applying interpolation (e.g., linear interpolation) between 0% and a first minimum score (e.g., 50%). (As described above, in some implementations, a coverage less than the first ROI may not result in a crop score that would exceed the minimum score threshold. Thus, the interpolation step may be avoided, and the determined crop region may be simply rejected as not encompassing enough of the essential portion of the image.) Similarly, if the amount of the ROI encompassed by the determined crop region is somewhere between the extent of the second ROI and the entire extent of the image, the crop score may be determined by applying interpolation (e.g., linear interpolation) between a second minimum score (e.g., 75%) and a score of 100%. It should be understood that other functions (e.g., non-linear functions), lookup tables (LUTs), thresholds, rules, etc. may be used to map from a value indicating the amount of the first image encompassed by the determined crop region to a crop score, as desired for a given implementation.
[0052] Now go to Figure 3B, which illustrates an exemplary interpolation technique 350 / 360 for determining a clipping score according to one or more embodiments. Figure 3A As described, according to some embodiments, the crop fraction for a given determined crop region may be determined via one or more interpolation processes. For example, looking at image 350, the determined crop region 352 encompasses the entire vertical extent of the first ROI 122 and the second ROI 132, but is positioned approximately halfway between the horizontal extents of the first ROI 122 and the second ROI 132. As shown below image 350, applying the above Figure 3A 300, if the crop region 352 encompasses only the entire first ROI 122, it will be assigned a crop score of 50% (354). Similarly, if the crop region 352 encompasses only the entire second ROI 132, it will be assigned a crop score of 75% (358).
[0053] However, as shown, the crop region 352 extends halfway between the left side of the first ROI 122 and the left side of the second ROI 132. Likewise, because the crop region 352 has been horizontally centered over the ROIs, it extends halfway between the right side of the first ROI 122 and the right side of the second ROI 132. Therefore, performing linear interpolation between the first lowest crop score of 50% (354) and the second lowest crop score of 75% (358), the determined crop region 352 may be assigned a crop score that is halfway between the first lowest crop score of 50% (354) and the second lowest crop score of 75% (358), i.e., a score of 62.5% (356).
[0054] Turning now to image 360, the determined crop region 362 again encompasses the entire vertical extent of the first ROI 122 and the second ROI 132 and the horizontal extent of the first ROI 122, but is positioned approximately halfway between the horizontal extent of the second ROI 132 and the outer extent of the image 360. As shown below image 360, applying the above Figure 3A 300, if the crop region 362 encompasses the entire second ROI 132, it will be assigned a crop score of 75% (364). Similarly, if the crop region 362 encompasses the entire image 360, it will be assigned a crop score of 100% (368).
[0055] However, as shown, the crop region 362 extends halfway between the left side of the second ROI 132 and the left side of the image 360. Likewise, because the crop region 362 has been horizontally centered above the ROI, it extends halfway between the right side of the second ROI 132 and the right side of the image 360. Therefore, by performing linear interpolation between the second lowest crop score of 75% (364) and the highest crop score of 100% (368), the determined crop region 362 may be assigned a crop score that is halfway between the second lowest crop score of 75% (364) and the highest crop score of 100% (368), i.e., a score of 87.5% (366).
[0056] like Figure 3B As shown, the determined crop fraction applies only to the horizontal extent of the determined crop region. It should be understood that a similar crop fraction may also be determined for the vertical extent of each determined crop region. Therefore, while a given image may have a 100% crop fraction in one dimension, it may not have a 100% fraction in another dimension (e.g., unless the desired aspect ratio of the crop exactly matches the image). Then, in some implementations, the final crop fraction of the image may be the smaller of the crop fractions calculated for the vertical extent and the horizontal extent of the image. In other implementations, the larger of the vertical crop fraction and the horizontal crop fraction, the average of the vertical crop fraction and the horizontal crop fraction, or some other combination may be used to determine the final crop fraction of the image.
[0057] It should be understood that the above reference Figure 3A and 3B The crop fraction scheme detailed is only one possible such scheme, and other methods may be employed to determine and / or use the crop fraction for a given crop region as desired for a given implementation.
[0058] For example, in some crop score schemes, content within an image may be assigned separate rankings and / or weighting factors (e.g., broken down by pixel, by ranked region, by object, etc.), and then a crop region may be determined to attempt to maximize the scores of pixels within the crop region (e.g., by summing together all of the determined scores for pixels, regions, etc. encompassed by the crop region). In such schemes, the final crop score for the determined crop region may be calculated, for example, as the sum of the percentage of each ranked region encompassed in the determined crop multiplied by the corresponding weighting factor for that region. For example, if the "food" objects in a given image are given the highest ranking and weighting factor of 100, while the "people" objects in the given image are given a secondary ranking and weighting factor of 25, then a cropped region determined to include all people in the image but only half of the food objects will receive a score of 75 (i.e., 25*1.0+100*0.5), while a cropped region determined to not include people in the image but include all food objects will receive a score of 100 (i.e., 25*0.0+100*1.0), and therefore become the higher-scoring cropped region, based on the specified scoring scheme in this example that favors food-based content in the image—even if all human subjects are omitted from the cropped region.
[0059] Based on the above examples, it will be appreciated that the example described above with two ROIs (i.e., an inner region and an outer region) is merely exemplary, and that any number of weighted score thresholds may be used, for example, to identify far more than two ROIs (e.g., a first ROI includes cropped regions that would have a score of 100 or higher, a second ROI includes cropped regions that would have a score of 75 or higher, a third ROI includes cropped regions that would have a score of 50 or higher, and so on), and that such ROIs may overlap, at least partially overlap, or not overlap at all within the image depending on the weighting scheme assigned and the layout of objects in the scene. Furthermore, the ROIs within a given image may change over time, e.g., if a given scheme assigns a weighting factor of 200 to areas of an image that include the face of a person identified in the image, then when the image is first captured, the area of the image that contains unknown "Person A" may not be part of the first ROI (i.e., the most important area), but if "Person A" is identified at a later time and added to the user's database of identified people, then when the crop score for the image is determined again at a later time, the area of the image that contains the now known "Person A" may become part of the first ROI because it will be scored much higher because this person is now included in the identified people.
[0060] In some embodiments, for example, if the areas of "essential" and / or "preferred" content within the image happen to be discontinuous (e.g., where a highly salient content area is at the left edge of the image, other equally highly salient content is at the right edge of the image, and less salient content is in the center portion of the image), multiple candidate areas may be identified for use as the first ROI and / or second ROI. In such cases, the final crop score may actually be viewed as the best score, worst score, or average score among all candidate selections for the first ROI and the second ROI. In other words, if the scoring scheme can accept a ranked and weighted list of ROIs, then in addition to the final crop score, the scoring scheme may also provide information about the extent to which each candidate ROI is captured by the final crop area.
[0061] In some embodiments, the cropping score of a given image may be used in real time to determine which type of cropping area (and / or the number of cropping areas) will be presented and incorporated into the specified content area of the device UI of each given image. For example, an application that presents graphical information to a device UI may face a decision about whether it should display a single rectangular cropping of an image in the specified content area of the application's UI or two square croppings of two different images that occupy the same total space as a single rectangular picture of the specified content area. If the square aspect ratio cropping scores of the two images in this example are relatively close (e.g., within a predetermined relative cropping score similarity threshold), one option may be to display the two images as side-by-side squares in the specified content area of the UI of the application or device. In contrast, if the rectangular aspect ratio cropping score of an image when cropped to a single rectangular image is significantly higher (e.g., greater than a predetermined relative cropping score difference threshold), it may be a better choice to display the one image as a single rectangular picture in the specified content area of the UI of the application or device. Note that in a given situation, display and application characteristics, such as those mentioned above (e.g., size, orientation, aspect ratio, resolution, etc.), may also play a role in this decision of how many images to display (and which crops of such images) in a specified content area. For example, if a single rectangular image is to be displayed on a high-resolution TV screen, the decision may be to display two square images in the specified content area, because the single image may not have a high enough resolution to be used as a single image on the TV. However, it may be determined that the same content (i.e., the same two images from the above example) should be displayed as a single image on the phone, because the resolution of the first of the two images may be of sufficient quality in the context of the specified content area on the relatively small display screen of the phone. Note also that the intelligent cropping techniques discussed herein may enable the image storage / management system to store only a single source version of each piece of multimedia content, and to select how to crop, layout, and display such content "on the fly" (i.e., in real time or near real time), for example, based on the specific display device, orientation, resolution, available screen space, specified content area, etc.
[0062] In other embodiments, the device and / or application may use the crop score to make intelligent decisions about which potential crops to use in a given situation, for example, based on the designated content area available for display in that situation. For example, if there is a large enough designated content area in which the device or application wishes to display content, it may be desirable to have a higher crop score quality threshold for content selected to appear in that area. In contrast, for smaller designated content areas, a lower crop score quality threshold may be used because such content is more likely to be accompanied by other content of equal or higher crop scores on the display UI.
[0063] In other embodiments, other auxiliary information (e.g., the user's likely familiarity with the location where the image was taken) may be used to determine and score the crop region. For example, if the image is of a scenic vacation location (e.g., a location that the user does not visit often or does not have a large number of images of), the crop score may be further penalized for determining a crop region that crops out a large portion of the original image, while if the image is of a scenic location near the user (e.g., a location that the user visits often or the user already has a large number of other images of the location in their multimedia library), the crop score may be assigned less penalization for determining a crop region that crops out a larger portion of the original image because the user is likely already familiar with the location shown in the image.
[0064] Exemplary Smart Image Cropping Operation
[0065] See now Figure 4 , which shows a flowchart of a method 400 of an exemplary method for performing automatic image cropping according to one or more embodiments. First, at step 402, method 400 may obtain a first image. Next, at step 404, method 400 may receive a first cropping request, wherein the first cropping request includes: a first target size, from which a first aspect ratio and a first orientation can be determined. Allowing specification of a target size (rather than an explicit aspect ratio and orientation) will allow the aspect ratio and orientation to be inferred. In addition, it will also allow the previously discussed minimum resolution cropping constraint scenario. Next, at step 406, method 400 may determine a first region of interest (ROI) of the first image, for example using any of the aforementioned saliency or object detection-based techniques.
[0066] Next, at step 408, method 400 may determine a first crop region of the first image based on the first crop request, for example, wherein the first crop region has a first width, a first height, a first position within the first image, and surrounds a first subset of content in the first image (step 410), and wherein at least one of the first width, the first height, and the first position is at least partially determined to maximize an amount of overlap between the first crop region and the first ROI (step 412).
[0067] Next, at step 414, method 400 may determine a first score for the first crop region, wherein the first score is determined at least in part based on an amount of overlap between the first crop region and the first ROI. Finally, at step 416, when it is determined that the first score is greater than a minimum score threshold, method 400 may crop the first crop region from the first image.
[0068] See now Figure 5, which shows a flow chart of method 500 illustrating another method of performing automatic image cropping according to one or more embodiments. Method 500 is similar to method 400, however, method 500 details a scenario in which multiple ROIs are defined on a first image, and an optional specification of a focus area within the determined cropping region.
[0069] First, at step 502, method 500 may obtain a first image. Next, at step 504, method 500 may receive a first crop request, wherein the first crop request includes: a first target size, and optionally a specification of a focus area, from which a first aspect ratio and a first orientation may be determined. Next, at step 506, method 500 may determine a first region of interest (ROI) and a second ROI of the first image, for example, wherein the second ROI may optionally be a superset of the first ROI (i.e., completely surround the first ROI).
[0070] Next, at step 508, method 500 may determine a first crop region of the first image based on the first crop request, for example, wherein the first crop region has a first width, a first height, a first position within the first image, and contains a first subset of content in the first image (step 510), and wherein at least one of the first width, the first height, and the first position is at least partially determined to maximize the amount of overlap between the first crop region and the first ROI and / or the second ROI (step 512). For example, as described above, given the constraints of the image size and the target size of the first crop request, some smart cropping schemes may prioritize overlapping with the entire first ROI and then also attempt to overlap with as much of the second ROI as possible.
[0071] Next, at step 514, method 500 may determine a first score for the first crop region, wherein the first score is determined at least in part based on an amount of overlap between the first crop region and the first ROI and the second ROI (and, optionally, an amount of the first ROI and the second ROI that can be contained in the first focus region), wherein if the first ROI is completely enclosed in the first crop region (and, optionally, also enclosed within a first focus region of the first crop region), then the first score is at least a first minimum score, wherein if the second ROI is completely enclosed in the first crop region (and, optionally, also enclosed within a first focus region of the first crop region), then the first score is at least a second minimum score, and wherein the second minimum score is greater than the first minimum score.
[0072] Finally, at step 516 , when it is determined that the first score is greater than the minimum score threshold, method 500 may crop a first crop region from the first image.
[0073] Exemplary Electronic Computing Devices
[0074] See now Figure 6 , which shows a simplified functional block diagram of an exemplary programmable electronic computing device 600 according to an embodiment. The electronic device 600 can be a system such as a mobile phone, a personal media device, a portable camera, or a tablet, a laptop or a desktop computer. As shown, the electronic device 600 may include a processor 605, a display 610, a user interface 615, a graphics hardware 620, a device sensor 625 (e.g., a proximity sensor / ambient light sensor, an accelerometer, an inertial measurement unit and / or a gyroscope), a microphone 630, an audio codec 635, a speaker 640, a communication circuit 645, an image capture device 650 (e.g., it may include multiple camera units / optical image sensors with different characteristics or capabilities (e.g., static image stabilization (SIS), high dynamic range (HDR), optical image stabilization (OIS) system, optical zoom and digital zoom, etc.), a video codec 655, a memory 660, a storage device 665 and a communication bus 670.
[0075] The processor 605 may execute instructions necessary for implementing or controlling the operation of various functions performed by the electronic device 600 (e.g., generation and / or processing of images according to various embodiments described herein). The processor 605 may, for example, drive the display 610 and may receive user input from the user interface 615. The user interface 615 may take a variety of forms, such as buttons, keypads, dials, click wheels, keyboards, display screens, and / or touch screens. The user interface 615 may, for example, be a conduit through which a user can view a captured video stream and / or indicate a specific image frame that the user wants to capture (e.g., by clicking a physical or virtual button at the moment when the desired image frame is being displayed on the display screen of the device). In one embodiment, the display 610 may display a video stream that is captured while the processor 605 and / or graphics hardware 620 and / or image capture circuitry are simultaneously generating the video stream and storing the video stream in the memory 660 and / or storage device 665. The processor 605 may be a system on a chip (SOC) such as those found in mobile devices, and may include one or more dedicated graphics processing units (GPUs). Processor 605 may be based on a reduced instruction set computer (RISC) or complex instruction set computer (CISC) architecture or any other suitable architecture, and may include one or more processing cores. Graphics hardware 620 may be specialized computing hardware for processing graphics and / or assisting processor 605 in performing computing tasks. In one embodiment, graphics hardware 620 may include one or more programmable graphics processing units (GPUs) and / or one or more specialized SOCs, for example, specifically designed to implement neural network and machine learning operations (e.g., convolutions) in a more energy-efficient manner than a host device central processing unit (CPU) or a typical GPU (such as Apple's Neural Engine processing core).
[0076] For example, according to the present disclosure, the image capture device 650 may include one or more camera units configured to capture images, for example, images that can be processed to generate smart cropped versions of the captured images. In some cases, the smart cropping techniques described herein may be integrated into the image capture device 650 itself, so that even before shooting, the camera unit may be able to convey high-quality framing selections of possible images to the user. The output from the image capture device 650 may be processed at least in part by the following devices: a video codec 655 and / or a processor 605 and / or graphics hardware 620, and / or a dedicated image processing unit or image signal processor incorporated in the image capture device 650. Such captured images may be stored in a memory 660 and / or a storage device 665. The memory 660 may include one or more different types of media used by the processor 605, the graphics hardware 620, and the image capture device 650 to perform device functions. For example, the memory 660 may include a memory cache, a read-only memory (ROM), and / or a random access memory (RAM). The storage device 665 can store media (e.g., audio files, image files, and video files), computer program instructions or software, preference information, device profile information, and any other suitable data. The storage device 665 may include one or more non-transitory storage media, including, for example, disks (fixed hard disks, floppy disks, and removable disks) and tapes, optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such as electrically programmable read-only memories (EPROMs), and electrically erasable programmable read-only memories (EEPROMs). The memory 660 and storage device 665 can be used to maintain computer program instructions or codes organized into one or more modules and written in any desired computer programming language. For example, when executed by the processor 605, such computer program code can implement one or more of the methods or processes described herein.
[0077] It should be understood that the above description is intended to be exemplary and not restrictive. For example, the above embodiments may be used in combination with each other. Upon reviewing the above description, many other embodiments will be apparent to those skilled in the art. Therefore, the scope of the present invention should be determined with reference to the attached claims and the full scope of equivalents to such claims.
Claims
1. A device comprising: Memory; monitor; User interface; and One or more processors operably coupled to the memory, wherein the one or more processors are configured to execute instructions that cause the one or more processors to: acquiring a first image; Receiving a first cropping request, wherein the first cropping request includes: a first target size, according to which a first aspect ratio and a first orientation can be determined; determining a first region of interest (ROI) and a second ROI of the first image using image processing, wherein the first ROI is determined to encompass all portions of the first image having a significance score greater than a first threshold, wherein the second ROI is determined to encompass all portions of the first image having a significance score greater than a second threshold, and wherein the second threshold significance score is lower than the first threshold significance score; determining a first crop region of the first image based on the first crop request, wherein the first crop region has a first width, a first height, and a first position within the first image, wherein the first crop region encloses a first subset of content in the first image, and wherein at least one of the first width, the first height, and the first position is determined to maximize an amount of overlap between the first crop region and the first ROI and an amount of overlap between the first crop region and the second ROI; determining a first score for the first crop region, wherein the first score is determined based at least in part on an amount of overlap between the first crop region and the first ROI and an amount of overlap between the first crop region and the second ROI; and When it is determined that the first score is greater than a minimum score threshold, the first cropping region is cropped from the first image.
2. The apparatus of claim 1 , wherein the instructions further cause the one or more processors to: The first crop region is used as at least one of: a screen saver, wallpaper, or background image on the display of the device; or content placed in a designated content area within a user interface of an application executing on the device.
3. The apparatus of claim 1 , wherein the instructions further cause the one or more processors to: A second cropping request is received, wherein the second cropping request comprises: a second target size, according to which the first aspect ratio and the second orientation can be determined; determining a second crop region of the first image based on the second crop request, wherein the second crop region has a second width, a second height, and a second position within the first image, wherein the second crop region surrounds a second subset of content in the first image, and wherein at least one of the second width, the second height, and the second position is determined to maximize an amount of overlap between the second crop region and the first ROI; determining a second score for the second crop region, wherein the second score is determined based at least in part on an amount of overlap between the second crop region and the first ROI; as well as When it is determined that the second score is greater than the minimum score threshold, the second cropping area is cropped from the first image. 4 . The apparatus of claim 1 , wherein the first score is further determined based at least in part on a relative distance of a boundary of the first crop region between corresponding boundaries of the first ROI and the second ROI.
5. The apparatus of claim 1 , wherein the first score is at least a first minimum score if the first ROI is completely enclosed in a first crop region, wherein the first score is at least a second minimum score if the second ROI is completely enclosed in the first crop region, and wherein the second minimum score is greater than the first minimum score.
6. The apparatus of claim 1 , wherein the first crop request further specifies a first focus region, wherein the first focus region includes a specified portion of the determined first crop region, and wherein the first score is further determined based at least in part on an amount of the first ROI enclosed by the first focus region. 7 . The apparatus of claim 1 , wherein the first ROI is determined based on one or more of: a saliency map generated based on the first image, an object detection box, a face detection box, or a face recognition box.
8. The apparatus of claim 1, wherein at least one of the first width or the first height is selected to match a corresponding dimension of the first image.
9. A non-transitory computer readable medium comprising computer readable instructions executable by one or more processors to: acquiring a first image; Receiving a first cropping request, wherein the first cropping request includes: a first target size, according to which a first aspect ratio and a first orientation can be determined; determining a first region of interest (ROI) and a second ROI of the first image using image processing, wherein the first ROI is determined to encompass all portions of the first image having a significance score greater than a first threshold, wherein the second ROI is determined to encompass all portions of the first image having a significance score greater than a second threshold, and wherein the second threshold significance score is lower than the first threshold significance score; determining a first crop region of the first image based on the first crop request, wherein the first crop region has a first width, a first height, and a first position within the first image, wherein the first crop region encloses a first subset of content in the first image, and wherein at least one of the first width, the first height, and the first position is determined to maximize an amount of overlap between the first crop region and the first ROI and an amount of overlap between the first crop region and the second ROI; determining a first score for the first crop region, wherein the first score is determined based at least in part on an amount of overlap between the first crop region and the first ROI and an amount of overlap between the first crop region and the second ROI; and When it is determined that the first score is greater than a minimum score threshold, the first cropping region is cropped from the first image.
10. The non-transitory computer readable medium of claim 9, wherein the instructions further cause the one or more processors to: A second cropping request is received, wherein the second cropping request comprises: a second target size, according to which the first aspect ratio and the second orientation can be determined; determining a second crop region of the first image based on the second crop request, wherein the second crop region has a second width, a second height, and a second position within the first image, wherein the second crop region surrounds a second subset of content in the first image, and wherein at least one of the second width, the second height, and the second position is determined to maximize an amount of overlap between the second crop region and the first ROI; determining a second score for the second crop region, wherein the second score is determined based at least in part on an amount of overlap between the second crop region and the first ROI; as well as When it is determined that the second score is greater than the minimum score threshold, the second cropping area is cropped from the first image. 11 . The non-transitory computer readable medium of claim 9 , wherein the first score is further determined based at least in part on a relative distance of a boundary of the first crop region between corresponding boundaries of the first ROI and the second ROI.
12. An image processing method, comprising: acquiring a first image; Receiving a first cropping request, wherein the first cropping request includes: a first target size, according to which a first aspect ratio and a first orientation can be determined; determining a first region of interest (ROI) and a second ROI of the first image using image processing, wherein the first ROI is determined to encompass all portions of the first image having a significance score greater than a first threshold, wherein the second ROI is determined to encompass all portions of the first image having a significance score greater than a second threshold, and wherein the second threshold significance score is lower than the first threshold significance score; determining a first crop region of the first image based on the first crop request, wherein the first crop region has a first width, a first height, and a first position within the first image, wherein the first crop region encloses a first subset of content in the first image, and wherein at least one of the first width, the first height, and the first position is determined to maximize an amount of overlap between the first crop region and the first ROI and an amount of overlap between the first crop region and the second ROI; determining a first score for the first crop region, wherein the first score is determined based at least in part on an amount of overlap between the first crop region and the first ROI and an amount of overlap between the first crop region and the second ROI; and When it is determined that the first score is greater than a minimum score threshold, the first cropping region is cropped from the first image.
13. The method of claim 12, wherein the first crop request further specifies a first focus region, wherein the first focus region includes a specified portion of the determined first crop region, and wherein the first score is further determined based at least in part on an amount of the first ROI enclosed by the first focus region.
14. The method of claim 12, wherein the first ROI is determined based on one or more of: a saliency map generated based on the first image, an object detection box, a face detection box, or a face recognition box.
15. The method of claim 12, wherein at least one of the first width or the first height is selected to match a corresponding dimension of the first image.
Citation Information
Patent Citations
Saliency of an object for image processing operations
US11308345B2
Facilitating preservation of regions of interest in automatic image cropping
US20180357803A1