Depth-aware brush size scaling for mask creation for inpainting
Depth-aware brush size scaling addresses the inaccuracies of fixed brush sizes in inpainting by dynamically adjusting the brush size based on depth maps, resulting in precise and noise-reduced image editing.
Patent Information
- Application Number
- PCT/EP2025/050124
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-17
AI Technical Summary
Existing inpainting techniques face challenges in accurately removing objects from images using fixed brush sizes, leading to inclusion of unnecessary pixels and noise due to varying object distances and dimensions, which is time-consuming and error-prone.
Implementing depth-aware brush size scaling based on depth maps generated by sensors or machine learning, dynamically adjusting the brush size according to perceived depth values as the cursor moves across the image.
Achieves more precise and less noisy inpainting by ensuring the mask region accurately covers the object, minimizing the inclusion of background pixels and reducing noise in the resulting image.
Smart Images

Figure EP2025050124_17072025_PF_FP_ABST
Abstract
Description
[0001] DEPTH-AWARE BRUSH SIZE SCALING FOR MASK CREATION FOR INPAINTING
[0002] FIELD OF THE INVENTION
[0003] This invention relates to digital image manipulation and more particularly to inpainting digital images.
[0004] BACKGROUND OF THE INVENTION
[0005] Inpainting is a technique for filling gaps in images with new values (e.g., by assigning new color values to one or more image pixels). Inpainting can be used for object removal, by creating a mask that covers the object, removing masked pixels from the original image to create a gap, and filling the gap with new color values. One way of creating a mask is by drawing onto an image using a brush. For filling a gap, techniques such as interpolation using the surrounding pixels or artificial intelligence (Al) model may be used.
[0006] SUMMARY OF THE INVENTION
[0007] Described herein are systems and methods for inpainting an image, including drawing onto the image using a brush with a rescaled size. The size of the brush is rescaled based on the depth value at one or more cursor positions on the image. The depth value is determined from a depth map, indicating depth values at different points within the image.
[0008] According to one aspect of the disclosure, a method for inpainting an image includes obtaining a depth map from the image, the depth map indicating depth values at different points within the image. In some embodiments, the method includes determining a first cursor position on the image and displaying a brush with an initial size on the image at the first cursor position. In some embodiments, in response to a cursor moving to one or more other cursor positions on the image, the method includes determining depth values at the one or more other cursor positions on the image using the depth map, and rescaling a size of the brush according to the depth values as the cursor moves to the one or more other cursor positions on the image. In some embodiments, the method includes determining a region of the image to mask based at least on the initial brush size and the rescaled brush size.
[0009] According to one aspect of the disclosure, an apparatus includes one or more processors and one or more memory elements, including instructions that cause the one or more processors to perform operations. In some embodiments, the apparatus includes obtaining a depth map from an image, the depth map indicating depth values at different points within the image. In some embodiments, the apparatus includes determining a first cursor position on the image and displaying a brush with an initial size on the image at the first cursor position. In some embodiments, in response to a cursor moving to one or more other cursor positions on the image, the apparatus includes determining depth values at the one or more other cursor positions on the image using the depth map and rescaling a size of the brush according to the depth values as the cursor moves to the one or more other cursor positions on the image. In some embodiments, the apparatus includes instructions for determining a region of the image to mask based at least on the initial brush size and the rescaled brush size.
[0010] In some embodiments, rescaling the size of the brush includes dividing the initial brush size by a depth value of a cursor position as the cursor moves to the one or more other cursor positions on the image. In some embodiments, obtaining the depth map comprises generating a depth map using a depth measuring sensor, using a stereo, and / or using machine learning (ML).
[0011] BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The manner and process of making and using the disclosed embodiments may be appreciated by reference to the figures of the accompanying drawings. It should be appreciated that the components and structures illustrated in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principals of the concepts described herein. Like reference numerals designate corresponding parts throughout the different views. Furthermore, embodiments are illustrated by way of example and not limitation in the figures, in which:
[0013] FIGS. 1 A and IB are pictorial diagrams showing an inpainting technique; FIGS. 2 A and 2B are pictorial diagrams showing an inpainting technique using brush size scaling for mask creation;
[0014] FIG. 3 is a flow chart of a method for inpainting an image; and
[0015] FIG. 4 is an example computing system for implementing inpainting using brush size scaling for mask creation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Turning to FIG. 1 A, according to a conventional inpainting technique, a user can draw a mask region 110 on an image 100 (e.g., a digital photograph) using paint brush 102 having a fixed brush size (e.g., a brush with a radius Rl). In this example, it is assumed that a user desires to remove a series of luminaires 122a, 122b, etc. (122 generally) from the image 100 using inpainting, while leaving other regions of the image (e.g., background region 120) unaltered. In general, inpainting can be used to remove any type of object from an image, with luminaires being merely one example.
[0017] The mask region 110 is generated as brush 102 is applied to one or more positions on the image 100. In the example shown, an approximately linear mask region 110 is generated as brush 102 moves (e.g., is dragged) between position 103 and position 104. Brush 102 can be controlled by user input, such as input generated by a mouse, trackball, touchscreen (e.g., a user applying pressure thereto), stylus, etc. Inpainting can be done using various brush types / shapes. In the example shown, brush 102 is a circular brush having a radius Rl. The brush center 130 may correspond to a cursor / pointer whose position is controlled by user input. Thus, for convenience, the terms “brush position” and “cursor position” may be used interchangeably herein. In other words, the position of the cursor / pointing, controlled by user input, determines the position of brush 102. Brush 102 may be initially displayed at position 103, with fixed brush size Rl, in response to user input at or near that position on the image.
[0018] When inpainting to remove one or more objects from an image, such as image 100 of FIG. 1 A, it may be desirable to mask only the pixels of the image corresponding to those objects. This prevents other pixels that do not corresponding to desired objects from being inpainted. The size of the brush determines the mask region 110. When a brush is too large, more of the surrounding pixels may be included in the mask region than are necessary. It is appreciated herein that this is especially problematic when inpainting objects that appear further away in the image, as they are relatively smaller and as a result a relatively smaller brush size must be used to precisely mask the smaller object. It is further appreciated that requiring a user to select between different brush sizes can be time consuming and error prone.
[0019] For example, mask region 110 of FIG. 1 A includes multiple luminaires 122, some of which may occupy more pixels within the image compared to others, e.g., a first luminaire 122a may occupy more pixels than second luminaire 122b. This could be the result of luminaire 122b appearing further away within the image compared to luminaire 122a (even though the two luminaires have the same physical dimensions) or because the luminaires have different physical dimensions. Even if the brush size is selected to perfectly match the size of first luminaire 122a within the image 100, the resulting mask region 110 will be larger than need to mask the second luminaire 122b (because the brush has a fixed size R1 in this example).
[0020] Referring to FIG. IB, an image 140 may be the result of inpainting the masked image 100 of FIG. 1A. In more detail, an inpainted region 160 of FIG. IB may result from a user inpainting some or all of mask region 162 (which may be the same as mask region 110 of FIG. 1A). In other words, images 100 and 140 may be identical (e.g., in terms of pixel values) except for implanted region 160. The image 140 comprises a background region 150 that is not included in the inpainted region 160. As seen, within the inpainted region 160, pixel values may be replaced so as to remove objects from the image, such as the luminaires 122a, 122b of FIG. 1 A. As can also been seen, the inpainted region 160 may be larger than necessary to remove certain luminaires 122 (e.g., luminaires that appear further away) due to the fixed brush size. This is exemplified in region 164, where more pixels of the image may be inpainted than necessary to remove the objects and may introduce noise into the resulting image 140. For example, portions of background region 150 may be undesirably removed.
[0021] Turning to FIG. 2A, according to embodiments of the present disclosure, a user can draw a mask region 210 on an image 200 (e.g., a digital photograph) using paint brush 202 that dynamically scales in a depth-aware manner. It is assumed that the user desires to remove a set of luminaires 222a, 222b, etc. (222 generally) from the image 200 using inpainting, while leaving other regions of the image (e.g., background region 120) unaltered. One or more luminaires 222a, 222b, are shown in the image 200, including a first luminaire 222a and a second luminaire 222b. The mask region 210 is generated as brush 202 moves between position 203 and position 204.
[0022] Illustrative brush 202 is a circular brush having a center 230. Brush 202 can be controlled (e.g., dragged between positions 203 and 204) using, for example, any of the user input techniques previously discussed.
[0023] The size of brush 202 may be automatically changed (i.e., scale up or down) as it is moved between positions 203 and 204. For example, at position 203 the brush 202 may have an initial size defined by radius R2, whereas at position 204 the brush 202 may have a different size defined by radius R3. Dynamically scaling brush size in this manner can allow for the creation of a mask region 210 that includes fewer undesired pixels compared to using a fixed brush size. That is, can allow for more accurate and less noisy inpainting. The initial brush size, in this case having a radius R2, may be determined a number of different ways. For example, the initial brush size may be determined based on a user selection. The brush size may be adjusted by the user (e.g., through a toggle or a sliding bar that adjusts the size of the brush), the size selected by the user may be used as the initial brush size.
[0024] A depth map may be used to dynamically scale the brush size as it is moved across an image. A depth map is a data structure that indicates the perceived depth at different points within the image 200 (e.g., at every pixel or at points corresponding to lower resolution features of the image). For each point, the depth map stores a distance between that point and a camera in three dimensional (3D) space. The camera can be a physical camera used to capture the image or a virtual camera used to generate the image (e.g., using ray tracing). The further the point is located from the mechanism forming the image, the larger its depth value can be. Alternatively, smaller depth values may be used to represent larger depths.
[0025] The depth map can be generated a number of different ways, such as through the use of a depth measuring sensor, stereoscopy, and / or machine learning (ML).
[0026] A depth measuring sensor may be incorporated into a digital camera used to capture image 200. A depth measuring sensor may include a light detection and ranging (LiDAR) sensor, which send out infrared radiation and measures how long it takes for this radiation to return (with a known speed of light, it can be determined how far the object was).
[0027] With stereoscopy, depth values can be found using multiple different images captured of the same scene and measuring the distance between the images. For example, two cameras are positioned at a known distance apart from each other and used to take two images of the same scene. By finding the same objects in both images, it can be determined how much said objects have shifted between the two images. Objects that is closer to the camera will have shifted more, while objects further away will have shifted less. This analysis can be performed for each pixel, resulting in a disparity map that is inversely proportional to a depth map. The depth map can be calculated from the disparity map, as the distance between the cameras is known. This process is called computer stereo vision or depth from stereo images. In some cases, a device having two or more cameras, a fixed distance apart, within the same housing may be used to generate a depth map.
[0028] With ML, the depth map can be generated from an image using an Al model that is trained to predict depths at different points (e.g., pixels) within the image. In the case of a depth measuring sensor or stereoscopy, the depth map may be generated at the time the image is captured and stored along with the captured image data (e.g., within the same image file). In the case of ML, a depth map may be generated “on the fly” when the image is accessed for inpainting (e.g., when the image is loaded by an image editing application).
[0029] The initial brush size may be determined by first determining the depth value of the image at the first position 203 using the depth map, then determining the initial brush size based on said depth value. As another example, the initial brush size may be determined based on a predetermined default size and a first depth value.
[0030] The brush size may be rescaled by moving the brush to one or more other positions 203, 204 on the image 200. Once the depth map has been obtained, the depth value of the positions 203, 204 can be determined on the image 200. Once determined, the brush size is rescaled according to the depth value. Rescaling the brush size may include dividing or multiplying the initial brush size (e.g., radius R2) by a depth value of the position 203 as the brush moves to the one or more other positions on the image 200. For example, once the brush moves the position 204 on the image 200, the depth value at the position 204 is determined and brush 202 may be rescaled according to that value (e.g., from radius R2 to radius R3).
[0031] The brush may be rescaled to account for other brush positions. For example, the mask region 210 includes multiple luminaires 222, some of which may occupy more pixels within the image compared to others, e.g., the first luminaire 222a may occupy more pixels than second luminaire 222b. This could be the result of luminaire 222b appearing further away within the image compared to luminaire 222a (even though the two luminaires have the same physical dimensions) or because the luminaires have different physical dimensions. The brush may be positioned at the first luminaire 222a. The size of the brush may then be rescaled in accordance with the determined depth value at the brush position. The brush may be positioned at the second luminaire 222b. The brush size at the second luminaire 222b may be smaller than the brush size at the first luminaire 222a.
[0032] Referring to FIG. 2B, an image 240 may be the result of inpainting the masked image 200 of FIG. 2A. In more detail, an inpainted region 260 of FIG. 2B may result from a user inpainting some or all of mask region 262 (which may the be same as mask region 210 of FIG. 2A). In other words, images 200 and 240 may be identical (e.g., in terms of pixel values) except for inpainted region 260. The image 240 comprises a background region 250 that is not included in the inpainted region 260. As seen, within the inpainted region 260, pixel values may be replaced so as to remove objects from the image, such as luminaires 222a, 222b of FIG. 2 A.
[0033] As can also be seen, the inpainted region 260 covers more appropriately the pixels values corresponding to the objects from the image, such as luminaires 222a, 222b of FIG. 2A, due to the rescaled brush size. This is exemplified in region 264, where an appropriate amount of pixel values of the image may be inpainted to remove the objects and may result in a reduction of noise in the resulting image 240 (reduced in comparison to the region 164 of FIG. IB). For example, less portions of the background region 250 may be undesirable removed (less portions in comparison to the region 164 of FIG. IB).
[0034] FIG. 3 shows an example of a method 300 for inpainting an image. At block 302, a depth map is obtained from an image, the depth map indicating perceived depths at different points within the image. At block 304, a first cursor position (or, equivalently, brush position) on the image can be determined . At block 306, a brush can be displayed with an initial size on the image at the first cursor position.
[0035] At block 308, in response to the cursor moving to one or more other cursor positions on the image, method 300 can perform blocks 310-314, described next.
[0036] At block 310, depth values can be determined at the one or more other cursor positions on the image using the depth map. At block 312, the size of the brush may be rescaled according to the depth values as the cursor moves to the one or more other cursor positions on the image. At block 314, a region of the image to mask may be determined based at least on the initial and rescaled brush sizes.
[0037] FIG. 4 shows a block diagram of a representative computing system 414 usable to implement the present disclosure. The method for inpainting to generate the image 240 of FIG. 2B may be implemented by the computing system 414. Computing system 414 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (for example, a smart watch, eyeglasses, or a head wearable display), desktop computer, laptop computer, or implemented with distributed computing devices. The computing system 414 may include conventional computer components, such as a processor 416, a storage device 418, a network interface 420, a user input device 422, and a user output device 424.
[0038] Network interface 420 can provide a connection to a wide area network (for example, the Internet) to which a WAN interface of a remote server system is also connected. Network interface 420 can include a wired interface (for example, ethernet) and / or a wireless interface implementing various RF data communication standards, such as Wi-Fi, Bluetooth, or cellular data network standards (for example, 3G, 4G, 5G, 60 GHz, or LTE).
[0039] User input device 422 can include any device via which a user can provide signals to computing system 414, which in turn can interpret the signals as indicative of particular user requests or information. User input device 422 can include any or all of a keyboard, a touch pad, a touch screen, a mouse or other pointing device, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, a microphone, a sensor (for example, a motion sensor or an eye tracking sensor).
[0040] User output device 424 can include any device via which computing system 414 can provide information to a user. For example, user output device 424 can include a display to display images generated by or delivered to computing system 414. The display can incorporate various image generation technologies, for example, a liquid crystal display (LCD), a light-emitting diode (LED), such as an organic light-emitting diode (OLED), a projection system, a cathode ray tube (CRT), or the like, together with supporting electronics (for example, digital-to-analog or analog-to-digital converters, or signal processors). A device such as a touch screen that functions as both input and output device can be used. User output devices 424 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.
[0041] Some implementations include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer readable storage medium (for example, a non-transitory computer readable medium).
[0042] Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer readable storage medium. When these program instructions are executed by one or more processors, they cause the processors to perform various operations indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processor 416 can provide various functionality for computing system 414, including any of the functionality described herein as being performed by a server or client, or other functionality associated with message management services.
[0043] It will be appreciated that computing system 414 is illustrative and that variations and modifications are possible. Computer systems used in connection with the present disclosure can have other capabilities not specifically described here. Further, while computing system 414 is described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, for example, by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Implementations of the present disclosure can be realized in a variety of apparatus, including electronic devices implemented using any combination of circuitry and software.
[0044] Although reference is made herein to particular materials, it is appreciated that other materials having similar functional and / or structural properties may be substituted where appropriate, and that a person having ordinary skill in the art would understand how to select such materials and incorporate them into embodiments of the concepts, techniques, and structures set forth herein without deviating from the scope of those teachings.
[0045] It is to be understood that the disclosed subject matter is not limited in its application to the details of construction and to the arrangements of the components set forth in the following description or illustrated in the drawings. The disclosed subject matter is capable of other embodiments and of being practiced and carried out in various ways. Also, it is to be understood that the phraseology and terminology employed herein are for the purpose of description and should not be regarded as limiting. As such, those skilled in the art will appreciate that the conception, upon which this disclosure is based, may readily be utilized as a basis for the designing of other structures, methods, and systems for carrying out the several purposes of the disclosed subject matter. Therefore, the claims should be regarded as including such equivalent constructions insofar as they do not depart from the spirit and scope of the disclosed subject matter.
[0046] Although the disclosed subject matter has been described and illustrated in the foregoing exemplary embodiments, it is understood that the present disclosure has been made only by way of example, and that numerous changes in the details of implementation of the disclosed subject matter may be made without departing from the spirit and scope of the disclosed subject matter.
Claims
CLAIMS1. A method for inpainting an image (300), comprising: obtaining a depth map from the image, the depth map indicating depth values at different points within the image (302); determining a first cursor position on the image (304); displaying a brush with an initial size on the image at the first cursor position (306); in response to a cursor moving to one or more other cursor positions on the image (308): determining depth values at the one or more other cursor positions on the image using the depth map (310), and rescaling a size of the brush according to the depth values as the cursor moves to the one or more other cursor positions on the image (312); and determining a region of the image to mask based at least on the initial brush size and the rescaled brush size (314).
2. The method of claim 1, further comprising generating another image by inpainting the region of the image to mask.
3. The method of claim 1, wherein rescaling the size of the brush includes dividing the initial brush size by a depth value of a cursor position as the cursor moves to the one or more other cursor positions on the image.
4. The method of claim 1, comprising: determining a first depth value of the image at the first cursor position using the depth map; and determining the initial brush size based on the first depth value.
5. The method of claim 4, comprising determining the initial brush size based on a predetermined default size and the first depth value.
6. The method of claim 1, comprising determining the initial brush size based on a user selection.
7. The method of claim 1, wherein obtaining the depth map comprises generating a depth map using a depth measuring sensor, using stereoscopy, and / or using machine learning (ML).
8. The method of claim 1, wherein the cursor control corresponds to a trackball, a mouse, or a user applying pressure to a touch screen.
9. An apparatus ( 14), comprising: one or more processors (416); one or more memory elements (418) including instructions that cause the one or more processors to perform operations, including: obtaining a depth map from an image, the depth map indicating depth values at different points within the image (302); determining a first cursor position on the image (304); displaying a brush with an initial size on the image at the first cursor position (306); in response to a cursor moving to one or more other cursor positions on the image (308): determining depth values at the one or more other cursor positions on the image using the depth map (310), and rescaling a size of the brush according to the depth values as the cursor moves to the one or more other cursor positions on the image (312); and determining a region of the image to mask based at least on the initial brush size and the rescaled brush size (314).
10. The apparatus of claim 9, further comprising generating another image by inpainting the region of the image to mask.
11. The apparatus of claim 9, wherein rescaling the size of the brush includes dividing the initial brush size by a depth value of a cursor position as the cursor moves to the one or more other cursor positions on the image.
12. The apparatus of claim 9, comprising: determining a first depth value of the image at the first cursor position using the depth map; and determining the initial brush size based on the first depth value.
13. The apparatus of claim 12, comprising determining the initial brush size based on a predetermined default size and the first depth value.
14. The apparatus of claim 9, comprising determining the initial brush size based on a user selection.
15. The apparatus of claim 9, wherein obtaining the depth map comprises generating a depth map using a depth measuring sensor, using stereoscopy and / or using machine learning (ML).
Citation Information
Patent Citations
Dynamically adjusted brush for direct paint systems on parameterized multi-dimensional surfaces
US20070115295A1
System and method for retexturing of images of three-dimensional objects
US20170032563A1