Depth estimation based on iris size

By identifying faces and estimating depth in the image of a single image capture component, the cost and complexity problem of using multiple image capture components in the prior art is solved, and efficient image depth estimation and image enhancement effects are achieved.

CN114945943BActive Publication Date: 2025-05-06GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080093126.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-13
Filing Date
2020-05-21
Publication Date
2025-05-06
Estimated Expiration
2040-05-21

AI Technical Summary

Technical Problem

Prior art When acquiring image depth information, it is often necessary to use multiple image capture components, resulting in increased cost and complexity.

Method used

By identifying faces in images of a single image capture component and generating a face grid, estimating eye pixel sizes, combining the camera's inherent matrix and average eye size, the depth of face relative to the camera is calculated.

Benefits of technology

This enables the use of a single image capture component to acquire image depth information, reducing cost and complexity, while enhancing image processing effects, such as generating images focused on people and blurring the background.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114945943B_ABST
    Figure CN114945943B_ABST
Patent Text Reader

Abstract

Example embodiments relate to estimating depth information based on iris size. A computing system may obtain an image depicting a person and determine a facial mesh for the face based on features of the person's face. In some cases, the facial mesh includes a combination of facial landmarks and eye landmarks. Thus, the computing system may estimate an iris pixel size of the eye based on the eye landmarks of the facial mesh, and estimate a distance of the eye of the face relative to the camera based on the iris pixel size, an average iris size, and an intrinsic matrix of the camera. The computing system may further modify the image based on the estimated distance.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 62 / 960,514, filed on January 13, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] Example embodiments presented herein relate to depth estimation techniques, and more particularly, to methods and systems for estimating depth based on iris size. Background Art

[0004] Many modern computing devices, such as mobile phones, personal computers, and tablet computers, include image capture devices (e.g., still and / or video cameras). The image capture devices can capture images that can depict various scenes (including scenes involving people, animals, landscapes, and / or objects).

[0005] When capturing an image, an image capture device typically generates a two-dimensional representation of a three-dimensional scene (3D). In order to obtain three-dimensional (3D) information about a scene, multiple components are typically used. For example, a stereo camera setup is a common technique for generating 3D information for a scene. A stereo camera involves the use of two or more image capture components that simultaneously capture multiple images to create or simulate a 3D stereo image. Although a stereo camera can generate depth information about a scene, the use of multiple image capture components may increase the cost and complexity associated with obtaining depth information. Summary of the invention

[0006] The example embodiments presented herein relate to depth estimation techniques that involve the use of a single image capture component. Specifically, a smartphone or another type of processing device (e.g., a computing system) can identify the presence of a person's face in an image captured by an image capture component, and then generate a facial mesh representing the outline and features of the person's face. Based on the eye landmarks of the facial mesh indicating the outline features of one or both eyes of the face, the device can estimate one or more eye pixel sizes of at least one eye. For example, the device can estimate the pixel size of the iris of the eye represented in the image. Using one or more estimated eye pixel sizes, the intrinsic matrix of the image capture component that captures the image, and the average eye size corresponding to the estimated eye pixel size, the depth representing the distance between the image capture component and the person's face can be estimated. The depth estimate can then be used to enhance the original image using various techniques, such as generating a new version of the original image focused on the person, while blurring other parts of the image in a manner similar to the bokeh effect.

[0007] Thus, in a first example embodiment, a method is provided. The method includes obtaining, at a computing system, an image depicting a person from a camera, and determining a face mesh for the person's face based on one or more features of the face. The face mesh includes a combination of face landmarks and eye landmarks. The method also includes estimating an eye pixel size of at least one eye of the face based on the eye landmarks of the face mesh, and estimating, by the computing system, a distance of the at least one eye relative to the camera based on the eye pixel size and an intrinsic matrix of the camera. The method also includes modifying the image based on the distance of the at least one eye relative to the camera.

[0008] In a second example embodiment, a system is provided. The system includes a camera having an intrinsic matrix and a computing system configured to perform operations. The operations include obtaining an image depicting a person from the camera and determining a face mesh for the face based on one or more features of the face of the person. The face mesh includes a combination of face landmarks and eye landmarks. The operations also include estimating an eye pixel size of at least one eye of the face based on the eye landmarks of the face mesh, and estimating a distance of the at least one eye relative to the camera based on the eye pixel size and the intrinsic matrix of the camera. The operations also include modifying the image based on the distance of the at least one eye relative to the camera.

[0009] In a third exemplary embodiment, a non-transitory computer readable medium configured to store instructions is provided. The program instructions may be stored in a data memory and, when executed by a computing system, may cause the computing system to perform operations according to the first and second exemplary embodiments.

[0010] In the fourth example embodiment, the system may include various means for performing each operation of the above-described example embodiments.

[0011] These and other embodiments, aspects, advantages and alternatives will become apparent to those of ordinary skill in the art by reading the following detailed description and referring appropriately to the accompanying drawings. In addition, it should be understood that this overview and other descriptions and drawings provided herein are intended to illustrate the embodiments only by way of example, and as such, many variations are possible. For example, structural elements and process steps may be rearranged, combined, distributed, eliminated, or otherwise changed while remaining within the scope of the claimed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1A Depicted are front and side views of a digital camera device according to an example embodiment.

[0013] Figure 1B Depicted is a rear view of a digital camera device according to an example embodiment.

[0014] Figure 2Depicted is a block diagram of a computing system with image capture capabilities in accordance with an example embodiment.

[0015] Figure 3 is a flow chart according to an example embodiment.

[0016] Figure 4 An eye portion of a facial mesh is shown according to an example embodiment.

[0017] Figure 5 Depicted is a simplified representation of an image capture assembly capturing an image of a person in accordance with an example embodiment.

[0018] Figure 6 Depicted is a simplified representation of depth estimation between a camera and an eye in accordance with an example embodiment.

[0019] Figure 7 Depicted is another simplified representation of depth estimation between a camera and an eye, according to an example embodiment.

[0020] Figure 8 Modifying an image captured using a single image capture component according to an example embodiment is depicted. DETAILED DESCRIPTION

[0021] Example methods, devices, and systems are described herein. It should be understood that the words "example" and "exemplary" as used herein mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example" or "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments or features. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.

[0022] Therefore, the example embodiments described herein are not meant to be limiting.As generally described herein and illustrated in the accompanying drawings, various aspects of the present disclosure may be arranged, substituted, combined, separated and designed in a variety of different configurations, all of which are contemplated herein.

[0023] In addition, unless the context suggests otherwise, the features shown in each of the figures in the drawings may be used in combination with each other. Therefore, the drawings should generally be viewed as constituent aspects of one or more overall embodiments, and it should be understood that not all of the features shown are necessary for each embodiment.

[0024] Depending on the context, a "camera" may refer to a separate image capture component, or a device that includes one or more image capture components. Generally, the image capture components may include an aperture, a lens, a recording surface, and a shutter, as described below. In addition, in some embodiments, the image processing steps described herein may be performed by a camera device, while in other embodiments, the image processing steps may be performed by a computing device that communicates with (and possibly controls) one or more camera devices.

[0025] 1. Sample Image Capture Device

[0026] As cameras have become more popular, they are available as standalone hardware devices or integrated into other types of devices. For example, still and video cameras are now commonly included in wireless computing devices (e.g., smartphones and tablet computers), laptop computers, wearable computing devices, video game interfaces, home automation devices, and automobiles and other types of transportation.

[0027] The image capture assembly of a camera may include one or more apertures through which light enters, one or more recording surfaces for capturing images represented by the light, and one or more lenses located in front of each aperture to focus at least a portion of the image onto the recording surface(s). The apertures may be fixed size or adjustable.

[0028] In an analog camera, the recording surface may be a photographic film. In a digital camera, the recording surface may include an electronic image sensor (e.g., a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) sensor) for transmitting and / or storing the captured image in a data storage unit (e.g., a memory). The image sensor may include an array of photosites configured to capture incident light passing through an aperture. When exposure occurs to capture an image, each photosite may collect photons from the incident light and store the photons as an electrical signal. Once the exposure is over, the camera may shut down each of the photosites and continue to measure the electrical signal of each photosite.

[0029] The signals from the array of photosites of the image sensor can then be quantized into digital values, the accuracy of which can be determined by the bit depth. The bit depth can be used to quantify how many unique colors are available in the image's palette, specifying the number of "bits" or 0s and 1s used to represent each color. This does not mean that the image must use all of these colors, but rather that the image can specify the colors with that level of accuracy. For example, for a grayscale image, the bit depth can quantify how many unique shades are available. Therefore, an image with a higher bit depth can encode more shades or colors because there are more combinations of 0s and 1s available.

[0030] To capture a scene in a color image, a color filter array (CFA) located near the image sensor can allow only one color of light to enter each photosite. For example, a digital camera may include a CFA (e.g., a Bayer array) that allows the photosites of the image sensor to capture only one of the three primary colors (red, green, and blue (RGB)). Other potential CFAs may use other color systems, such as a cyan, magenta, yellow, and black (CMYK) array. As a result, the photosites can measure the color of the scene for subsequent display in a color image.

[0031] In some examples, the camera can utilize a Bayer array consisting of alternating rows of red-green and green-blue filters. Within the Bayer array, each primary color does not receive an equal portion of the total area of ​​the array of photosites of the image sensor because the human eye is more sensitive to green light than to both red and blue light. In particular, the redundancy of green pixels can produce an image that appears less noisy and more detailed. Therefore, when configuring a color image of a scene, the camera can approximate the other two primary colors so as to have full color at each pixel. For example, the camera can perform a Bayer demosaicing or interpolation process to convert the primary color array into an image containing full color information for each pixel. The Bayer demosaicing or interpolation can depend on the image format, size, and compression technology used by the camera.

[0032] One or more shutters may be coupled to or near the lens or recording surface. Each shutter may be in a closed position (in which it blocks light from reaching the recording surface), or in an open position (in which light is allowed to reach the recording surface). The position of each shutter may be controlled by a shutter button. For example, the shutter may be in a closed position by default. When the shutter button is triggered (e.g., pressed), the shutter may change from a closed position to an open position for a period of time (referred to as a shutter cycle). During the shutter cycle, an image may be captured on the recording surface. At the end of the shutter cycle, the shutter may be changed back to a closed position.

[0033] Alternatively, the shutter process can be electronic. For example, before the electronic shutter of a CCD image sensor is "opened", the sensor can be reset to remove any residual signal in its photosites. While the electronic shutter remains open, the photosites can accumulate charge. When or after the shutter closes, this charge can be transferred to long-term data storage. Combinations of mechanical and electronic shutters are also possible.

[0034] Regardless of the type, a shutter can be activated and / or controlled by something other than a shutter button. For example, a shutter can be activated by a soft key, a timer, or some other trigger. As used herein, the term "image capture" can refer to any mechanical and / or electronic shutter process that can result in one or more images being recorded, regardless of how the shutter process is triggered or controlled.

[0035] The exposure of the captured image can be determined by a combination of the size of the aperture, the brightness of the light entering the aperture, and the length of the shutter period (also referred to as shutter length or exposure length). Additionally, digital and / or analog gain can be applied to the image, thereby affecting the exposure. In some embodiments, the terms "exposure length," "exposure time," or "exposure time interval" can refer to the shutter length multiplied by the gain of a particular aperture size. Therefore, these terms can be used somewhat interchangeably and should be interpreted as being shutter length, exposure time, and / or any other measure of the amount of signal response generated by the light reaching the recording surface.

[0036] A still camera may capture one or more images each time image capture is triggered. A video camera may continuously capture images at a particular rate (e.g., 24 images or frames per second) as long as image capture remains triggered (e.g., when a shutter button is pressed). Some digital still cameras may open a shutter when a camera device or application is activated, and the shutter may remain in this position until the camera device or application is deactivated. While the shutter is open, the camera device or application may capture and display a representation of the scene on the viewfinder. When image capture is triggered, one or more different digital images of the current scene may be captured.

[0037] A camera may include software that controls one or more camera functions and / or settings, such as aperture size, exposure time, gain, etc. Additionally, some cameras may include software that digitally processes images during or after capturing those images.

[0038] As mentioned above, a digital camera can be a stand-alone device or integrated with other devices. As an example, Figure 1A The form factor of the digital camera device 100 as seen from the front view 101A and the side view 101B is shown. Figure 1B Also shown are the form factors of the digital camera device 100 as seen from a rear view 101C and another rear view 101D. The digital camera device 100 may be a mobile phone, a tablet computer, or a wearable computing device. Other embodiments are also possible.

[0039] like Figure 1A and 1B As shown, digital camera device 100 may include various elements, such as body 102, front camera 104, multi-element display 106, shutter button 108, and additional buttons 110. Front camera 104 may be located on the side of body 102 that generally faces the user during operation, or on the same side as multi-element display 106.

[0040] like Figure 1BAs shown, the digital camera device 100 also includes a rear camera 112. In particular, the rear camera 112 is shown as being located on a side of the body 102 opposite to the front camera 104. In addition, Figure 1B The rear views 101C and 101D shown in the figure represent two alternative arrangements of the rear camera 112. However, other arrangements are also possible. In addition, calling a camera front or rear is arbitrary, and the digital camera device 100 may include one or more cameras located on each side of the body 102.

[0041] The multi-element display 106 may represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, or any other type of display known in the art. In some embodiments, the multi-element display 106 may display a digital representation of a current image captured by the front camera 104 and / or the rear camera 112, or an image that may be captured or recently captured by any one or more of these cameras. Thus, the multi-element display 106 may act as a viewfinder for the camera. The multi-element display 106 may also support a touch screen and / or presence-sensitive functionality that enables adjustment of settings and / or configuration of any aspect of the digital camera device 100.

[0042] The front camera 104 may include an image sensor and associated optical elements, such as a lens. Thus, the front camera 104 may provide a zoom capability or may have a fixed focal length. In other embodiments, an interchangeable lens may be used with the front camera 104. The front camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. The front camera 104 may also be configured to capture still images, video images, or both. The rear camera 112 may be an image capture component of a similar type, and may include an aperture, a lens, a recording surface, and a shutter. Specifically, the rear camera 112 may operate similarly to the front camera 104.

[0043] Either or both of the front camera 104 and the rear camera 112 may include or be associated with an illumination assembly that provides a light field to illuminate the target object. For example, the illumination assembly may provide a flash or constant illumination of the target object. The illumination assembly may also be configured to provide a light field that includes one or more of structured light, polarized light, and light with specific spectral content. Other types of light fields that are known and used to recover a 3D model from an object are possible in the context of the embodiments herein.

[0044] Either or both of the front camera 104 and / or the rear camera 112 may include or be associated with an ambient light sensor that may continuously or from time to time determine the ambient brightness of the scene that the camera may capture. In some devices, the ambient light sensor may be used to adjust the display brightness of a screen (e.g., a viewfinder) associated with the camera. When the determined ambient brightness is high, the brightness level of the screen may be increased to make the screen easier to view. When the determined ambient brightness is low, the brightness level of the screen may be reduced, again to make the screen easier to view and potentially save power. The ambient light sensor may also be used to determine the exposure time for image capture.

[0045] The digital camera device 100 can be configured to capture images of a target object using the multi-element display 106 and the front camera 104 or the rear camera 112. The captured images can be multiple static images or a video stream. Image capture can be triggered by activating the shutter button 108, pressing a soft key on the multi-element display 106, or by some other mechanism. Depending on the implementation, images can be automatically captured at specific time intervals, for example, when the shutter button 108 is pressed, under appropriate lighting conditions for the target object, when the digital camera device 100 is moved a predetermined distance, or according to a predetermined capture schedule.

[0046] In some examples, one or both of the front camera 104 and the rear camera 112 are calibrated monocular cameras. A monocular camera can be an image capture component configured to capture a 2D image. For example, a monocular camera can use a modified refracting telescope to magnify the image of a distant object by passing light through a series of lenses and prisms. Therefore, a monocular camera and / or other types of cameras can have an intrinsic matrix that can be used for the depth estimation techniques presented herein. The intrinsic matrix of the camera is used to transform 3D camera coordinates into 2D homogeneous image coordinates.

[0047] As described above, the functionality of the digital camera device 100 (or another type of digital camera) can be integrated into a computing device (such as a wireless computing device, a mobile phone, a tablet computer, a wearable computing device, a robotic device, a laptop computer, a car camera, etc.). For example purposes, Figure 2 is a simplified block diagram illustrating some components of an example computing system 200 that may include a camera component 224 .

[0048] By way of example and not limitation, computing system 200 may be a cellular mobile telephone (e.g., a smartphone), a still camera, a video camera, a fax machine, a computer (such as a desktop, notebook, tablet, or handheld computer), a personal digital assistant (PDA), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a robotic device, a vehicle, or some other type of device equipped with at least some image capture and / or image processing capabilities. It should be understood that computing system 200 may represent a physical camera device, such as a digital camera, a specific physical hardware platform on which a camera application operates in software, or other combinations of hardware and software configured to perform camera functions.

[0049] like Figure 2 As shown, computing system 200 may include a communication interface 202 , a user interface 204 , a processor 206 , data storage 208 , and a camera assembly 224 , all of which may be communicatively linked together via a system bus, network, or other connection mechanism 210 .

[0050] The communication interface 202 may allow the computing system 200 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation. Thus, the communication interface 202 may facilitate circuit-switched and / or packet-switched communications, such as plain old telephone service (POTS) communications and / or Internet Protocol (IP) or other packetized communications. For example, the communication interface 202 may include a chipset and an antenna arranged to communicate wirelessly with a radio access network or access point. Furthermore, the communication interface 202 may take the form of or include a wired interface, such as an Ethernet, Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI) port. The communication interface 202 may also take the form of a wireless interface, such as Wi-Fi, The communication interface 202 may be in the form of or include a wireless interface such as a global positioning system (GPS) or a wide area wireless interface (e.g., WiMAX or 3GPP Long Term Evolution (LTE)). However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used on the communication interface 202. In addition, the communication interface 202 may include multiple physical communication interfaces (e.g., Wifi interfaces, interface and wide area wireless interface).

[0051] User interface 204 can be used to allow computing system 200 to interact with the human or non-human user, such as receiving input from the user and providing output to the user. Therefore, user interface 204 can include input components, such as keypad, keyboard, touch-sensitive or display-sensitive panel, computer mouse, trackball, joystick, microphone etc. User interface 204 can also include one or more output components, such as one or more display screens, which for example can be combined with display-sensitive panel. Display screen can be based on CRT, LCD and / or LED technology or other technologies known now or developed later. User interface 204 can also be configured to generate (multiple) auditory outputs via loudspeaker, loudspeaker insert aperture, audio output port, audio output device, earphone and / or other similar devices.

[0052] In some embodiments, user interface 204 may include a display that serves as a viewfinder for still camera and / or video camera functions supported by computing system 200. Additionally, user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate configuration and focusing of camera functions and capture of images (e.g., capturing a picture). Some or all of these buttons, switches, knobs, and / or dials may be implemented via a display sensitive panel.

[0053] Processor 206 may include one or more general purpose processors (e.g., microprocessors) and / or one or more special purpose processors (e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, or application specific integrated circuits (ASICs). In some cases, the special purpose processors may be capable of image processing, image alignment, and merging images, etc. Data storage 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated with processor 206 in whole or in part. Data storage 208 may include removable and / or non-removable components.

[0054] The processor 206 may be capable of executing program instructions 218 (e.g., compiled or non-compiled program logic and / or machine code) stored in the data storage device 208 to perform various functions described herein. Therefore, the data storage device 208 may include a non-transitory computer-readable medium having program instructions stored thereon, which, when executed by the computing system 200, causes the computing system 200 to perform any method, process, or operation disclosed in this specification and / or the accompanying drawings. The execution of the program instructions 218 by the processor 206 may cause the processor 206 to use the data 212.

[0055] For example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, (multiple) device drivers, and / or other modules) and one or more applications 220 (e.g., camera functions, address book, email, web browsing, social networking, image applications, and / or game applications) installed on computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. Operating system data 216 may be primarily accessed by operating system 222, and application data 214 may be primarily accessed by one or more of application programs 220. Application data 214 may be arranged in a file system that is visible or hidden to a user of computing system 200.

[0056] Applications 220 may communicate with operating system 222 via one or more application programming interfaces (APIs). These APIs may facilitate, for example, application 220 reading and / or writing application data 214, sending or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and the like.

[0057] In some dialects, the application 220 may be referred to simply as an "app". Additionally, the application 220 may be downloaded to the computing system 200 through one or more online application stores or application markets. However, the application may also be installed on the computing system 200 in other ways, such as via a web browser or through a physical interface (e.g., a USB port) on the computing system 200.

[0058] Camera assembly 224 may include, but is not limited to, an aperture, a shutter, a recording surface (e.g., a photographic film and / or an image sensor), a lens, and / or a shutter button. Thus, camera assembly 224 may be controlled, at least in part, by software executed by processor 206. In some examples, camera assembly 224 may include one or more image capture components, such as a monocular camera. Although camera assembly 224 is shown as part of computing system 200, in other embodiments they may be physically separated. For example, camera assembly 224 may capture an image and provide the image to computing system 200 via a wired or wireless connection for subsequent processing.

[0059] 2. Example Operation

[0060] As described above, an image captured by a camera may include intensity values, where bright pixels have higher intensity values ​​and dark pixels have lower intensity values. In some cases, an image may also represent the depth of objects within a scene, where the depth indicates the distance of one or more objects relative to the camera device that captured the image. For example, depth information can be used to guide an observer to a specific aspect of an image, such as a person, while also blurring the background to enhance the overall image. A conventional type of camera for generating depth information is a stereo camera, which may involve the use of two or more image capture components to simultaneously capture multiple images that can be used to generate depth information. Although a stereo camera can produce depth information about a scene, the use of multiple image capture components increases the cost and complexity associated with generating the depth information.

[0061] The example embodiments presented herein relate to depth estimation techniques that can be performed using a single camera. Specifically, a smartphone or another type of processing device (e.g., a computing system) can identify the presence of a person's face in an image captured by a single camera (e.g., a monocular camera), and then generate a facial mesh representing the characteristic outline of the person's face. The facial mesh may include facial landmarks arranged according to the facial features of the person and eye landmarks arranged according to the position and size of one or both eyes of the person's face. For example, the eye landmarks may include a set of iris landmarks positioned relative to the iris of the eye to provide information about the iris. For example, the iris landmarks may be located around the iris of the eye and at other locations relative to the iris (e.g., an iris landmark marking the center of the iris). As a result, the facial mesh can convey information about the face of the person depicted in the image. In some embodiments, the facial mesh may also provide information about the position of the person's face relative to other objects within the scene captured by the camera.

[0062] Thus, the device can use the eye landmarks of the facial mesh to estimate one or more eye pixel sizes of at least one eye. For example, the estimated eye size can represent the pixel size of the iris of the eye (e.g., the vertical or horizontal diameter of the iris quantified in pixels as represented in the image). Another estimated eye size can be a pixel eye size, which represents the pixel size of the entire eye depicted in the image captured by the camera. Other examples of eye size estimation are also possible. For example, the eye size can represent the diameter of the cornea of ​​the eye.

[0063] In addition, the device may use an average eye size corresponding to the determined eye pixel size for depth estimation. The average eye size may represent an average eye size and may be based on ophthalmological information or other measurements from multiple eyes. Ophthalmological information may provide dimensions applicable to various aspects of the eyes of various people. For example, ophthalmological information may specify an average mean horizontal diameter, average vertical diameter, or other measurement of an adult's iris. Ophthalmological information may represent standardized measurements obtained from many people.

[0064] In one embodiment, when the eye pixel size corresponds to the iris size measured in pixels, the average eye size may represent the average iris size measured in millimeters or another unit. For example, when the device estimates the number of pixels representing the horizontal diameter of the iris depicted in the image, the device may further use the average eye size representing the average horizontal diameter of the iris in millimeters or another unit. In another example, when the eye pixel size conveys the number of pixels representing the vertical diameter of the iris depicted in the image, the device may use the average eye size representing the average vertical diameter of the iris.

[0065] The device may obtain the average eye size from another computing system or memory. For example, the device may access eye data (e.g., average eye size) through a wireless connection to a database. In some embodiments, the device may have locally stored average eye values, which may reduce the time required to estimate the depth of a person's face in an image.

[0066] The device can then use one or more eye pixel sizes and one or more corresponding average eye sizes and the intrinsic matrix of the camera that captured the image to estimate the depth of the person's face relative to the camera. The estimated depth of the person can then be used to enhance the original image. For example, the device can generate a new image based on the original image, which enhances the presence of the person in the image via simulation of a bokeh effect and / or using other image enhancement techniques. By estimating the depth of a person within an image based on the estimated eye pixel size, the average eye size, and the intrinsic matrix of the camera, the need for multiple cameras is eliminated, so the overall cost and complexity associated with generating depth information for the scene can be reduced.

[0067] To further illustrate, Figure 3 A flow chart for depth estimation using images obtained via a single image capture component is shown. Figure 3 The illustrated embodiment may be performed by a computing system such as the digital camera device 100 shown in FIG1 . However, the embodiment may also be performed by other types of devices or device subsystems, such as a computing system remote from the camera. In addition, the embodiment may be combined with any aspect or feature disclosed in the specification or drawings.

[0068] At block 302, method 300 may include obtaining an image depicting a person from a camera. The computing system (or a subsystem of the computing system) may obtain the image from the camera via a wired or wireless connection. For example, the computing system may be a smartphone configured with a camera for capturing the image. Alternatively, the computing system may be located remotely from the camera and obtain the image via a wireless or wired connection.

[0069] A camera may represent any type of image capturing component. For example, a camera may include one or more apertures, lenses, and recording surfaces. Thus, a camera may have an intrinsic matrix that enables the transformation of 3D camera coordinates into 2D homogeneous image coordinates.

[0070] The image may depict one or more people in the scene. For example, a smartphone or another type of device may use a camera to capture a portrait of a person. After obtaining the image from the camera, the computing system may perform an initial inspection of the image to detect the face of the person before performing other functions of method 300. For example, the computing system may use image processing techniques, machine learning, a trained neural network, or another process (e.g., machine learning) to recognize the face of the person. In addition, in some examples, before performing one or more functions of method 300, the computing system may require that the image be captured when the camera is in a specific camera mode (e.g., portrait mode).

[0071] In further embodiments, the image may depict an animal or another type of face (e.g., a painting of a person). Thus, the device may detect the face of the animal or artwork before continuing to perform other functions of method 300.

[0072] At block 304, method 300 may include determining a facial mesh for a person's face based on one or more features of the person's face. The facial mesh may include a combination of facial landmarks and eye landmarks arranged to represent information about the face. In some embodiments, the facial mesh may be generated based on an image, and the facial mesh provides information about at least one person depicted within the image. For example, the facial mesh may indicate which pixels represent portions of a person's face.

[0073] The computing system may use one or more image processing techniques to perform a cursory inspection of the image to identify the presence of one or more persons and subsequently generate a facial mesh of at least one person. For example, the computing system may be configured to generate a facial mesh for a person approximately located in the center of the image when analyzing an image with multiple persons. In examples, various image processing techniques may be used to identify information about the image, such as detecting the presence of a person and / or a person's face. Example image processing techniques include, but are not limited to, feature extraction, pattern recognition, machine learning, and neural networks.

[0074] The facial mesh generated by the computing system for a face may include a set of facial landmarks and a set of eye landmarks. Landmarks may be circular dots (or dots having another shape) arranged in a particular manner to mark information about a person's face. In particular, the computing device may mark the facial landmarks as representing the outline of the face, and mark the eye landmarks as representing the outline of one or both eyes of the face. For example, facial landmarks may be dots arranged to represent facial features of a person's face, such as cheeks, lips, eyebrows, ears, etc. As a result, facial landmarks may convey the layout and outline of a person's face.

[0075] Eye landmarks are similar to facial landmarks, but are specific to one or both eyes of a person's face. In particular, eye landmarks may specify the location and outline size of one or both eyes. In some embodiments, eye landmarks may include a set of iris landmarks that clearly define the location, positioning, and size of the iris. The combination of facial landmarks and eye landmarks together may provide a representation of the arrangement and features of a person's face captured within an image.

[0076] The total number of facial landmarks and eye landmarks used to generate a facial mesh may vary in examples. For example, a facial mesh may include approximately 500 facial landmarks and 30 eye landmarks per eye, with 5 iris landmarks defining the center and outline of the iris. These numbers may vary for other facial meshes and may depend on the number of pixels used to represent the face of the person in the image. In addition, in examples, the size, color, and style of the dots used for facial landmarks and eye landmarks (as well as iris landmarks) may be different. In some cases, facial landmarks and eye / iris landmarks may be circular dots of uniform size. Alternatively, dots of different shapes and sizes may be used. In addition, facial landmarks and eye landmarks of different colors may be used to generate a facial mesh. For example, the eye landmarks may be a first color (e.g., green), and the iris landmarks may be a second color (e.g., red). In other embodiments, the facial mesh may include lines or other structures for conveying information about a person's face.

[0077] In some examples, a neural network is trained to determine a facial mesh for a person's face. For example, a neural network may be trained using pairs of images depicting faces with and without facial meshes. As a result, a computing system may use the neural network to determine a facial mesh and further measure facial features based on the facial mesh. Alternatively, one or more image processing techniques may be used to generate a facial mesh for a face captured in one or more images. For example, a computing system may use a combination of a neural network and edge detection to generate one or more facial meshes for a face represented within an image.

[0078] At block 306, method 300 may include estimating eye pixel dimensions of at least one eye of the face based on the eye landmarks of the face mesh. After generating a face mesh for a face depicted within an image, the computing system may use the eye landmarks of the face mesh to estimate one or more eye dimensions of one or both eyes of the person.

[0079] In some embodiments, estimating the eye size of the eye may include estimating the pixel size of the iris of the eye based on an iris landmark located around the iris of the eye. For example, the pixel size of the iris may correspond to the number of pixels representing the horizontal diameter of the iris depicted in the image. Alternatively, the pixel size of the iris may correspond to the number of pixels representing the vertical diameter of the iris depicted in the image. Thus, when estimating the eye pixel size (e.g., the iris pixel size), the computing system may use the eye landmark and / or the iris landmark. In some examples, the computing system may perform multiple eye pixel size estimates and use an average of the estimates as the output eye size for use at block 308.

[0080] The computing system may also make a comparison between the pixel size of the iris of the eye and an average pixel iris size. The average pixel iris size may be based on multiple pixel measurements of the iris and / or another value (e.g., a threshold range). Based on the comparison, the computing system may determine whether the pixel size of the iris satisfies a threshold difference before estimating the distance of the eye relative to the camera. When the estimated eye size fails to satisfy the threshold difference (i.e., is significantly different from the target eye size), the computing system may repeat one or more processes to derive a new eye size estimate.

[0081] In addition, the computing system may also use an average eye size corresponding to the eye pixel size. Specifically, the average eye size may indicate the size of a person's eyes based on the eye data. When the iris pixel size is used, the average eye size may correspond to the matching average iris size. For example, the iris pixel size and the average iris size may represent the same parameter of the iris, such as the horizontal diameter of the iris.

[0082] At block 308, method 300 may include estimating a distance of at least one eye relative to the camera based on the eye pixel size and an intrinsic matrix of the camera. The computing device may use the intrinsic matrix of the camera and one or more estimated eye pixel sizes (e.g., pixel iris sizes) to calculate an estimated depth of the eye relative to the camera. For example, the camera may be a calibrated monocular camera with an intrinsic matrix, and the computing system may utilize the intrinsic matrix in conjunction with one or more eye sizes during depth estimation. Additionally, the device may also use an average eye size to represent the person's eyes in millimeters or another unit when estimating the distance between the camera and the person's face. For example, the device may use a combination of pixel iris estimates, an intrinsic matrix, and an average iris size to estimate the distance from the camera to the person's face.

[0083] The intrinsic matrix is ​​used to transform 3D camera coordinates into 2D image coordinates. An example intrinsic matrix can be parameterized as follows:

[0084]

[0085] Each of the intrinsic parameters shown above describes a geometric property of the camera. The focal length, also called pixel focal length, is given by f x 、f y denoted by , and corresponds to the distance between the camera aperture and the image plane. Focal length is usually measured in pixels, and when the camera simulates a real pinhole camera that produces square pixels, f x and f y have the same value. In practice, f x and f y may differ for various reasons, such as imperfections in the digital camera sensor, the image being non-uniformly scaled in post-processing, unintentional distortion caused by the camera lens, or errors in camera calibration. x and f y Otherwise, the resulting image may consist of non-square pixels.

[0086] Additionally, the principal point offset is represented by x0 and y0. The principal axis of the camera is the line perpendicular to the image plane that passes through the camera aperture. The intersection of the principal axis with the image plane is called the principal point. Therefore, the principal point offset is the position of the principal point relative to the origin of the image plane. Additionally, axis tilt is represented by s in the matrix above and causes shear distortion in the projected image. From this, the computing system can estimate the depth of a person relative to the camera using the camera's intrinsic matrix and the estimated eye size derived from the facial mesh. In some examples, the facial mesh can be further used during depth estimation.

[0087] At block 310, method 300 may include modifying the image based on a distance of at least one eye relative to the camera. Modifying the image may include adjusting aspects of an original image provided by the camera and / or generating a new enhanced image corresponding to the original image.

[0088] The computing system can use depth estimation to segment the image into background and foreground portions. Thus, when generating an enhanced final image, the computing system can blur one or more pixels of the background portion in the original image. In particular, focused features of a scene present in the foreground (e.g., a person in the center of the image) can retain sharp pixels, while other features that are part of the background of the scene can be blurred to enhance the final image. In some cases, pixels of features in the background of the image can be blurred proportionally based on how far each background feature is from the focus plane (e.g., from the camera). The estimated depth map can be used to estimate the distance between the background features and the focus plane.

[0089] In one embodiment, the computing system may generate a final image that includes one or more enhancements compared to the original captured image. For example, the final image may utilize a blur effect to help draw the viewer's attention to the primary focal point in the image (e.g., a person or a person's face) in a manner similar to an image captured with a shallow depth of field. In particular, an image with a shallow depth of field may help draw the viewer's attention to the focal point of the image and may also help suppress cluttered backgrounds, thereby enhancing the overall presentation of the image.

[0090] In another embodiment, blurring pixels of background features may include replacing the pixels with a semi-transparent disk of the same color but a different size. By compositing all of these disks in depth order in a manner similar to averaging the disks, the result of the enhanced final image resembles a real optical blur produced using a single-lens reflex (SLR) camera with a large lens. The synthetic defocus applied using the above techniques can produce a disk-shaped bokeh effect in the final image without requiring the extensive equipment that other cameras often use to achieve the effect. Additionally, unlike SLR cameras, because the bokeh effect is applied in a synthetic manner, the bokeh effect can be modified to have other shapes in the final image using the above techniques.

[0091] 3. Example facial mesh determination

[0092] As described above, a computing system (e.g., a smartphone) may generate a facial mesh for one or more faces depicted within an image prior to performing depth estimation techniques. For illustration, Figure 4 Shown is the eye portion of a facial mesh generated for the face of a person captured within an image.

[0093] Eye portion 400 may be part of a face mesh that includes facial landmarks (not shown), eye landmarks 402, and iris landmarks 404, 405. In some embodiments, eye landmarks 402 and iris landmarks 404, 405 may be combined so that they appear generally as eye landmarks. Alternatively, in some examples, the face mesh may include only eye landmarks 402 and / or iris landmarks 404, 405. Facial landmarks may be arranged to mark features of a person's face, such as cheeks, lips, eyebrows, the location of ears, and the overall outline of the face.

[0094] The eye indicia shown in the eye portion 400 of the facial diagram are arranged to mark various aspects of the eye 406 (shown in dashed boxes). Figure 4 In the illustrated embodiment, the eye mark includes an outline eye mark 402 shown as a solid dot and iris marks 404, 405 shown as dotted dots located around the iris 408. The outline eye mark 402 is arranged to outline the entire eye 406. Figure 4 As shown, the eye outline marker 402 can be a set of points (e.g., 17 points) arranged to enable the computing system to recognize the outline of the eye 406.

[0095] In addition to the outline eye marker 402, iris markers 404, 405 are arranged to mark the location of an iris 408 within an eye 406. In particular, four iris markers 404 are located around the iris 408, and one iris marker 405 is located at the approximate center of the iris 408. As shown, the iris markers 404, 405 may divide the iris 408 into quadrants, and may enable a computing system to estimate a horizontal diameter 410 of the iris 408 and / or a vertical diameter 412 of the iris 408. One or both of the estimated eye dimensions may be used to estimate the depth of the eye 406 relative to a camera capturing an image of the person.

[0096] 4. Example Depth Estimation

[0097] Figure 5 502 and 504, as well as other components not shown. During image capture, light representing person 506 and other elements of the scene (not shown) may pass through lens 504, enabling the camera to subsequently create an image of person 506 on recording surface 502. As a result, the camera may display a digital image of person 506 on a viewfinder. Figure 5 In the illustrated embodiment, the image of person 506 appears upside down on recording surface 502 due to the optical properties of lens 504, but image processing techniques can invert the image for display.

[0098] For some camera configurations, lens 504 may be adjustable. For example, lens 504 may be moved left or right, adjusting the lens position and focal length of the camera for image capture. The position of lens 504 may be controlled by moving a motor ( Figure 5 The motor 504 may be adjusted by applying a voltage to the recording surface 502 (not shown). Thus, the motor may move the lens 504 closer to or further away from the recording surface 502, enabling the camera to focus on objects (e.g., person 506) within a certain distance range. The distance between the lens 504 and the recording surface 502 at any point in time may be referred to as the lens position, and may be measured in millimeters or other units. By extension, the distance between the lens 504 and its focal area may be referred to as the focal length, which may similarly be measured in millimeters or other units.

[0099] As described above, depth estimation techniques may involve using information obtained from one or more eyes of a person captured within an image. In particular, a smartphone, server, or another type of computing system may obtain an image and generate a facial mesh for one or more faces depicted in the image. For example, the device may use at least one image processing technique to perform a cursory inspection to detect and identify the presence of a person's face, and then generate a facial mesh with facial landmarks and eye landmarks that mark out the face.

[0100] Facial landmarks can mark the contours of the face, such as the curvature and position of the cheeks, lips, eyebrows, and facial contours. Similar to facial landmarks, eye landmarks can mark the contours of at least one eye on the face. For example, Figure 3 As shown, the eye landmarks may include a first set of eye landmarks that outline the eyes of the face and a second set of eye landmarks that mark the size of the irises of the eyes.

[0101] Using the facial mesh, the device can determine information that can be used for depth estimation. For example, the facial mesh can be used to determine the location of a person's eyes in an image captured by a camera. The location of a person's eyes can be determined relative to other objects in the scene and can include using the camera's intrinsic matrix. The facial mesh can also be used to estimate the pixel size of the eyes depicted in the image. For example, the facial mesh can be used to determine the number of pixels (e.g., 5 pixels, tens of pixels) that depict the horizontal (or vertical) diameter of an iris represented within the image. In particular, depth estimation techniques can include creating a distance formula based on a combination of eye measurements and the camera's intrinsic matrix. To further illustrate, the camera's intrinsic matrix can be expressed as follows:

[0102]

[0103] Similar to the matrix shown in Equation 1 above, this matrix uses f when the image has square pixels x and f yThe focal length is expressed in pixels, and their values ​​are equal, and O x and O y is used to represent the position of the principal point on the image sensor of the camera. In addition, for the purpose of illustration, matrix 2 has the axis tilt values ​​set to zero.

[0104] To further illustrate how the above matrix, eye pixel estimates, and corresponding average eye size can be used to estimate depth estimates for people in a scene, Figure 6 and 7 Simplified representations of depth estimation between a camera and an eye are shown. In particular, each representation 600, 700 provides a diagram showing example distances and estimates that can be used with the camera's intrinsic matrix to estimate the depth of the eye. Although the representations 600, 700 are simplified for illustrative purposes, showing only limited camera components and an eye, the depth estimation techniques described herein can be used to depict images of more complex scenes.

[0105] In simplified representation 600, a camera may use a lens 602 with a pinhole O 612 to capture an image of an eye 604 and generate an image of the eye 604 on a recording surface 606. Using the image, the device may perform one or more estimations of the eye 604 for use in conjunction with the camera's intrinsic matrix for depth estimation. The device may determine a first eye size 604 based on the size of the eye. For example, the device may determine Figure 6 The average iris size is represented in 608 as iris size AB, which may be expressed in millimeters or another unit. The device may determine the average iris size using stored data and / or by obtaining an average from another computing system.

[0106] In addition, the device can also estimate the size of the iris depicted in the image on the recording surface 606, which is Figure 6 610. Thus, the iris pixel size A'B' 610 can be determined based on the number of pixels (or sub-pixels) within the image and is therefore represented as the total number of pixels. The iris pixel size can be different depending on the measurement that the device is estimating (e.g., vertical diameter or horizontal diameter).

[0107] Simplified representation 600 also shows focal length OO' 616 (also called pixel focal length) extending between the center of the lens (pinhole O 612) and a principal point 614 on recording surface 606. Focal length OO' 616 may be determined based on the intrinsic matrix of the camera and may be expressed in millimeters or other units.

[0108] As further shown in the simplified representation 600, the distances from the center of the lens (pinhole O 612) to point A and point B of the eye 604 may be determined. Specifically, a first triangle OAB may be determined based on the distances from the pinhole O 612 to points A, B using an average eye size (e.g., iris size AB 608). Additionally, a second triangle OA'B' may be determined based on the distances from the pinhole O 612 to points A' and B' of the image using an estimated iris pixel size A'B' 610.

[0109] To construct a depth estimation equation that can be used to estimate the depth of eye 604 relative to lens 602, triangles OAB and OA'B' shown in simplified representation 600 can then be used to construct the following equation:

[0110]

[0111]

[0112] Equations 3 and 4 derived based on triangles OAB and OA'B' can be used to calculate the distance to the camera as follows:

[0113]

[0114]

[0115]

[0116] As shown above, equation 7 can be determined based on equations 5 and 6 and can be used to estimate the distance of the eye 604 (and typically the face of a person) relative to the camera. In an example, the distance can be estimated in millimeters or other units. Therefore, the device can use equation 7 to estimate the depth of the face of a person relative to the camera that captured the image using the iris size estimate and the intrinsic matrix of the camera. In equation 7, pupil to focal center represents the pixel distance in image space from the center of the eye (pupil) to the focal origin from the intrinsic matrix.

[0117] Similar to Figure 6Simplified representation 600 is shown, simplified representation 700 represents another view of depth estimation between a camera and an eye. Specifically, representation 700 shows a camera using a lens 702 to capture an image 706 depicting an eye 704. The device can obtain the image 706 and then determine an average iris size AB 708 in millimeters or other units based on stored eye data. In addition, the device can estimate an iris pixel size A'B' 710, which represents the size of the iris of the eye 704 depicted within the image 706 in terms of the number of pixels. Using one or both of the average iris size AB 708, the estimated iris pixel size A'B' 710, and information from the camera's intrinsic matrix, the device can estimate the distance of the eye 704 relative to the camera. In particular, the distance between the principal point O' and the center O of the lens 702 can correspond to the focal length OO' 712 of the camera expressed in pixels.

[0118] 5. Example Image Modification

[0119] Figure 8 A simplified image modification according to one or more example embodiments is shown. The simplified image modification shows an input image 800 and an output image 804. For purposes of illustration, both images depict a person 802 in a simplified portrait arrangement. Other examples may involve more complex scenes, such as images depicting multiple people with various backgrounds.

[0120] The input image 800 is a Figure 1A An image of a person 802 captured by the front camera 104 of the digital camera device 100 is shown. In an embodiment, the input image 800 may represent an image captured by a monocular camera having an intrinsic matrix. Thus, the computing system may receive the input image 800 from the camera and may also access the intrinsic matrix of the camera.

[0121] In response to receiving input image 800, the computing system may perform depth estimation techniques to determine the depth of person 802 relative to the camera that captured input image 800. Thus, the depth estimation may be used to modify input image 800.

[0122] In some examples, modifying input image 800 may include directly enhancing input image 800 or generating an enhanced image (e.g., output image 804) based on input image 800. For example, the device may generate a new version of the initial image that includes a focus on person 802 while also blurring out other portions of the scene, such as the background of the scene (represented by black in output image 804). The camera device may identify the portion of the image to focus on depending on the overall layout of the scene captured within the image or based on additional information (e.g., user input specifying a focus point during image capture).

[0123] In another example, the camera may use an image segmentation process to segment the image into multiple segments. Each segment may include a group of pixels of the image. The camera may identify segments that share respective features (e.g., representing person 802) and further identify boundaries of features in the scene based on the segments that share respective features. One of the depth estimation techniques described above may be used to generate a new version of the image that focuses on the person in the scene and blurs one or more other portions of the scene.

[0124] In another implementation, the camera may receive input that specifies a particular person in the scene that the camera focused on when capturing the image. For example, the camera may receive input via a touch screen that displays a view of the scene from the camera's viewpoint. Thus, the camera may estimate the depth of the person in response to receiving the input.

[0125] 6. Conclusion

[0126] The present disclosure is not limited to the specific embodiments described in this application, which are intended to illustrate various aspects. It will be apparent to those skilled in the art that many modifications and variations may be made without departing from the scope thereof. In addition to those listed herein, functionally equivalent methods and devices within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.

[0127] The above detailed description describes various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. The example embodiments described herein and in the accompanying drawings are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and shown in the accompanying drawings, may be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein.

[0128] With respect to any or all message flow diagrams, scenarios and flow charts in the accompanying drawings, and as discussed herein, according to example embodiments, each step, frame and / or communication can represent the processing of information and / or the transmission of information. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, the functions described as steps, frames, transmissions, communications, requests, responses and / or messages may not be performed in the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functions involved. In addition, more or less frames and / or functions can be used together with any ladder diagram, scenario and flow chart discussed herein, and these ladder diagrams, scenarios and flow charts can be combined with each other in part or in their entirety.

[0129] The steps or boxes representing information processing may correspond to circuits that may be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, the steps or boxes representing information processing may correspond to modules, fragments, or portions of program code (including associated data). The program code may include one or more instructions that may be executed by a processor to implement specific logical functions or actions in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium (such as a storage device including a disk, hard drive, or other storage medium).

[0130] Computer readable media may also include non-transitory computer readable media, such as computer readable media that store data in the short term, such as register memory, processor cache, and random access memory (RAM). Computer readable media may also include non-transitory computer readable media that store program code and / or data in the long term. Therefore, computer readable media may include secondary or permanent long-term storage, such as read-only memory (ROM), optical disk or disk, compact disk read-only memory (CD-ROM). Computer readable media may also be any other volatile or non-volatile storage system. For example, computer readable media may be considered to be computer readable storage media, or tangible storage devices.

[0131] In addition, the steps or boxes representing one or more information transfers may correspond to information transfers between software and / or hardware modules in the same physical device. However, other information transfers may be performed between software modules and / or hardware modules in different physical devices.

[0132] The particular arrangements shown in the accompanying drawings should not be considered limiting. It should be understood that other embodiments may include more or less of each element shown in a given figure. In addition, some of the elements shown may be combined or omitted. In addition, example embodiments may include elements not shown in the figures.

[0133] In addition, any listing of elements, frames or steps in this specification or claims is for clarity purposes. Therefore, such listing should not be interpreted as requiring or implying that these elements, frames or steps follow a specific arrangement or are performed in a specific order.

[0134] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and not limitation, with the true scope being indicated by the appended claims.

Claims

1. An image processing method, comprising: obtaining, at a computing system, from a camera, an image depicting a person; determining a facial mesh for the face based on one or more features of the face of the person, wherein the facial mesh includes a combination of facial landmarks and eye landmarks; estimating an eye pixel size of at least one eye of the face based on eye landmarks of the face mesh; estimating, by the computing system, a distance of the at least one eye relative to the camera based on the eye pixel size, an intrinsic matrix of the camera, and an average eye size, wherein the average eye size indicates a size of a person's eye based on the eye data; and The image is modified based on a distance of the at least one eye relative to the camera.

2. The method according to claim 1, wherein: The camera is a monocular camera.

3. The method according to claim 1, further comprising: In response to obtaining an image depicting the person, a face of the person in the image is recognized using at least one image processing technique.

4. The method according to claim 3, wherein: The at least one image processing technique includes using a trained neural network to detect one or more features of the face.

5. The method according to claim 1, wherein: Determining a facial mesh of the face includes: marking the facial landmarks as an outline representing the face based on the image; and The eye landmark is marked as an outline representing the at least one eye of the face based on the image.

6. The method according to claim 5, wherein: Marking the eye landmark as an outline representing at least one eye of the face comprises: The eye landmark is positioned around the iris of the at least one eye.

7. The method according to claim 6, wherein: Estimating an eye pixel size of at least one eye of the face comprises: A pixel size of the iris of the at least one eye is estimated based on eye landmarks located around the iris of the at least one eye.

8. The method according to claim 7, wherein: Estimating the pixel size of the iris of the at least one eye comprises: A number of pixels indicative of a horizontal diameter of an iris of the at least one eye represented in the image is estimated.

9. The method according to claim 1, wherein: Modifying the image based on a distance of the at least one eye relative to the camera comprises: Based on a distance of the at least one eye relative to the camera, a partial blur is applied to one or more background portions of the image.

10. An image processing system, comprising: Camera, with intrinsic matrix; and A computing system configured to: obtaining an image depicting a person from the camera; determining a facial mesh for the face based on one or more features of the face of the person, wherein the facial mesh comprises a combination of facial landmarks and eye landmarks; estimating an eye pixel size of at least one eye of the face based on eye landmarks of the face mesh; estimating a distance of the at least one eye of the face relative to the camera based on the eye pixel size, an intrinsic matrix of the camera, and an average eye size, wherein the average eye size indicates a size of a person's eye based on the eye data; as well as The image is modified based on a distance of the at least one eye relative to the camera.

11. The system according to claim 10, wherein: The eye pixel size is based on a pixel size of an iris of the at least one eye represented by the image.

12. The system according to claim 11, wherein: The intrinsic matrix indicates a pixel focal length of the camera and a principal point on an image sensor of the camera; and Wherein the computing system is further configured to estimate a distance of the at least one eye relative to the camera based on a combination of the pixel focal length, a principal point on the image sensor, the eye pixel size, and an average eye size.

13. The system according to claim 10, wherein: Based on the image, the facial landmarks of the facial mesh represent an outline of the face, and the eye landmarks of the facial mesh represent an outline of at least one eye of the face.

14. The system according to claim 13, wherein: The eye landmark is located around the iris of the at least one eye.

15. The system of claim 14, wherein: The computing system is configured to estimate a pixel size of an iris of the at least one eye based on an eye landmark located around the iris of the at least one eye, wherein the pixel size of the iris corresponds to the eye pixel size.

16. The system of claim 15, wherein: The computing system is also configured to: making a comparison between a pixel size of the iris of the at least one eye and an average pixel iris size; as well as Based on the comparison, determining whether the pixel size of the iris satisfies a threshold difference; and Wherein the computing system is configured to estimate a distance of the at least one eye relative to the camera in response to determining that a pixel size of the iris satisfies the threshold difference.

17. The system according to claim 10, wherein: The computing system is configured to: The image is modified by applying a partial blur to one or more background portions of the image based on a distance of the at least one eye relative to the camera.

18. A non-transitory computer-readable medium configured to store instructions that, when executed by a computing system comprising one or more processors, cause the computing system to perform operations comprising: Obtaining an image depicting a person from a camera; determining a facial mesh for the face based on one or more features of the face of the person, wherein the facial mesh includes a combination of facial landmarks and eye landmarks; estimating an eye pixel size of at least one eye of the face based on eye landmarks of the face mesh; estimating a distance of the at least one eye relative to the camera based on the eye pixel size, an intrinsic matrix of the camera, and an average eye size, wherein the average eye size indicates a size of a person's eye based on eye data; and The image is modified based on a distance of the at least one eye relative to the camera.

19. The non-transitory computer readable medium of claim 18, wherein: Estimating an eye pixel size of at least one eye of the face comprises: A number of pixels indicative of a horizontal diameter of an iris of the at least one eye represented in the image is estimated.

Citation Information

Patent Citations

  • Background modification in video conferencing

    US20150195491A1

  • Image Classification Based On Camera-to-Object Distance

    US20170169570A1