Multi-camera video capture for accurate skin tone measurements
A multi-camera system with machine learning models accurately estimates skin color by analyzing lighting conditions and facial images, addressing challenges in online beauty product recommendations.
Patent Information
- Application Number
- JP2025540888
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-17
- Filing Date
- 2024-01-12
- Publication Date
- 2026-01-29
AI Technical Summary
Existing technologies face challenges in accurately estimating skin color from images captured by consumers due to variations in lighting conditions and user-induced movements, making it difficult to recommend beauty products like foundations online.
A system using multiple cameras on a mobile device captures videos with different fields of view, employing machine learning models to analyze lighting conditions and facial images, combining determinations to provide accurate skin color estimates.
This approach enhances the accuracy of skin color estimation, enabling personalized product recommendations by averaging skin color determinations across varied lighting conditions and reducing user-induced errors.
Smart Images

Figure 2026503444000001_ABST
Abstract
Description
[Technical Field]
[0001] Multi-camera video capture for accurate skin tone measurements. [Background technology]
[0002] Multi-camera video capture for accurate skin tone measurements. Summary of the Invention [Means for solving the problem]
[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0004] In some embodiments, a computing device performs a method for estimating skin color based on facial images, the method including: acquiring a first video recorded by a first camera having a first field of view in an illuminated environment; acquiring a second video recorded by a second camera having a second field of view different from the first field of view in the illuminated environment, the second field of view being directed toward a face of a live subject; extracting a plurality of lighting environment images from the first video; extracting a plurality of face images including the face of the live subject from the second video; processing the face images and the lighting environment images to obtain a plurality of determinations of facial skin color; combining, by the computing device, the determinations of facial skin color to determine a combined determination of skin color; and presenting the combined determination of skin color. In some embodiments, the method further includes constructing a panoramic image from the first video.
[0005] In some embodiments, processing the face images and lighting environment images includes running at least one machine learning model using the face images and lighting environment images as inputs to generate as output a facial skin color determination. In some embodiments, this includes running a first machine learning model using the face images and lighting environment images as inputs to generate as output lighting condition information, such as a light source color for each face image, and running a second machine learning model using the face images and corresponding lighting condition information as inputs to generate as output a facial skin color determination.
[0006] In some embodiments, the skin color determinations are evaluated for suitability in further processing steps before being combined to obtain an integrated skin color determination. In an exemplary scenario, the skin color determinations each include a corresponding confidence level. The computing device compares the corresponding confidence level to a threshold confidence level and excludes the skin color determination from the combining step if the corresponding confidence level is less than the threshold confidence level.
[0007] In some embodiments, the system comprises a skin color estimation unit including a computational circuit configured to: acquire a first video recorded by a first camera having a first field of view in an illuminated environment; acquire a second video recorded by a second camera having a second field of view different from the first field of view in the illuminated environment, the second field of view being directed toward a face of a real subject; extract a plurality of lighting environment images from the first video; extract a plurality of face images including the face of the real subject from the second video; process the face images and the lighting environment images to obtain a plurality of determinations of facial skin color; integrate the determinations of facial skin color to determine an integrated determination of facial skin color; and present the integrated determination of facial skin color.
[0008] In some embodiments, the first camera is a rear-facing camera, and the computing circuitry is further configured to construct a panoramic image from the first video. In some embodiments, the computing circuitry is further configured to run at least one machine learning model using the face image and the lighting environment image as inputs to generate as output a determination of facial skin color.
[0009] In some embodiments, the computing circuitry is further configured to run a first machine learning model using the facial images and the lighting environment image as inputs to generate lighting condition information for each facial image as output, and to run a second machine learning model using the facial images and corresponding lighting condition information as inputs to generate a facial skin color determination as output.
[0010] In some embodiments, a non-transitory computer-readable medium stores computer-executable instructions that, upon execution by one or more processors of the computer system, cause the computer system to perform actions including presenting a user interface on a mobile computing device that instructs a user to capture a panoramic image in an illuminated environment; in response to user input via the user interface, recording a first video with a rear-facing camera of the mobile computing device in the illuminated environment and recording a second video with a front-facing camera of the mobile computing device in the illuminated environment; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images including the user's face from the second video; processing the facial images and the lighting environment images to obtain a plurality of determinations of facial skin color; and aggregating the plurality of determinations of facial skin color to determine an integrated determination of facial skin color.
[0011] The foregoing aspects and many of the attendant advantages of this invention will become more readily appreciated as the same become better understood by reference to the following detailed description, when considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram of a non-limiting, exemplary embodiment of a system for automatically estimating skin color using video recorded by multiple cameras, according to various aspects of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating a non-limiting, exemplary embodiment of a mobile computing device and a skin color determination device, according to various aspects of the present disclosure. [Figure 3] FIG. 10 is a schematic diagram illustrating a non-limiting exemplary embodiment of a user interface that prompts a user to record a first video with a rear camera having a first field of view while a second video is being recorded with a front camera having a second field of view aimed at the face of a real subject, according to various aspects of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram illustrating a non-limiting exemplary embodiment of normalizing an image containing a face, according to various aspects of the present disclosure. [Figure 5] 1 is a flowchart illustrating a non-limiting exemplary embodiment of a method for automatically estimating skin color using video recorded by multiple cameras, according to various aspects of the present disclosure. [Figure 6] FIG. 1 is a block diagram illustrating an exemplary computing device suitable for use as a computing device of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] There are some types of products for which it is difficult to replace an in-person experience with an online interaction. For example, beauty products, such as foundation or other makeup products, are difficult to browse online and difficult to recommend in an automated manner. This is primarily due to the difficulty of automatically estimating skin color. While it is possible to capture images or videos of consumers, it is difficult to ensure that such images can be reliably used to determine skin color. First, it is difficult to instruct consumers on how to capture such images in a manner that is both easy and accurate for them. Second, even if consumers are successfully instructed on how to accurately capture such images, variations in lighting conditions in the consumer's environment often result in inaccurate skin color estimates. A technique that can overcome these technical limitations is needed to accurately estimate skin color regardless of lighting conditions.
[0014] In some embodiments of the present disclosure, one or more machine learning models are trained to accurately estimate skin color in an image regardless of lighting conditions. In some embodiments, these models can then be used to estimate skin color for new images, and the estimated skin color can be used for various purposes. For example, the skin color can be used to generate recommendations for foundation shades that accurately match the skin color or recommendations for other cosmetic products that complement the estimated skin color.
[0015] In some embodiments, the computer system uses video recorded by multiple cameras to automatically estimate skin color using one or more machine learning models, as described in more detail below, thus eliminating the need for in-person product testing and significantly improving the computer system's ability to automatically determine skin color from images taken by average consumers who lack photography skills or technical expertise.
[0016] 1 is a schematic diagram of a non-limiting, exemplary embodiment of a system for generating an automatic skin color estimate according to various aspects of the present disclosure. As shown, a user 90 has a mobile computing device 102. The user 90 takes a first video with a rear-facing camera and a second video with a front-facing camera in an illuminated environment (e.g., a room with artificial or natural light). The rear-facing camera has a field of view 120 facing away from the user 90, and the front-facing camera (sometimes referred to as a "selfie" camera) has a field of view 122 pointed toward the user 90.
[0017] The mobile computing device 102 transmits the first video and the second video to a skin color determination device 104. In some embodiments, the skin color determination device 104 makes a skin color determination 108 based on the images in the video using one or more machine learning models 106. In some embodiments, the skin color 108 may then be used to make a recommendation of one or more products 110 that complement or are otherwise appropriate for the skin color 108.
[0018] To accurately estimate skin color, the system acquires images of the user's face under different lighting conditions. Lighting conditions can be changed by moving the mobile computing device 102 while capturing video. In some embodiments, different lighting conditions are achieved by instructing the user 90 to change orientation while recording a self-portrait or "selfie" video with the mobile computing device 102. As the user changes orientation, the orientation of the user's face changes relative to nearby light sources, and individual frames of the video capture the user's face in these different lighting conditions.
[0019] However, in a typical “selfie video” usage scenario, the user 90 is presented with a preview of the video via the display of the mobile computing device 102. This presentation can be distracting, and the user 90 may change their facial expression, turn their head from side to side, or tilt their head back and forth. Even if the lighting environment is well controlled, these types of movements or changes in facial expression can make it difficult for the system to determine true skin color. Such movements can cause shadows or other effects that are present in some frames but absent or changing in others. Facial features may appear and disappear in frames depending on the user's head movements. Additionally, such movements can affect the accuracy of face detection techniques, which may be used to normalize extracted video frames, such as by ensuring that the user's face remains within the field of view of the front camera or by centering the user's face within the frame. In such a scenario, if the user 90 moves their head and, for example, obscures both eyes, this can adversely affect the face detection or image normalization process.
[0020] To reduce the likelihood of such problems, in some embodiments, the system prompts the user 90 to record a panoramic image using the rear camera while simultaneously recording the user's face with the front camera. The system prompts the user 90 to slowly rotate in place while holding the mobile computing device 102 horizontally. This approach directs the user's 90 focus toward accurately capturing the panoramic image, rather than the selfie video. The typical "preview" image of a "selfie" video is also omitted from the mobile computing device 102 display. By redirecting the user's focus and removing distracting elements, the user is more likely to maintain a neutral facial expression and head orientation, keeping the camera in a position that better captures more accurate and consistent facial images with the front camera.
[0021] 1, a user 90 rotates in place while holding a mobile computing device 102 in a "selfie" position to change the lighting conditions experienced. As shown, the user captures a panoramic image video with the rear camera while holding the mobile computing device 102 at a distance sufficient to allow the user's face to be completely captured by the front camera.
[0022] In addition to encouraging the user 90 to maintain a stable head and a neutral facial expression, this approach allows the user 90 to measure lighting conditions from multiple viewpoints with multiple cameras without further effort. Measuring different lighting conditions in this manner provides the machine learning model 106 with additional input to improve the accuracy of its individual skin color estimates, which can then be averaged or otherwise combined to further improve the accuracy of the overall skin color determination 108.
[0023] In some embodiments, the first machine learning model is trained to take images extracted from a video (e.g., a face image and a lighting environment image) as input and generate lighting condition information as output. In some embodiments, the lighting condition information generated as output by the first machine learning model includes an estimated illuminant (light source) color. In some embodiments, the illuminant color of a face image is estimated by predicting the illuminant color from image information of the face image, predicting the illuminant color in a corresponding lighting environment image captured simultaneously with the face image, and then averaging or otherwise combining these predicted illuminant colors. In this way, the illuminant color prediction can be more accurate for a given face image than using only the face image itself (from a single camera, such as a front-facing camera). The lighting condition information may also include information other than color, such as predicted light source intensity or direction.
[0024] Once the lighting condition information is generated, the corresponding facial image can be adjusted to account for the color of the light source(s) in the lighting environment. This helps ensure that the system can accurately estimate skin color in everyday lighting environments where the color, intensity, and location of the light source vary depending on the user's location. For example, if the system determines that the lighting environment includes a light source that has a yellow color, the system can adjust the corresponding facial image to account for the particular yellow color of the light source, thereby avoiding inaccurately estimating skin color as yellower than it appears under a white light source.
[0025] In some embodiments, the second machine learning model is trained to take the extracted face image and corresponding lighting condition information as input and output an estimated skin color. In other embodiments, a single machine learning model may be trained to take the extracted image as input and output an estimated skin color. Using a first machine learning model and a second machine learning model may be beneficial in that the lighting condition information can be used to correct the presented color of the image or detect the color of other objects in the image, while using a single machine learning model may be beneficial in that it reduces training time and complexity.
[0026] In some embodiments, the machine learning model may be a neural network, including but not limited to a feedforward neural network, a convolutional neural network (CNN), a recurrent neural network, and a generative adversarial network (GAN). In some embodiments, any suitable training technique may be used, including gradient descent, which may include but is not limited to stochastic gradient descent, batch gradient descent, and mini-batch gradient descent.
[0027] 2 is a block diagram illustrating a non-limiting, exemplary embodiment of a mobile computing device and a skin color determination device according to various aspects of the present disclosure. As described above, the mobile computing device 102 is used to capture video of the user 90 using a front-facing camera from which a facial image is extracted and to capture additional video using a rear-facing camera from which a lighting environment image is extracted. The mobile computing device 102 transmits the video data to the skin color determination device 104 to determine the skin color of the user 90. The mobile computing device 102 and the skin color determination device 104 may communicate using any suitable communication technology, such as wireless communication technologies, including but not limited to Wi-Fi, Wi-MAX, Bluetooth, 2G, 3G, 4G, 5G, and LTE, or wired communication technologies, including but not limited to Ethernet, FireWire, and USB. In some embodiments, communication between the mobile computing device 102 and the skin color determination device 104 may occur at least in part over the Internet.
[0028] In some embodiments, the mobile computing device 102 is a smartphone, a tablet computing device, or another computing device having at least the components shown. In the illustrated embodiment, the mobile computing device 102 includes a front-facing camera 202A, a rear-facing camera 202B, a data collection engine 204, and a user interface engine 206.
[0029] In some embodiments, the data collection engine 204 is configured to acquire video captured by the front camera 202A and extract facial images of the user 90 from the video, and to acquire video captured by the rear camera 202B and extract lighting environment images from the video. The data collection engine 204 may also be configured to capture training images for training the machine learning model 106.
[0030] In some embodiments, the user interface engine 206 is configured to present a user interface for capturing video data from which facial images and lighting environment images are extracted. In some embodiments, the user interface engine 206 uses a graphical user interface (GUI), a voice interface, or other interface design to guide the user in capturing video, while omitting the typical “selfie preview” mode in favor of a panoramic image capture interface. In some embodiments, the panoramic image capture interface includes one or more graphical guides to assist the user 90 in capturing video and stabilizing the rear camera (and thus the front camera) during video capture. Such features, in some embodiments, include horizontal guidelines and a progress bar.
[0031] 3 is a schematic diagram illustrating a non-limiting example embodiment of a graphical user interface presentation according to various aspects of the present disclosure. As shown, the mobile computing device 102 displays a graphical user interface 300 including instructions for capturing a panoramic image as well as a graphical guide to assist the user in capturing video and stabilizing the rear camera (and thus the front camera) while capturing the video. In the illustrated example, the user is instructed to hold the mobile computing device 102 level and move it continuously to capture the panoramic image. The user is instructed to keep the center of the graphical arrow aligned on a horizontal line during this movement. In some embodiments, the movement and positioning of the graphical arrow depends on data received from sensors, such as an accelerometer and gyroscope, of the mobile computing device 102, which are used to detect the movement and orientation of the mobile computing device 102. In some embodiments, the graphical arrow progresses along a horizontal line as the video is captured. In some embodiments, the progression of the graphical arrow along the horizontal line depends at least in part on the number of suitable images extracted from the video. For example, once a predetermined number (e.g., 6, 10, etc.) of images suitable for use in the image processing stage have been extracted, the graphical arrow may be moved to the end of the line (or otherwise moved or modified to indicate that processing is complete). Exemplary techniques for determining whether images are suitable for use are described below.
[0032] Referring again to FIG. 2 , in some embodiments, the skin color determination device 104 is a desktop computing device, a server computing device, a cloud computing device, or another computing device that provides the illustrated components. In the illustrated embodiment, the skin color determination device 104 includes a training engine 208, a skin color determination engine 210, an image normalization engine 212, and a product recommendation engine 214. Generally, the term “engine” as used herein refers to logic embodied in hardware or software instructions, which may be written in a programming language such as C, C++, COBOL, JAVA, PHP, Perl, HTML, CSS, JavaScript, VBScript, ASPX, Microsoft .NET™, etc. An engine may be compiled into an executable program or written in an interpreted programming language. Software engines may be callable from other engines or from themselves. Generally, engines described herein refer to logical modules that may be combined with other engines or divided into sub-engines. The engine may be stored in any type of computer-readable medium or computer storage device and may be stored on and executed by one or more general-purpose computers, thus forming the engine or special-purpose computers configured to provide its functionality.
[0033] As shown, the skin color determination device 104 also includes a training data store 216, a model data store 218, and a product data store 220. As will be appreciated by those skilled in the art, a “data store” as described herein may be any suitable device configured to store data for access by a computing device. One example of a data store is a reliable, high-speed relational database management system (DBMS) running on one or more computer devices and accessible over a high-speed network. Another example of a data store is a key-value store. However, any other suitable storage technique and / or device capable of quickly and reliably providing stored data in response to a query may be used, and the computer devices may be accessible locally instead of over a network or may be provided as a cloud-based service. A data store may also include data organized and stored on a computer-readable storage medium, as described further below. Those skilled in the art will recognize that the separate data stores described herein may be combined into a single data store and / or that the single data store described herein may be separated into multiple data stores without departing from the scope of the present disclosure.
[0034] In some embodiments, the training engine 208 is configured to access training data stored in the training data store 216 and generate one or more machine learning models using the training data. The training engine 208 can store the generated machine learning models in the model data store 218. In some embodiments, the skin color determination engine 210 is configured to process images using the one or more machine learning models stored in the model data store 218 to estimate skin colors depicted in the images. In some embodiments, the image normalization engine 212 is configured to preprocess images before providing them to the training engine 208 or the skin color determination engine 210 to improve the accuracy of the determinations made. In some embodiments, the product recommendation engine 214 is configured to recommend one or more products stored in the product data store 220 based on the determined skin colors.
[0035] 4 is a schematic diagram illustrating a non-limiting, exemplary embodiment of normalizing an image including a face, according to various aspects of the present disclosure. As shown, image 402 includes an off-center face that occupies a small portion of the overall image 402. Therefore, it may be difficult to estimate skin color based on image 402. It may be desirable to reduce the amount of non-face areas in the image and consistently position face areas within the image.
[0036] In a first normalization action 404, the image normalization engine 212 uses a face detection algorithm to detect portions of the image 404 that depict a face. The image normalization engine 212 may use the face detection algorithm to find a bounding box 406 that contains the face. In a second normalization action 408, the image normalization engine 212 may modify the image 408 so that the bounding box 406 is centered within the image 408. In a third normalization action 410, the image normalization engine 212 may zoom the image 410 so that the bounding box 406 is as large as possible within the image 410. By performing normalization actions, the image normalization engine 212 may reduce layout and size differences between multiple images, thereby improving the training and accuracy of machine learning models and the accuracy of results when new images are applied to the machine learning models. In some embodiments, different normalization actions may be performed. For example, in some embodiments, instead of centering and zooming the bounding box 406, the image normalization engine 212 may crop the image to the bounding box 406. As another example, in some embodiments, the image normalization engine 212 may decrease or increase the bit depth or undersample or oversample the pixels of an image to match other images collected by a different camera 202 or a different mobile computing device 102.
[0037] In some embodiments, the normalization process and machine learning models described herein may be customized to the particular type of mobile computing device 102 or camera 202A, 202B used to collect training images or images of real-world subjects. In some embodiments, differences between images captured during normalization may be minimized.
[0038] 2 illustrates various components as being provided by the mobile computing device 102 or the skin color determination device 104, in some embodiments, the layout of the components may differ. For example, in some embodiments, the skin color determination engine 210 and the model data store 218 may reside on the mobile computing device 102 such that the mobile computing device 102 can determine skin color in an image captured by the front camera 202A without transmitting the image to the skin color determination device 104. As another example, in some embodiments, all components may be provided by a single computing device. As yet another example, in some embodiments, multiple computing devices may work together to provide the functionality illustrated as being provided by the skin color determination device 104.
[0039] 5 is a flowchart illustrating a non-limiting, exemplary embodiment of a method for estimating facial skin color of a real subject in an image using images extracted from multiple camera views, according to various aspects of the present disclosure. Method 500 refers to the subject as a “real subject” to distinguish it from training subjects that may be used in the training process of a machine learning model that may be used in method 500. While ground truth skin color information is available for training subjects, such information is not available for real subjects.
[0040] In block 502, the computing device acquires a first video recorded by a first camera having a first field of view in an illuminated environment. In some embodiments, the data collection engine 204 of the mobile computing device 102 (e.g., an Apple iPhone® 14) uses the rear-facing camera 202B of the mobile computing device 102 to capture video of an environment in which a real subject is located. In block 504, the computing device acquires a second video recorded by a second camera having a second field of view in the illuminated environment. The second field of view is different from the first field of view and is directed toward the face of the real subject. In some embodiments, the data collection engine 204 of the mobile computing device 102 uses the front-facing camera 202A of the mobile computing device 102 to capture video of the face of the real subject. In some embodiments, the two videos can start and stop recording simultaneously or otherwise cover the same or substantially the same time period. This allows the illuminated environment image to better match the lighting conditions present when the face image is extracted.
[0041] In block 506, the computing device extracts lighting environment images from the first video, and in block 508, the computing device extracts facial images containing the faces of real subjects from the second video. In some embodiments, each facial image is paired with a corresponding lighting environment image from the same or substantially the same time instance. In some embodiments, the extraction of facial images includes one or more preprocessing steps in which frames of the second video are evaluated for facial image suitability in further processing steps. In an exemplary scenario, the computing device measures the luminance level of each frame and excludes frames from the extracted facial images if the measured luminance level is below a minimum threshold luminance level or above a maximum threshold luminance level. Alternatively or additionally, other suitability assessments may be performed, such as verifying the presence of an entire face in the frame or checking for blurring or other image quality issues. In some embodiments, the data collection engine 204 sends the extracted images to the skin color determination device 104 for further processing.
[0042] In block 510, the computing device processes the extracted face image and the lighting environment image to obtain a determination of the facial skin color of the real subject. In some embodiments, the computing device performs image normalization as part of step 510. In some embodiments, the normalization actions illustrated in FIG. 4 may be applied to the face image.
[0043] In some embodiments, processing the face images and lighting environment images includes running at least one machine learning model using the face images and lighting environment images as inputs to generate, as output, a determination of facial skin color. In the example shown in Figure 5, at block 512, the computing device runs a first machine learning model using the face images and lighting environment images as inputs to generate, as output, lighting condition information for each face image, and at block 514, the computing device runs a second machine learning model using the face images and corresponding lighting condition information as inputs to generate, as output, a determination of facial skin color. In some embodiments, the lighting condition information includes a light source color.
[0044] At block 520, the computing device combines the individual determinations of skin color to obtain an integrated determination of the skin color of the real subject. In some embodiments, the estimated skin colors for the normalized images may be averaged or otherwise combined to generate a final skin color determination.
[0045] In some embodiments, the skin color determinations are evaluated for suitability in further processing steps before being combined to obtain an integrated skin color determination. In an exemplary scenario, the skin color determinations each include a corresponding confidence level. The confidence levels may be generated by a machine learning model for each skin color determination. The computing device compares the corresponding confidence level to a threshold confidence level and, if the corresponding confidence level is less than the threshold confidence level, excludes or gives a low weighting to the corresponding skin color determination in the combining step.
[0046] At block 512, the computing device presents the integrative determination of skin color. In some embodiments, the skin color determination engine 210 sends the integrative determination of skin color to the mobile computing device 102. In some embodiments, the user interface engine 206 can present a category, classification, numeric value, or other indication of skin color to the user. In some embodiments, the user interface engine 206 can recreate the color for presentation to the user on the display of the mobile computing device 102.
[0047] In some embodiments, the product recommendation engine 214 of the skin color determination device 104 determines one or more products to recommend based on skin color. The user interface engine 206 can then present representations of such products to the user and enable the user to purchase or otherwise acquire the products.
[0048] In some embodiments, the product recommendation engine 214 can determine one or more products in the product data store 220 that match the determined skin tone. This can be particularly useful for products such as foundations that are intended to match the user's skin tone. In some embodiments, the product recommendation engine 214 can determine one or more products in the product data store 220 that complement the skin tone but do not match the skin tone. This can be particularly useful for products such as eye shadows and lip colors where an exact skin tone match is less desirable. In some embodiments, the product recommendation engine 214 can use another machine learning model, such as a recommender system, to determine products from the product data store 220 based on products preferred by other users with matching or similar skin tones. In some embodiments, if existing products do not match the skin tone, the product recommendation engine 214 can determine ingredients for creating a product that matches the skin tone and provide the ingredients to a formulation system for creating a custom product that matches the skin tone.
[0049] Figure 6 is a block diagram illustrating aspects of an exemplary computing device 600 suitable for use as a computing device of the present disclosure. While several different types of computing devices are described above, the exemplary computing device 600 represents various elements common to many different types of computing devices. While Figure 6 is described with reference to a computing device implemented as a device on a network, the following description may apply to servers, personal computers, mobile phones, smartphones, tablet computers, embedded computing devices, and other devices that may be used to implement some embodiments of the present disclosure. Moreover, those skilled in the art and others will recognize that computing device 600 may be any number of currently available or yet to be developed devices.
[0050] In its most basic configuration, computing device 600 includes at least one processor 602 and a system memory 604 connected by a communications bus 606. Depending on the exact configuration and type of device, the system memory 604 may be volatile or non-volatile memory, such as read-only memory (“ROM”), random-access memory (“RAM”), EEPROM, flash memory, or similar memory technologies. Those skilled in the art and others will recognize that system memory 604 typically stores data and / or program modules that are immediately accessible to and / or presently being processed by the processor 602. In this regard, the processor 602 may act as the computational center of the computing device 600 by supporting the execution of instructions.
[0051] 6, computing device 600 may include a network interface 610 comprising one or more components for communicating with other devices over a network. Embodiments of the present disclosure may utilize network interface 610 to access basic services that communicate using a common network protocol. Network interface 610 may include a wireless network interface configured to communicate via one or more wireless communication protocols, such as WiFi, 2G, 3G, 4G, LTE, 5G, WiMAX, Bluetooth, Bluetooth low energy, etc. As will be appreciated by those skilled in the art, network interface 610 shown in FIG. 6 may represent one or more wireless or physical communication interfaces described and illustrated above with respect to particular components of system 100.
[0052] In the exemplary embodiment shown in Figure 6, computing device 600 also includes storage medium 608. However, services may be accessed using computing devices that do not include means for persisting data to local storage medium. Accordingly, storage medium 608 shown in Figure 6 is represented using a dotted line to indicate that storage medium 608 is optional. In either case, storage medium 608 may be volatile or non-volatile, removable or non-removable, and may be implemented using any technology capable of storing information, including, but not limited to, a hard drive, solid state drive, CD ROM, DVD, or other disk storage, magnetic cassette, magnetic tape, or magnetic disk storage.
[0053] The term "computer-readable medium" as used herein includes volatile and nonvolatile media, and removable and non-removable media implemented in any method or technology capable of storing information such as computer-readable instructions, data structures, program modules, or other data. In this regard, system memory 604 and storage media 608 shown in Figure 6 are examples of computer-readable media.
[0054] Suitable implementations of a computing device including a processor 602, system memory 604, communication bus 606, storage medium 608, and network interface 610 are known and commercially available. For ease of explanation and because they are not important to an understanding of the claimed subject matter, FIG. 6 does not show some of the representative components of many computing devices. In this regard, computing device 600 may include input devices such as a keyboard, keypad, mouse, microphone, touch input device, touch screen, tablet, etc. Such input devices may be coupled to computing device 600 by wired or wireless connections, including RF, infrared, serial, parallel, Bluetooth, Bluetooth low energy, USB, or other suitable connection protocols using wireless or physical connections. Similarly, computing device 600 may include output devices such as a display, speakers, printer, etc. These devices are well known in the art and will not be further shown or described herein.
[0055] While exemplary embodiments have been illustrated and described, it will be understood that various modifications may be made therein without departing from the spirit and scope of the present invention. For example, while the embodiments described above train and use models to estimate skin color, in some embodiments, skin characteristics other than skin color may be estimated. For example, in some embodiments, one or more machine learning models may be trained to estimate Fitzpatrick skin type using techniques similar to those described above with respect to skin color, and such models may then be used to estimate Fitzpatrick skin type for images of real subjects. The embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows: [Explanation of symbols]
[0056] 90 users 102 Mobile Computing Devices 104 Skin Color Determination Device 106 Machine Learning Models 108 Skin color determination 120 field of view 122 Field of view 202A Front Camera 202B rear camera 204 Data Collection Engine 206 User Interface Engine 208 Training Engine 210 Skin Color Determination Engine 212 Image Normalization Engine 214 Product Recommendation Engine 216 Training Data Store 218 Model Data Store 220 Product Data Store 300 Graphical User Interface 402 images 404 first canonical action 404 images 406 Bounding Box 408 Second Canonicalization Action 408 images 410 Third Normalization Action 410 images 600 computing devices 602 processor 604 system memory 606 Communication Bus 608 Storage medium 610 Network Interface
Claims
1. 1. A method for estimating skin color based on a facial image, comprising: acquiring, by a computing device, a first video recorded by a first camera having a first field of view in an illuminated environment; acquiring, by the computing device, a second video recorded by a second camera having a second field of view in the illuminated environment that is different from the first field of view, the second field of view being directed toward a face of a real-world subject; extracting, by the computing device, a plurality of lighting environment images from the first video; extracting, by the computing device, a plurality of facial images from the second video, the facial images including the face of the real-world subject; processing, by the computing device, the face image and the lighting environment image to obtain a plurality of determinations of skin color of the face; aggregating, by the computing device, the determinations of skin color of the face to determine an aggregate determination of skin color; presenting, by the computing device, the integrated determination of skin color; A method comprising:
2. The method of claim 1 , further comprising constructing a panoramic image from the first video.
3. 10. The method of claim 1, wherein processing the face image and the lighting environment image comprises running at least one machine learning model using the face image and the lighting environment image as inputs to generate the determination of the skin color of the face as output.
4. executing the at least one machine learning model running a first machine learning model using the face images and the lighting environment image as inputs to generate lighting condition information for each of the face images as output; running a second machine learning model using the facial image and the corresponding lighting condition information as input to generate the determination of the facial skin color as output; 4. The method of claim 3, comprising:
5. The method of claim 4 , wherein the lighting condition information includes a light source color.
6. The method of claim 4 , wherein the determinations of skin color each include a corresponding confidence level.
7. comparing the corresponding confidence level to a threshold confidence level; excluding the skin color determination from the combining step if the corresponding confidence level is less than the threshold confidence level; The method of claim 6 further comprising:
8. The step of extracting the face image from the second video includes, for each frame of the second video: measuring the luminance level of said frame; excluding the frame from the extracted face image if the measured luminance level is below a minimum threshold luminance level or above a maximum threshold luminance level; 2. The method of claim 1, comprising:
9. The method of claim 1 , wherein the first video is recorded simultaneously with the second video.
10. acquiring a first video recorded by a first camera having a first field of view in an illuminated environment; acquiring a second video recorded by a second camera having a second field of view different from the first field of view in the illuminated environment, the second field of view being directed toward a face of a real-world subject; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images from the second video, the facial images including the face of the real-world subject; processing the face image and the lighting environment image to obtain a plurality of determinations of skin color of the face; aggregating the determinations of the facial skin color to determine an aggregate determination of the facial skin color; presenting the integrated determination of the skin color of the face; and a skin color estimation unit including a computation circuit configured to perform A system comprising:
11. 11. The system of claim 10, wherein the first camera is a rear-facing camera, and the computing circuitry is further configured to construct a panoramic image from the first video.
12. 11. The system of claim 10, wherein the computing circuitry is further configured to execute at least one machine learning model using the face image and the lighting environment image as inputs to generate the determination of the facial skin color as an output.
13. The calculation circuitry running a first machine learning model using the face images and the lighting environment image as inputs to generate lighting condition information for each of the face images as output; running a second machine learning model using the facial image and the corresponding lighting condition information as input to generate the determination of the facial skin color as output; The system of claim 10, further configured to:
14. The system of claim 13 , wherein the lighting condition information includes a light source color.
15. The system of claim 13 , wherein the determinations of skin color each include a corresponding confidence level.
16. The computing circuitry, for each of the determinations of skin color: comparing the corresponding confidence level to a threshold confidence level; excluding the skin color determination from the combining step if the corresponding confidence level is less than the threshold confidence level; 16. The system of claim 15, further configured to aggregate the determinations of skin color of the face to determine an aggregate determination of skin color by:
17. The computing circuitry calculates, for each frame of the second video: measuring the luminance level of said frame; If the measured luminance level is less than a minimum threshold luminance level or more than a maximum threshold luminance level, the frame is excluded from the extracted face image. The system of claim 10, further configured to:
18. In response to execution by one or more processors of a computer system, the computer system: presenting a user interface on the mobile computing device that instructs a user to capture a panoramic image in an illuminated environment; recording a first video with a rear camera of the mobile computing device in the illuminated environment and a second video with a front camera of the mobile computing device in the illuminated environment in response to user input via the user interface; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images including a face of the user from the second video; processing the face image and the lighting environment image to obtain a plurality of determinations of skin color of the face; aggregating the plurality of determinations of the skin color of the face to determine an aggregate determination of the skin color of the face; A non-transitory computer-readable medium having stored thereon computer-executable instructions for causing the computer to perform actions including:
19. Processing the face image and the lighting environment image includes: running, by the mobile computing device, a first machine learning model using the face images and the lighting environment image as inputs to generate lighting condition information for each of the face images as output; executing, by the mobile computing device, a second machine learning model using the facial image and the corresponding lighting condition information as input to generate, as output, the determination of skin color of the face; 20. The non-transitory computer-readable medium of claim 18, comprising:
20. 20. The non-transitory computer-readable medium of claim 19, wherein the lighting condition information includes a light source color.