Multi-camera video capture for accurate skin hue measurement
By using rear and front cameras on mobile computing devices combined with machine learning models to process facial and lighting environment images, the accuracy of online estimation of skin tones is solved, improving the accuracy of cosmetic recommendations.
Patent Information
- Application Number
- CN202480007394.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-17
- Filing Date
- 2024-01-12
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to accurately estimate the skin tone of consumers online, especially in variable lighting environments, resulting in inaccuracy in cosmetic recommendations.
By simultaneously recording video using the rear and front cameras of the mobile computing device, processing facial and lighting environment images in combination with the machine learning model, estimating skin tones, including training the first machine learning model to generate lighting status information and the second machine learning model to generate skin tone determination.
Accurate estimation of skin tones under different lighting conditions is achieved, the accuracy of cosmetic recommendations is improved, and the need for offline testing is reduced.
Smart Images

Figure CN120457462A_ABST
Abstract
Description
Summary of the Invention
[0001] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0002] In some embodiments, a computing device performs a method for estimating skin tone based on facial images, the method comprising: obtaining a first video recorded in a lighting environment by a first camera having a first field of view; obtaining a second video recorded in the lighting environment by a second camera having a second field of view different from the first field of view, wherein the second field of view is directed toward the face of a live subject; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images including the face of the live subject from the second video; processing the facial images and the lighting environment images to obtain a plurality of skin tone determinations of the face; combining the skin tone determinations of the face by the computing device to determine a combined skin tone determination; and presenting the combined skin tone determination. In some embodiments, the method further comprises constructing a panoramic image from the first video.
[0003] In some embodiments, processing the facial image and the lighting environment image includes executing at least one machine learning model using the facial image and the lighting environment image as input to generate a skin tone determination for the face as an output. In some embodiments, this includes executing a first machine learning model using the facial image and the lighting environment image as input to generate lighting condition information (such as illuminant color) for each of the facial images as an output; and executing a second machine learning model using the facial image and the corresponding lighting condition information as input to generate a skin tone determination for the face as an output.
[0004] In some embodiments, the skin tone determinations are evaluated for suitability for further processing before being combined to obtain the combined skin tone determination. In one exemplary scenario, the skin tone determinations each include a corresponding confidence level. The computing device compares the corresponding confidence level to a threshold confidence level and, if the corresponding confidence level is less than the threshold confidence level, omits the skin tone determination from the combining step.
[0005] In some embodiments, a system includes a skin tone estimation unit, which includes computing circuitry configured to: obtain a first video recorded in a lighting environment by a first camera having a first field of view; obtain a second video recorded in the lighting environment by a second camera having a second field of view different from the first field of view, wherein the second field of view is directed toward the face of a live subject; extract a plurality of lighting environment images from the first video; extract a plurality of facial images including the face of the live subject from the second video; process the facial images and the lighting environment images to obtain a plurality of skin tone determinations of the face; combine the skin tone determinations of the faces to determine a combined skin tone determination of the face; and present the combined skin tone determination of the face.
[0006] In some embodiments, the first camera is a rear-facing camera, and the computing circuitry is further configured to construct a panoramic image from the first video. In some embodiments, the computing circuitry is further configured to execute at least one machine learning model using the facial image and the lighting environment image as input to generate a skin tone determination of the face as an output.
[0007] In some embodiments, the computing circuit is further configured to: execute a first machine learning model using the facial images and the lighting environment images as input to generate lighting condition information as output for each of the facial images; and execute a second machine learning model using the facial images and the corresponding lighting condition information as input to generate a skin tone determination of the face as output.
[0008] In some embodiments, a non-transitory computer-readable medium includes computer-executable instructions stored thereon that, in response to execution by one or more processors of a computer system, cause the computer system to perform actions including: presenting a user interface on a mobile computing device instructing a user to capture a panoramic image in a lit environment; in response to user input via the user interface, recording a first video with a rear-facing camera of the mobile computing device in the lit environment and a second video with a front-facing camera of the mobile computing device in the lit environment; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images including the user's face from the second video; processing the facial images and the lighting environment images to obtain a plurality of skin tone determinations for the face; and combining the plurality of skin tone determinations for the face to determine a combined skin tone determination for the face. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The foregoing aspects of the present invention, as well as many of the attendant advantages, will become more readily appreciated as the same becomes better understood by reference to the following detailed description when considered in conjunction with the accompanying drawings, in which: Figure 1 is a schematic illustration of one non-limiting example embodiment of a system for automatically estimating skin tone using video recorded by multiple cameras in accordance with various aspects of the present disclosure; Figure 2 is a block diagram illustrating a non-limiting example embodiment of a mobile computing device and a skin tone determination device according to various aspects of the present disclosure; Figure 3 is a schematic diagram illustrating one non-limiting example embodiment of a user interface instructing a user to record a first video with a rear-facing camera having a first field of view while simultaneously recording a second video with a front-facing camera having a second field of view directed toward a face of a live subject in accordance with various aspects of the present disclosure; Figure 4 is a schematic diagram illustrating one non-limiting example embodiment of normalization of an image including a face in accordance with various aspects of the present disclosure; Figure 5 is a flow chart illustrating one non-limiting example embodiment of a method for automatically estimating skin tone using videos recorded by multiple cameras according to various aspects of the present disclosure; and Figure 6 is a block diagram of an illustrative computing device suitable for use as a computing device in accordance with the present disclosure. DETAILED DESCRIPTION
[0010] There are certain types of products for which the in-person experience has been difficult to replace with an online interaction. For example, beauty products such as foundation and other makeup products are difficult to browse online and difficult to recommend automatically. This is primarily due to the difficulty in automatically estimating skin tone. While it is possible to capture images or videos of consumers, it is difficult to ensure that such images can be reliably used to determine skin tone. First, it is difficult to instruct consumers on how to capture such images in a manner that is both accurate and user-friendly. Second, even if consumers are successfully instructed on how to accurately capture such images, the variable lighting conditions in the consumer's environment often lead to inaccurate estimates of skin tone. What is desired is a technology that can overcome these technical limitations and accurately estimate skin tone regardless of the lighting conditions.
[0011] In some embodiments of the present disclosure, one or more machine learning models are trained to accurately estimate skin tone in an image regardless of lighting conditions. In some embodiments, the model can then be used to estimate skin tone in new images, and the estimated skin tone can be used for a variety of purposes. For example, the skin tone can be used to generate recommendations for foundation shades that accurately match the skin tone, or to generate recommendations for other makeup products that complement the estimated skin tone.
[0012] In some embodiments, a computer system uses video recorded by multiple cameras to automatically estimate skin tone using one or more machine learning models, as described in further detail below. Thus, the need for offline testing of products is eliminated, and the computer system's ability to automatically determine skin tone based on images captured by an average consumer who does not have photographic skills or technical expertise is significantly improved.
[0013] Figure 1 FIG2 is a schematic illustration of one non-limiting example embodiment of a system for generating an automatic estimate of skin tone according to various aspects of the present disclosure. As shown, a user 90 has a mobile computing device 102. The user 90 captures a first video using a rear-facing camera and a second video using a front-facing camera in a lit environment (e.g., in a room with artificial or natural lighting). The rear-facing camera has a field of view 120 facing away from the user 90, and the front-facing camera (sometimes referred to as the "selfie" camera) has a field of view 122 pointing toward the user 90.
[0014] Mobile computing device 90 transmits the first video and the second video to skin tone determination device 104. In some embodiments, skin tone determination device 104 uses one or more machine learning models 106 to make a skin tone determination 108 based on the images in the video. In some embodiments, skin tone 108 can then be used to make recommendations for one or more products 110 that complement or are otherwise suitable for skin tone 108.
[0015] To accurately estimate skin tone, the system obtains images of the user's face under different lighting conditions. While capturing video, the lighting conditions can be changed by moving the mobile computing device 102. In some embodiments, different lighting conditions are obtained by instructing the user 90 to turn around while recording a self-portrait or "selfie" video using the mobile computing device 102. As the user turns around, the orientation of the user's face changes relative to nearby light sources, and various frames of the video capture the user's face under these different lighting conditions.
[0016] However, in a typical "selfie video" usage scenario, user 90 is presented with a preview of the video via the display of mobile computing device 102. This presentation can be distracting and may cause user 90 to change their facial expression, turn their head left or right, or tilt their head forward or backward. These types of movements or changes in expression can make it difficult for the system to determine true skin tone, even in well-controlled lighting environments. Such movement can create shadows or other effects that may be present in some frames but absent or altered in others. As the user's head moves, facial features may appear or disappear within the frame. Furthermore, such movement may affect the accuracy of facial detection technology, which can be used to ensure that user 90's face remains within the field of view of the front-facing camera or to normalize the extracted frames of the video, such as by centering the user's face in the frame. In such scenarios, if user 90 moves their head, for example, so that both eyes are no longer clearly visible, facial detection or image normalization may be adversely affected.
[0017] To reduce the likelihood of such issues, in some embodiments, the system instructs user 90 to record a panoramic image using the rear-facing camera while also recording the user's face using the front-facing camera. The system instructs user 90 to keep mobile computing device 102 level while slowly turning it in place. This approach directs user 90's attention to accurately capture the panoramic image, rather than the selfie video. The typical "preview" image of the "selfie" video is also omitted from the display of mobile computing device 102. By redirecting the user's attention and removing distractions, the user is more likely to maintain a neutral expression and head orientation, and keep the camera in a good position to capture a more accurate and consistent facial image using the front-facing camera.
[0018] exist Figure 1 In the example shown in , user 90 rotates in place while holding mobile computing device 102 in a "selfie" position to change the lighting conditions being experienced. As shown, the user holds mobile computing device 102 at a distance sufficient to allow the user's face to be fully captured by the front-facing camera while also capturing a panoramic image video using the rear-facing camera.
[0019] In addition to encouraging the user 90 to maintain a steady head and neutral facial expression, this approach also allows for the measurement of lighting conditions from multiple viewpoints via multiple cameras without requiring any further effort from the user. By measuring different lighting conditions in this way, additional input is provided to the machine learning model 106 to improve the accuracy of individual skin tone estimates, which can then be averaged or otherwise combined to further improve the accuracy of the overall skin tone determination 108.
[0020] In some embodiments, a first machine learning model is trained to take images extracted from a video (e.g., facial images and lighting environment images) as input and to generate lighting condition information as output. In some embodiments, the lighting condition information generated as output by the first machine learning model includes estimated illuminant (light source) color. In some embodiments, the illuminant color for a facial image is estimated by predicting the illuminant color based on image information in the facial image, predicting the illuminant color in a corresponding lighting environment image captured simultaneously with the facial image, and averaging or otherwise combining these predicted illuminant colors. In this way, the illuminant color prediction for a given facial image can be more accurate than using only the facial image itself (from a single camera, such as a front-facing camera). The lighting condition information may also include information other than color, such as the predicted intensity or direction of the light source.
[0021] Once the lighting condition information has been generated, the corresponding facial image can be adjusted to account for the color of the light source (or light sources) in the lighting environment. This helps ensure that the system can accurately estimate skin tone in everyday lighting environments, where the color, intensity, and positioning of the light sources will vary depending on the user's location. For example, if the system determines that the lighting environment includes a light source with a yellow color, the system can adjust the corresponding facial image to account for the specific yellow color of the light source and thereby avoid inaccurately estimating skin tone as more yellow than it would appear under a white light source.
[0022] In some embodiments, a second machine learning model is trained to take the extracted facial image and corresponding lighting condition information as input and to output an estimated skin tone. In other embodiments, a single machine learning model can be trained to take the extracted image as input and to output an estimated skin tone. Using a first machine learning model and a second machine learning model can be beneficial because lighting condition information can be used to correct the rendered color of an image or detect the color of other objects in the image, and using a single machine learning model can be beneficial in reducing training time and complexity.
[0023] In some embodiments, the machine learning model can be a neural network, including but not limited to a feedforward neural network, a convolutional neural network (CNN), a recurrent neural network, and a generative adversarial network (GAN). In some embodiments, any suitable training technique can be used, including but not limited to gradient descent, which can include stochastic, batch, and mini-batch gradient descent.
[0024] Figure 2is a block diagram illustrating a non-limiting example embodiment of a mobile computing device and a skin tone determination device according to various aspects of the present disclosure. As discussed above, mobile computing device 102 is used to capture video of user 90 using a front-facing camera (from which a facial image is extracted), and to capture additional video using a rear-facing camera (from which an image of the lighting environment is extracted). Mobile computing device 102 transmits the video data to skin tone determination device 104 for use in determining user 90's skin tone. Mobile computing device 102 and skin tone determination device 104 can communicate using any suitable communication technology, such as wireless communication technologies including, but not limited to, Wi-Fi, Wi-MAX, Bluetooth, 2G, 3G, 4G, 5G, and LTE; or wired communication technologies including, but not limited to, Ethernet, FireWire, and USB. In some embodiments, communication between mobile computing device 102 and skin tone determination device 104 can occur at least partially over the Internet.
[0025] In some embodiments, the mobile computing device 102 is a smartphone, a tablet computing device, or another computing device having at least the components illustrated. In the illustrated embodiment, the mobile computing device 102 includes a front-facing camera 202A, a rear-facing camera 202B, a data collection engine 204, and a user interface engine 206.
[0026] In some embodiments, the data collection engine 204 is configured to obtain a video captured by the front camera 202A and extract a facial image of the user 90 from the video, and to obtain a video captured by the rear camera 202B and extract a lighting environment image from the video. The data collection engine 204 may also be configured to capture training images for training the machine learning model 106.
[0027] In some embodiments, the user interface engine 206 is configured to present a user interface for capturing video data from which facial images and lighting environment images are extracted. In some embodiments, the user interface engine 206 uses a graphical user interface (GUI), a voice interface, or some other interface design to guide the user in capturing the video, bypassing the typical "selfie preview" mode in favor of a panoramic image capture interface. In some embodiments, the panoramic image capture interface includes one or more graphical guides to assist the user 90 in capturing the video and maintaining a stable rear-facing camera (and, consequently, front-facing camera) during video capture. In some embodiments, such features include horizontal guide lines and a progress bar.
[0028] Figure 3is a schematic diagram illustrating one non-limiting example embodiment of a graphical user interface presentation according to various aspects of the present disclosure. As shown, a mobile computing device 102 displays a graphical user interface 300 that includes instructions for capturing a panoramic image and graphical guidance to assist the user in capturing video and maintaining a stable rear camera (and, in turn, front camera) during video capture. In the illustrated example, the user is instructed to hold the mobile computing device 102 level and continuously move it to capture the panoramic image. The user is instructed to maintain a graphical arrow centered on a horizontal line during this movement. In some embodiments, the movement and positioning of the graphical arrow is dependent on data received from sensors of the mobile computing device 102 (such as an accelerometer and a gyroscope), which are used to detect the movement and orientation of the mobile computing device 102. In some embodiments, the graphical arrow advances along a horizontal line while the video is captured. In some embodiments, the progress of the graphical arrow along the horizontal line depends, at least in part, on the number of suitable images extracted from the video. For example, when a predetermined number of images suitable for use in the image processing stage have been extracted (e.g., 6, 10, etc.), the graphical arrow is moved to the end of the line (or otherwise moved or altered to indicate that the process is complete). An exemplary technique for determining whether an image is suitable for use is described below.
[0029] Reference again Figure 2 In some embodiments, the skin tone determination device 104 is a desktop computing device, a server computing device, a cloud computing device, or another computing device that provides the illustrated components. In the illustrated embodiment, the skin tone determination device 104 includes a training engine 208, a skin tone determination engine 210, an image normalization engine 212, and a product recommendation engine 214. Generally, as used herein, the term "engine" refers to logic embodied in hardware or software instructions, and the program can be written in a programming language such as C, C++, COBOL, JAVA, etc. TM , PHP, Perl, HTML, CSS, JavaScript, VBScript, ASPX, Microsoft .NET TM Etc. The engine can be compiled into an executable program or written in an interpreted programming language. Software engines can be callable from other engines or from themselves. Generally, the engines described herein refer to logical modules that can be merged with other engines or divided into sub-engines. The engine can be stored in any type of computer-readable medium or computer storage device and stored on and executed by one or more general-purpose computers, thereby creating a special-purpose computer configured to provide the engine or its functionality.
[0030] As shown, the skin tone determination device 104 also includes a training data store 216, a model data store 218, and a product data store 220. As one skilled in the art will appreciate, a "data store," as described herein, may be any suitable device configured to store data for access by a computing device. One example of a data store is a highly reliable, high-speed relational database management system (DBMS) executed on one or more computing devices and accessible via a high-speed network. Another example of a data store is a key-value store. However, any other suitable storage technology and / or device capable of quickly and reliably providing stored data in response to queries may be used, and the computing device may be locally accessible rather than via a network, or may be provided as a cloud-based service. A data store may also include data stored in an organized manner on a computer-readable storage medium, as further described below. One skilled in the art will recognize that the separate data stores described herein may be combined into a single data store, and / or that a single data store described herein may be divided into multiple data stores, without departing from the scope of this disclosure.
[0031] In some embodiments, the training engine 208 is configured to access training data stored in the training data store 216 and use the training data to generate one or more machine learning models. The training engine 208 may store the generated machine learning models in the model data store 218. In some embodiments, the skin tone determination engine 210 is configured to process an image using the one or more machine learning models stored in the model data store 218 to estimate the skin tone depicted in the image. In some embodiments, the image normalization engine 212 is configured to pre-process the image before providing it to the training engine 208 or the skin tone determination engine 210 to improve the accuracy of the determination made. In some embodiments, the product recommendation engine 214 is configured to recommend one or more products stored in the product data store 220 based on the determined skin tone.
[0032] Figure 4 FIG4 is a diagram illustrating a non-limiting example embodiment of normalizing an image including a face, according to various aspects of the present disclosure. As shown, image 402 includes an off-center face that occupies only a small portion of the overall image 402. As such, it may be difficult to estimate skin tone based on image 402. It may be desirable to reduce the amount of non-face areas in the image and place the facial area of the image in a consistent location.
[0033] In a first normalization action 404, image normalization engine 212 uses a facial detection algorithm to detect a portion of image 404 depicting a face. Image normalization engine 212 may use the facial detection algorithm to find a bounding box 406 that includes the face. In a second normalization action 408, image normalization engine 212 may modify image 408 so that bounding box 406 is centered within image 408. In a third normalization action 410, image normalization engine 212 may scale image 410 so that bounding box 406 is as large as possible within image 410. By performing these normalization actions, image normalization engine 212 may reduce layout and size differences between multiple images, thereby improving the training and accuracy of a machine learning model and the accuracy of results when applying new images to the machine learning model. In some embodiments, different normalization actions may occur. For example, in some embodiments, image normalization engine 212 may crop the image to bounding box 406 rather than centering and scaling bounding box 406. As another example, in some embodiments, the image normalization engine 212 may reduce or increase the bit depth, or may undersample or oversample the pixels of an image to match other images collected by a different camera 202 or a different mobile computing device 102 .
[0034] In some embodiments, the normalization process and machine learning models described herein can be customized for the specific type of mobile computing device 102 or camera 202A, 202B used to collect training images or images of live broadcast objects. In some embodiments, the differences between the captured images can be minimized during normalization.
[0035] although Figure 2 Various components are illustrated as being provided by either the mobile computing device 102 or the skin tone determination device 104, but in some embodiments, the layout of the components may be different. For example, in some embodiments, the skin tone determination engine 210 and the model data storage 218 may reside on the mobile computing device 102, allowing the mobile computing device 102 to determine skin tone in an image captured by the front-facing camera 202A without sending the image to the skin tone determination device 104. For another example, in some embodiments, all components may be provided by a single computing device. For another example, in some embodiments, multiple computing devices may work together to provide the functionality illustrated as being provided by the skin tone determination device 104.
[0036] Figure 5is a flow chart illustrating one non-limiting example embodiment of a method for estimating skin tone of a face of a live object in an image using images extracted from multiple camera views, in accordance with various aspects of the present disclosure. In method 500, an object is referred to as a "live object" to distinguish it from a training object that may be used in training a machine learning model for method 500. While ground truth skin tone information is available for training objects, such information is not available for live objects.
[0037] At block 502, a computing device obtains a first video recorded in a lit environment by a first camera having a first field of view. In some embodiments, the data collection engine 204 of a mobile computing device 102 (e.g., an Apple iPhone 14) uses the rear-facing camera 202B of the mobile computing device 102 to capture video of the environment in which a live subject is located. At block 504, the computing device obtains a second video recorded in the lit environment by a second camera having a second field of view. The second field of view is different from the first field of view and is directed toward the face of the live subject. In some embodiments, the data collection engine 204 of the mobile computing device 102 uses the front-facing camera 202A of the mobile computing device 102 to capture video of the face of the live subject. In some embodiments, the two videos may start and stop recording simultaneously or otherwise cover the same or substantially the same time period. This allows the image of the lit environment to be more relevant to the lighting conditions present when the facial image was extracted.
[0038] At block 506, the computing device extracts a lighting environment image from the first video, and at block 508, the computing device extracts a facial image including the face of the live subject from the second video. In some embodiments, the facial image is paired with a corresponding lighting environment image from the same or substantially the same time instance. In some embodiments, extraction of the facial image includes one or more pre-processing steps, in which frames of the first video are evaluated for suitability as facial images in further processing steps. In one exemplary scenario, the computing device measures the brightness level of each frame and, if the measured brightness level is less than a minimum threshold brightness level or greater than a maximum threshold brightness level, omits the frame from the extracted facial image. Alternatively or additionally, other suitability assessments may be performed, such as confirming the presence of a complete face in the frame or checking for blur or other image quality issues. In some embodiments, the data collection engine 204 transmits the extracted image to the skin tone determination device 104 for further processing.
[0039] At block 510, the computing device processes the extracted facial image and the lighting environment image to obtain a skin tone determination of the face of the live subject. In some embodiments, the computing device performs normalization of the image as part of step 510. In some embodiments, Figure 4 The standardized actions exemplified in can be applied to facial images.
[0040] In some embodiments, processing the facial image and the lighting environment image includes executing at least one machine learning model using the facial image and the lighting environment image as input to generate a skin tone determination of the face as output. Figure 5 In the example shown in FIG, at block 512, the computing device executes a first machine learning model using the facial images and the lighting environment images as input to generate lighting condition information for each of the facial images as output, and at block 514, the computing device executes a second machine learning model using the facial images and the corresponding lighting condition information as input to generate a skin tone determination of the face as output. In some embodiments, the lighting condition information includes illuminant color.
[0041] At block 520, the computing device combines the individual skin tone determinations to obtain a combined skin tone determination for the live object. In some embodiments, the estimated skin tones for the normalized images are averaged or otherwise combined to generate a final skin tone determination.
[0042] In some embodiments, before the skin tone determinations are combined to obtain a combined skin tone determination, the skin tone determinations are evaluated for suitability in further processing steps. In one exemplary scenario, each skin tone determination includes a corresponding confidence level. The confidence level can be generated for each skin tone determination by a machine learning model. The computing device compares the corresponding confidence level to a threshold confidence level and, if the corresponding confidence level is less than the threshold confidence level, ignores or assigns less weight to the corresponding skin tone determination in the combining step.
[0043] At block 512, the computing device presents the combined skin tone determination. In some embodiments, the skin tone determination engine 210 transmits the combined skin tone determination to the mobile computing device 102. In some embodiments, the user interface engine 206 can present a category, classification, numerical value, or other indication of the skin tone to the user. In some embodiments, the user interface engine 206 can recreate the color for presentation to the user on the display of the mobile computing device 102.
[0044] In some embodiments, the product recommendation engine 214 of the skin tone determination device 104 determines one or more products to recommend based on the skin tone. The user interface engine 206 can then present representations of such products to the user and can allow the user to purchase or otherwise obtain the products.
[0045] In some embodiments, the product recommendation engine 214 can identify one or more products in the product data store 220 that match the determined skin tone. This can be particularly useful for products intended to match the user's skin tone, such as foundation. In some embodiments, the product recommendation engine 214 can identify one or more products in the product data store 220 that complement the skin tone but do not match it. This can be particularly useful for products where an exact match to the skin tone is less desirable, such as eye shadow or lip gloss. In some embodiments, the product recommendation engine 214 can use a separate machine learning model, such as a recommender system, to identify products from the product data store 220 based on other user-preferred products with matching or similar skin tones. In some embodiments, if an existing product does not match the skin tone, the product recommendation engine 214 can determine ingredients for creating a product that will match the skin tone and can provide the ingredients to a compounding system for creating a customized product that will match the skin tone.
[0046] Figure 6 is a block diagram illustrating aspects of an exemplary computing device 600 suitable for use as a computing device for the present disclosure. Although various different types of computing devices have been discussed above, the exemplary computing device 600 describes various elements that are common to many different types of computing devices. Although reference is made to a computing device implemented as a device on a network, Figure 6 , but the following description is applicable to servers, personal computers, mobile phones, smart phones, tablet computers, embedded computing devices, and other devices that can be used to implement various parts of the embodiments of the present disclosure. In addition, those skilled in the art and others will recognize that the computing device 600 can be any one of any number of currently available or yet to be developed devices.
[0047] In its most basic configuration, computing device 600 includes at least one processor 602 and system memory 604 connected via a communication bus 606. Depending on the exact configuration and type of device, system memory 604 can be volatile or non-volatile memory, such as read-only memory ("ROM"), random access memory ("RAM"), EEPROM, flash memory, or similar memory technology. Those skilled in the art and others will recognize that system memory 604 typically stores data and / or program modules that are immediately accessible to and / or currently being operated on by processor 602. In this regard, processor 602 can serve as the computational center of computing device 600 by supporting the execution of instructions.
[0048] like Figure 6 As further illustrated in , the computing device 600 may include a network interface 610 that includes one or more components for communicating with other devices over a network. Embodiments of the present disclosure may access basic services that utilize the network interface 610 to perform communications using common network protocols. The network interface 610 may also include a wireless network interface configured to communicate via one or more wireless communication protocols (such as WiFi, 2G, 3G, 4G, LTE, 5G, WiMAX, Bluetooth, Bluetooth Low Energy, etc.). As will be understood by one of ordinary skill in the art, Figure 6 The network interface 610 illustrated in may represent one or more wireless interfaces or physical communication interfaces described and illustrated above with respect to the particular components of the system 100 .
[0049] exist Figure 6 In the exemplary embodiment depicted in FIG, computing device 600 also includes storage media 608. However, the service may be accessed using a computing device that does not include means for persisting data to a local storage medium. Figure 6 The storage medium 608 depicted in FIG is represented by dashed lines to indicate that the storage medium 608 is optional. Regardless, the storage medium 608 can be volatile or non-volatile, removable or non-removable, implemented using any technology capable of storing information, such as, but not limited to, a hard drive, a solid-state drive, a CD ROM, a DVD or other disk storage device, a magnetic cassette, a magnetic tape, a magnetic disk storage device, etc.
[0050] As used herein, the term "computer-readable media" includes volatile and nonvolatile, and removable and non-removable media implemented in any method or technology capable of storing information such as computer-readable instructions, data structures, program modules or other data. In this regard, Figure 6 The system memory 604 and storage media 608 depicted in FIG. 6 are examples of computer-readable media.
[0051] Suitable implementations of computing devices including processor 602, system memory 604, communication bus 606, storage media 608, and network interface 610 are known and commercially available. For ease of illustration and because they are not important to understanding the claimed subject matter, Figure 6 Some typical components of many computing devices are not shown. In this regard, computing device 600 may include input devices such as a keyboard, keys, a mouse, a microphone, a touch input device, a touch screen, a tablet computer, etc. Such input devices may be coupled to computing device 600 via a wired or wireless connection, including RF, infrared, serial, parallel, Bluetooth, Bluetooth Low Energy, USB, or other suitable connection protocols using wireless or physical connections. Similarly, computing device 600 may also include output devices such as a display, speakers, a printer, etc. Because these devices are well known in the art, they are not further illustrated or described herein.
[0052] While exemplary embodiments have been illustrated and described, it will be appreciated that various changes may be made therein without departing from the spirit and scope of the invention. For example, while the embodiments described above train and use models to estimate skin tone, in some embodiments, skin features other than skin tone may be estimated. For example, in some embodiments, one or more machine learning models may be trained using techniques similar to those discussed above with respect to skin tone to estimate Fitzpatrick skin type, and such models may then be used to estimate Fitzpatrick skin type for images of live subjects.
Claims
1. A method for estimating skin tone based on a facial image, the method comprising: obtaining, by a computing device, a first video recorded by a first camera having a first field of view in a lighting environment; obtaining, by the computing device, a second video recorded in the lighting environment by a second camera having a second field of view different from the first field of view, wherein the second field of view is directed toward a face of a live subject; extracting, by the computing device, a plurality of lighting environment images from the first video; extracting, by the computing device, a plurality of facial images including the face of the live subject from the second video; processing, by the computing device, the facial image and the lighting environment image to obtain a plurality of skin tone determinations of the face; combining, by the computing device, the skin tone determinations for the face to determine a combined skin tone determination; and The combined skin tone determination is presented by the computing device. The method of claim 1 , further comprising constructing a panoramic image from the first video.
3. The method according to claim 1, wherein Processing the facial image and the lighting environment image includes executing at least one machine learning model using the facial image and the lighting environment image as input to generate a skin tone determination of the face as output.
4. The method according to claim 3, wherein: Executing the at least one machine learning model comprises: executing a first machine learning model using the facial images and the lighting environment images as input to generate lighting condition information as output for each of the facial images; and A second machine learning model is executed using the facial image and corresponding lighting condition information as input to generate a skin tone determination of the face as output.
5. The method according to claim 4, wherein The lighting condition information includes the color of the light source.
6. The method according to claim 4, wherein: The skin tone determinations each include a corresponding confidence level.
7. The method according to claim 6, further comprising: comparing the corresponding confidence level to a threshold confidence level; as well as In case the corresponding confidence level is less than the threshold confidence level, the skin tone determination is omitted from the combining step.
8. The method according to claim 1, wherein For each frame of the second video, extracting the facial image from the second video includes: measuring a brightness level of the frame; and In the event that the measured brightness level is less than a minimum threshold brightness level or greater than a maximum threshold brightness level, the frame is omitted from the extracted facial image.
9. The method according to claim 1, wherein The first video and the second video are recorded simultaneously.
10. A system comprising: A skin tone estimation unit, comprising a calculation circuit configured to: obtaining a first video recorded by a first camera having a first field of view in a lit environment; obtaining a second video recorded in the illuminated environment by a second camera having a second field of view different from the first field of view, wherein the second field of view is directed toward a face of a live subject; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images including the face of the live subject from the second video; processing the facial image and the lighting environment image to obtain a plurality of skin tone determinations of the face; combining the skin tone determinations for the faces to determine a combined skin tone determination for the face; and The combined skin tone of the face present is determined.
11. The system according to claim 10, wherein: The first camera is a rear-facing camera, and wherein the computing circuit is further configured to construct a panoramic image from the first video.
12. The system according to claim 10, wherein: The computing circuitry is further configured to execute at least one machine learning model using the facial image and the lighting environment image as input to generate the skin tone determination of the face as output.
13. The system according to claim 10, wherein: The computing circuit is further configured to: executing a first machine learning model using the facial images and the lighting environment images as input to generate lighting condition information as output for each of the facial images; as well as A second machine learning model is executed using the facial image and corresponding lighting condition information as input to generate a skin tone determination of the face as output.
14. The system according to claim 13, wherein: The lighting condition information includes the color of the light source.
15. The system according to claim 13, wherein: The skin tone determinations each include a corresponding confidence level.
16. The system according to claim 15, wherein: The computational circuitry is further configured to, for each of the skin tone determinations, determine a combined skin tone determination by combining the skin tone determinations for the face in the following manner: comparing the corresponding confidence level to a threshold confidence level; as well as In case the corresponding confidence level is less than the threshold confidence level, the skin tone determination is omitted from the combining step.
17. The system according to claim 10, wherein: The computing circuit is further configured to: for each frame of the second video, measuring a brightness level of the frame; and In case the measured brightness level is less than a minimum threshold brightness level or greater than a maximum threshold brightness level, the frame is omitted from the extracted facial image.
18. A non-transitory computer-readable medium having computer-executable instructions stored thereon, the computer-executable instructions, in response to being executed by one or more processors of a computer system, causing the computer system to perform actions, the actions comprising: presenting, on the mobile computing device, a user interface instructing the user to capture a panoramic image in the illuminated environment; In response to user input via the user interface, recording a first video with a rear-facing camera of the mobile computing device in the lit environment and recording a second video with a front-facing camera of the mobile computing device in the lit environment; extracting a plurality of lighting environment images from the first video; extracting a plurality of facial images including the user's face from the second video; processing the facial image and the lighting environment image to obtain a plurality of skin tone determinations of the face; and The multiple skin tone determinations for the face are combined to determine a combined skin tone determination for the face.
19. The non-transitory computer-readable medium of claim 18, wherein: Processing the facial image and the lighting environment image includes: executing, by the computing device, a first machine learning model using the facial images and the lighting environment images as input to generate lighting condition information as output for each of the facial images; and A second machine learning model is executed by the computing device using the facial image and corresponding lighting condition information as input to generate a skin tone determination of the face as output.
20. The non-transitory computer-readable medium of claim 19, wherein: The lighting condition information includes the color of the light source.