Improved face quality of captured images
By generating heatmaps and facial quality scores, and utilizing machine learning models to select high-quality facial images from image sequences, this solves the problem of difficulty in selecting overall facial quality in existing technologies, and achieves automated and efficient image capture.
Patent Information
- Application Number
- CN202080039744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2020-05-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Existing technologies struggle to select high-quality images based on overall facial quality, especially in image sequences, resulting in images of interest to the user not being captured or selected.
By generating heatmaps and facial quality scores, machine learning models are used to evaluate facial location and quality in image sequences, and images with peak facial quality scores are selected.
It improves the accuracy of image selection, ensures that facial images of greater interest to the user are captured or stored, and enables automatic image capture without user intervention.
Smart Images

Figure CN113906437B_ABST
Abstract
Description
Background Art
[0001] The subject matter disclosed herein relates to the field of digital imaging and is not limited to techniques for improving facial quality of captured images.
[0002] Digital imaging systems (such as video or still imaging cameras) are capable of capturing a large number of images in a relatively short period of time. Cameras are increasingly capable of capturing dozens or even hundreds of images per second. Image capture can also occur before or after other user interactions. For example, an image can be captured while the camera is active but before the capture button is pressed (or just released) to compensate for the user pressing the capture button too late (or too quickly).
[0003] In many cases, a user may only want to retain a single image or a relatively small subset of these images. Existing techniques for selecting images from an image sequence (such as a video clip or image burst) include the use of face detection, expression detection, and / or motion detection. However, such techniques may not be optimally suited for capturing a full range of images with high-quality representations of human faces. For example, previous implementations have attempted to use face and expression detection to locate faces and then detect smiles or other similar expressions on these faces in order to capture or select images. However, smiling faces are not sufficient to define the range of facial images that a user may want to capture. A user may be less interested in an image of a person smiling than, for example, a person looking a certain way, having a certain expression, or having a face presented (or posed) in a visually pleasing manner. Therefore, techniques for detecting and capturing images based on overall facial quality or "picture value" are desired. Summary of the Invention
[0004] The present disclosure generally relates to the field of image processing. More specifically, but not by way of limitation, aspects of the present disclosure relate to a computer-implemented method for image processing. In some embodiments, the method includes: obtaining an image sequence; detecting a first face in one or more images in the image sequence; determining a first position of the detected first face in each of the one or more images in the image sequence having the detected first face; generating a heat map based on the first position of the detected first face in each of the images in the image sequence; determining a facial quality score of the detected first face for each of the one or more images in the image sequence having the detected first face; determining a peak facial quality score of the detected first face based at least in part on the facial quality score and the generated heat map; and selecting a first image in the image sequence corresponding to the peak facial quality score of the detected first face.
[0005] Aspects of the present disclosure relate to selecting images based on the relative position and quality (e.g., in terms of "picture value") of faces within images in an image sequence. The relative position of faces can be determined based on a temporal heat map that accumulates heat map values for one or more faces across the image sequence. This accumulation in the form of a heat map can be used to help de-emphasize transient faces, i.e., faces that are absent from the image sequence for relatively long periods of time, which may be less likely to be the desired subject of the image, thereby allowing selection of images that are more likely to be desired by the user, i.e., photos that include faces that are present in a larger portion of the image sequence.
[0006] As described above, artificial intelligence (AI) techniques (such as machine learning or deep learning) can be used to evaluate the "picture quality" of faces detected in images in an image sequence to generate a face quality score. Heat map values can be used to further weight or scale the face quality score, for example, based on the temporal prevalence of a given face across the image sequence. Facial recognition can also be used to further weight or scale the face quality score, for example, based on how closely related the recognized face is to the image owner.
[0007] Images from the image sequence can then be selected based on the resulting facial quality scores (i.e., including any desired weighting scheme). The image selection methods disclosed herein can also be used to implement a "shutterless" mode for an imaging device, for example, allowing the imaging device to determine the best moment to capture an image automatically (i.e., without user intervention), such as within an open or predetermined time interval. Image selection can be applied to single-person images or multi-person images. In some cases, an intent classifier can be used to rate images based on the extent to which the images are likely to reflect "picture-worthy" versions of single-person images and / or multi-person group images.
[0008] In other embodiments, each of the above methods and variations thereof can be implemented as a series of computer-executable instructions. Such instructions can use any one or more convenient programming languages. Such instructions can be collected into an engine and / or program and can be stored in any medium readable and executable by a computer system or other programmable control device. In other embodiments, such instructions can be implemented by an electronic device (e.g., an image capture device), which includes a memory, one or more image capture devices, and one or more processors operably coupled to the memory, wherein the one or more processors are configured to execute instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 An exemplary sequence of portrait images according to aspects of the present disclosure is shown.
[0010] Figure 2 An exemplary visualization of a heat map of an image sequence according to aspects of the present disclosure is shown.
[0011] Figures 3A to 3B is an exemplary graph of facial quality scores according to aspects of the present disclosure.
[0012] Figure 4 is a flow chart illustrating a technique for improving facial quality of a captured image 400 in accordance with aspects of the present disclosure.
[0013] Figure 5 is a block diagram of an image selection module according to aspects of the present disclosure.
[0014] Figure 6 A functional block diagram of a programmable electronic device according to one embodiment is shown. DETAILED DESCRIPTION
[0015] The present disclosure relates to systems, methods, and computer-readable media for improving the operation of digital imaging systems. More particularly, aspects of the present disclosure relate to improving the selection of images in an image sequence based on the relative position and quality of faces within the images of the image sequence.
[0016] Generally speaking, selecting images from an image sequence can be used for both post-capture processing and processing of images observed by an imaging device. For example, this processing can help identify the best or most visually appealing image or set of images in an image sequence. In the post-capture scenario, an image sequence can be captured by an imaging device, the resulting images processed, and images selected from the image sequence. These selected images can be presented to the user for indexing, creating summaries, thumbnails, slideshows, selecting favorite images from the image sequence, and so on. In the case of processing images being observed, the image sequence can be processed to select images to be stored (e.g., images that may be of interest to the user) without requiring user interaction, such as by pressing a shutter or capture button. In some cases, selecting images to be stored can select when to begin storing a set of images or a video sequence. For example, an image sequence can be processed as it is observed by an imaging device, and image sets stored based on the selected images. These image sets can be based on the selected images, and the image sets can begin at a point in time before, at, or after the selected images.
[0017] For a particular image sequence, a scene classifier can also be used to determine whether the image sequence is a portrait image sequence, as opposed to, for example, an action image sequence capturing an action scene, a landscape image sequence, or some other related image sequence. For portrait images, rather than examining specific predefined image attributes such as contrast, color, exposure, composition, the presence of a smile, etc., portrait images can be captured based on an overall facial quality score determined by an artificial intelligence system, such as a machine learning or deep learning model trained to identify high-quality or visually pleasing images containing human faces, which takes into account the context in which the detected faces appear. For example, a user taking a selfie with a particular background can position themselves within the frame and stabilize the camera. Another person can then pass through the frame. Given such an image sequence, the user may be primarily interested in people who appear consistently within the image sequence. Prioritizing facial quality for people who appear consistently within the image sequence (and, in some cases, may be specifically known to the user (e.g., as determined via techniques such as facial recognition)) can be used to help select images from the image sequence that the user will perceive as more relevant and higher quality.
[0018] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the disclosed concepts. As part of this description, some of the figures in the drawings of the present disclosure represent structures and devices in block diagram form to avoid obscuring the novel aspects of the disclosed concepts. For the sake of clarity, not all features of the actual specific implementation are described. In addition, the language used in this disclosure is selected primarily for readability and instructional purposes, and may not be selected to describe or define the claimed subject matter, so that it may be necessary to rely on the claims to determine such claimed subject matter. Reference to "one embodiment" or "an embodiment" in this disclosure means that the specific features, structures or characteristics described in conjunction with the embodiment are included in at least one embodiment of the disclosed subject matter, and multiple references to "one embodiment" or "an embodiment" should not be understood as necessarily all referring to the same embodiment.
[0019] It will be appreciated that in the development of any actual implementation (as in any software and / or hardware development project), many decisions must be made to achieve the developer's specific goals (e.g., to meet system and business-related constraints), and that these goals may vary from one implementation to another. It will also be appreciated that such a development effort may be complex and time-consuming, but nonetheless would be a routine undertaking for a person of ordinary skill having the benefit of this disclosure and designing and implementing graphics processor interface software or a graphics processing system.
[0020] A portrait image sequence contains one or more faces within the images in the image sequence. Figure 1, which illustrates an exemplary portrait image sequence 100 according to aspects of the present disclosure. The exemplary portrait image sequence 100 includes four images 102A-102D (collectively, 102). While the portrait image sequence 100 includes four images, it should be understood that other portrait image sequences may include at least two images. In this example, the portrait image sequence 100 includes a desired subject 104, which remains largely stationary relative to the image frames of the portrait image sequence 100. The portrait image sequence 100 also includes a second person 106 entering the frame in image 102A, moving through the frame in image 102B, and leaving the frame in image 102C. The portrait image sequence 100 also includes non-human objects, such as shrubs 108. In some cases, the portrait image sequence may include a real-time or near real-time sequence of images from a camera, and selected images from the image sequence may be stored and used to represent the image sequence as thumbnails or images displayed to a user. Unselected images may be discarded or stored. For example, all images in the image sequence may be stored, or selected images may be stored. In some cases, multiple images or multiple seconds of images before, after, or before and after the selected image may be stored. In some cases, the image sequence may begin after a certain event is detected by the imaging device. For example, a shutter release may be triggered on the imaging device or the imaging device may be switched to a particular mode. In some cases, input from another sensor or process of the imaging device may be used to begin the image sequence. For example, the image sequence may begin based on input from a gyroscope or accelerometer indicating that the imaging device has been stabilized, held in a particular pose, etc. In accordance with aspects of the present disclosure, facial detection and / or recognition may be used to track and / or identify faces present in the captured scene.
[0021] Generally, at least one of the faces in the image sequence is a target subject of the portrait image sequence. Face detection and tracking can be performed on the images of the image sequence to locate and track the faces in the images. Faces can be detected and tracked in the images via any known technique, such as model-based matching, low-level feature classifiers, deep learning, etc. Generally, faces can be detected and positional information from the detected faces can be used to generate a heat map of the locations of the detected faces across the image sequence. In some cases, the detected faces can be tracked such that the face of a particular person is tracked from one frame to the next throughout the image sequence. In some cases, the face can be identified or recognized, for example, by matching the detected face to the face of a person known to the user and tracking it across the frames of the image sequence.
[0022] Figure 2 An exemplary visualization of a heat map 200 of an image sequence according to aspects of the present disclosure is shown. It should be understood that Figure 2The heat map 200 is a visual representation of an object (such as a data object) in memory and may not be displayed or accessed by a user or program. Generally speaking, the heat map plots the locations of detected faces within the frames of an image sequence into a single plot. In the case where the image sequence represents real-time images from an imaging device, the heat map plots may be accumulated as the images of the image sequence arrive. In some cases, objects in the images that are not detected as faces may be ignored. In the heat map 200, the more frequently a particular detected face appears in an area of an image frame in the image sequence, the darker the area of the heat map. In this example, Figure 1 The first region 202 of the heat map 200 corresponding to the expected object 104 appears to be larger than Figure 1 The second region 204 of the heat map 200 corresponding to the second person 106 is darker because the object 104 is expected to appear in the same region of the frame on more images of the image sequence than the second person 106 .
[0023] Generally speaking, subjects in a portrait image sequence tend not to exhibit a significant amount of expected motion relative to the frame across the images in the image sequence. That is, outside of a portrait-type photograph that initially stabilizes or frames the intended subject, the intended subject tends to remain stationary and generally does not move extensively around the image frames in the image sequence. Therefore, detected faces can be evaluated based on their relative motion within the frames as indicated by the heat map. For example, a detected face that remains relatively consistently positioned across the frames (corresponding to the darker areas of heat map 200) may be more likely to be associated with the user and be valued higher than a detected face that moves around the frames in the image sequence. In some cases, the heat map value may also be adjusted as a function of the size of the detected face, where larger detected faces may generally be valued higher. This size function may be based on absolute size or relative size compared to other detected faces. This heat map evaluation helps distinguish the intended subject of the image sequence from, for example, passersby or people in the background.
[0024] According to aspects of the present disclosure, heatmap values may be used in conjunction with facial quality scores for detected faces. For example, heatmap values may be used to determine whether a facial quality score is obtained for a particular detected face. As another example, heatmap values may be used as weights to be applied to the facial quality score for the corresponding detected face. These heatmap values for a detected face may be absolute values or assigned relative to another detected face. In such cases, a threshold or relative heatmap value may be set such that a detected face should appear in a relatively similar position in a minimum number of frames before a facial quality score is obtained for that detected face. For a position of a detected face in an image, the heatmap value associated with the position of the detected face is incremented as the detected face remains at that position. In some cases, heatmap values may be associated with detected and tracked faces, and a separate heatmap may be created for each tracked face.
[0025] The heatmap values are aggregated across the images in the image sequence. If a detected face remains in a particular position sufficient for the heatmap value to exceed a threshold heatmap value, a facial quality score for the detected face can be obtained. In some cases, the threshold value can be relative to the heatmap values associated with other detected faces in the image sequence. The heatmap value threshold can help enhance performance by filtering out faces that are unlikely to be objects of the image sequence (e.g., because they appear only briefly within the duration of the image sequence), so that the facial quality analysis can focus on detected faces that are more likely to be associated with the user. In other cases, other criteria can also be used in conjunction with the heatmap to determine whether to perform facial quality analysis on a particular face, such as, for example, facial size, a confidence score determined by a face detector, a confidence score from a face recognizer as to whether the face is likely known to the user, etc.
[0026] In some cases, heatmap values may also be evaluated as part of determining the relevance of a face quality score. For example, heatmap values may be used as weights to be applied to the face quality scores of corresponding detected faces. As a more detailed example, the longer a detected face remains in a particular position in a frame, the higher the heatmap value (and therefore the weight) associated with that detected face may be. Return to Combining Figure 2 In the example discussed, the first region 202 of the heat map 200 appears in the same region of the frame across each image and may be associated with a maximum (or near-maximum) heat map value, such as a weight of "1" on a normalized scale of 0 to 1. Another face that moves around the frame or is present for a shorter duration may then be associated with a lower weight. Figure 2 In the example discussed, the second region 204 of each image, which is shifted relative to the frame, may be associated with a lower heatmap value than the first region 202, such as 0.1 (on the aforementioned normalized scale of 0 to 1). The facial quality score of each detected face may be weighted, for example, by multiplying the facial quality score of each detected face in a given image from the image sequence by its corresponding heatmap value. In some cases, the facial quality score may be adjusted based on the heatmap value of the detected face associated with at least a threshold heatmap value. Detected faces associated with heatmap values that do not meet the threshold heatmap value may be ignored.
[0027] According to aspects of the present disclosure, a detected face in an image sequence can be assigned a facial quality score for each frame in which the face was detected. In some cases, a machine learning model (MLM) classifier can be used to assign facial quality scores. It should be understood that MLM can encompass deep learning and / or neural network systems. Rather than training an MLM classifier based on specific image qualities (such as color, exposure, framing, etc.), the MLM classifier can be trained holistically based on the overall "image value" of a face. For example, an MLM classifier can be trained on pairs of images of a person, with annotations indicating which image in the pair should be retained, for example, based on which image is more visually pleasing to the annotator overall. Generally speaking, a holistic assessment of a face focuses on whether the face is visually interesting or the overall aesthetics of the face, rather than attempting to assess any specific aspect or set of aspects of the face, such as lighting, framing, focus, presence of a smile, gaze direction, etc., and then combining these aspects into a score. Other training methods can also be used, but these should measure facial quality holistically, rather than measuring individual aspects of facial quality, such as using detectors to detect lighting, smiles, wrinkles, emotion, gaze, attention, etc. Generally speaking, for a given input, such as a detected face in an image of an image sequence, an MLM classifier machine learning model outputs a prediction as a confidence score that indicates the machine learning model's confidence that the input corresponds to the specific set for which the classifier was trained. For example, given an MLM classifier trained to detect the overall picture value of a detected face, the MLM classifier can output a score that indicates the MLM classifier's confidence in the picture value of the detected face. This score can be used as a facial quality score for the corresponding detected face in a given image.
[0028] In some cases, the MLM can be updated based on images selected by the user. In some cases, these updates can be per-user or per-device updates. For example, a user can choose to retain images other than those selected. These images can be used for additional training of the machine learning model, either through the imaging device or through a separate electronic device (such as a server). In some cases, these images can be tagged and uploaded to a repository for training and creating updates to the machine learning model. In some cases, updates can also be made for groups of users based on the additional training. These updates can then be applied to the machine learning model on the imaging device or other imaging devices.
[0029] Once facial quality scores for detected faces are generated for the images of the image sequence, the facial quality scores for each image may be plotted. Figures 3A to 3B is an exemplary graph of facial quality scores 300 according to aspects of the present disclosure. The facial quality scores of detected faces may be sorted for a sequence of images to select images having the relatively highest corresponding facial quality scores. For example, Figure 3A Graph 302 plots detected faces on the Y-axis and facial quality scores 304 for image i in an image sequence, with the facial quality scores increasing in numerical order from 1 to 10 on the x-axis. In some cases, the facial quality scores associated with detected faces can be adjusted based on the heat map values associated with the detected faces. For example, the heat map values associated with the respective detected faces can be used to adjust the facial quality scores of the respective detected faces prior to image selection (e.g., via a weighting process). This adjustment of the facial quality scores can be performed for each detected face in each image of the image sequence.
[0030] Images can be selected based on, for example, a peak in facial quality scores within a sliding window. The sliding window typically starts with the first image and finds a peak in facial quality scores across N images. The sliding window then shifts to the next image, finds a peak, and repeats until the process stops. The sliding window can find a peak across a number of images and select the highest peak within the sliding window.
[0031] As a more detailed example, the sliding window of graph 302 may begin by searching for a peak between images 1 and 2, and if none is found, it may be moved to include image 3, then image 4, and so on. After the sliding window is moved to include image 5, peak 306 may be detected, and image 4 may be selected from graph 302. The sliding window then moves to the next image, and once the sliding window reaches a position that includes image 6, image 1 is no longer included in the sliding window. Since image 4 has already been selected, no further images are selected. As the sliding window continues to move, after the sliding window moves to include image 8 and peak 308 is detected, image 7 may be selected from graph 302. The sliding window may be moved, for example, to consider images as they are received at the imaging device. In some cases where a sliding window is not utilized, the selected image may be the highest-scoring image in the image sequence, in which case image 7 may be selected from graph 302. In other cases, peak detection may be applied after the entire image sequence is obtained. In some cases, the image corresponding to the detected peak may be selected. In some cases, an image with the lowest facial quality score (eg, a facial quality score identified based on a detected valley in facial quality scores) may also be selected as, for example, a candidate image or compared to other selected images.
[0032] Generally speaking, images containing faces of known people may be more important than images containing unknown faces. In some cases, a facial recognition system may be used in place of or in conjunction with a facial quality score. A facial recognition system can identify faces known to the user and help prioritize images containing faces known to the user. The facial recognition system can be configured to recognize faces known to the user. For example, the facial recognition system can be an MLM classifier trained on labeled images containing faces of people known to the user. These images may be provided by the user or associated with the user, such as through a metadata network associated with the user's digital asset (DA) library. The user's DA library may include media items in a collection associated with the user, such as photos, videos, image sequences, etc. In some cases, the facial recognition system can be a separate MLM classifier, or alternatively, the facial quality MLM classifier can be configured to generate both a facial quality score and a second score indicating whether a particular detected face is associated with or known to the user. This second score can be used to adjust the facial quality score in a manner similar to a heatmap value. In some cases, the facial quality MLM classifier can be configured to perform this adjustment as part of generating the facial quality score.
[0033] Figure 3BGraph 310 plots facial quality scores for a first tracked detected face 312 and a second tracked detected face 314 on the y-axis, and for image i in the image sequence, with the facial quality scores increasing numerically from 1 to 10 on the x-axis. If an image includes more than one tracked detected face, a facial quality score may be determined for each tracked detected face, and an image may be selected based on, for example, detected peaks in the facial quality scores of the faces. For example, the selected image may include facial quality peaks for multiple tracked detected faces. In this case, image 4 may be selected based on the coincidence of peak 316 in the facial quality score for first tracked detected face 312 and peak 318 in the facial quality score for second tracked detected face 314. In some cases, images 6 and 7 may also be selected based on the temporal proximity of peak 322 for first tracked detected face 312 and peak 320 for second tracked detected face 314. Additionally, images may be selected that include facial quality peaks for at least one tracked detected face. For example, image 6 may be selected based on peak 320 in the facial quality score for second tracked detected face 314, and image 7 may also be selected based on peak 322 in the facial quality score for first tracked detected face 312. In some cases, facial recognition may be performed on multiple tracked detected faces, and the selected image may be associated with the identified person based in part on the peak in the facial quality score of the identified person. In some cases, image selection may also be based on aggregated facial quality scores. For example, the facial quality scores of the first tracked detected face and the second tracked detected face may be summed, and peak detection may be performed on the aggregated facial quality score. In some cases, this aggregation of facial quality scores may be performed in addition to and after individually weighting the facial quality scores of the tracked detected faces.
[0034] As discussed above in connection with a single detected face, image selection with multiple detected faces can be performed based on a sliding window. According to certain aspects, larger groups of people, such as 3 or more people, may result in too many images being selected based on peaks in face quality scores. In some cases, face quality scores can be enhanced, for example, by one or more intent classifiers. These intent classifiers can be MLM classifiers trained to detect group intent. For example, an intent classifier can be trained to detect gaze direction to help select images in which a group of faces is looking in the same direction. This direction may or may not be at the image capture device, for example, where everyone is looking at a particular object. Similarly, an intent classifier can be trained to detect actions, such as pointing, jumping, panting, etc.
[0035] Figure 44 is a flow chart illustrating a technique for improving facial quality in a captured image 400 according to aspects of the present disclosure. At step 402, the technique begins by first obtaining an image sequence. Generally, the image sequence may be obtained from an imaging device in the form of a burst of still images, video, slow-motion images, and the like. The images in the image sequence are typically ordered in time. At step 404, the technique includes detecting a first face in one or more images in the image sequence. For example, any known facial detection technique may be used to detect faces in some images in the image sequence. At step 406, the technique includes determining a first position of the detected first face in each of one or more images in the image sequence having the detected first face. For example, once a face is detected, the position of the detected face may be determined. The detected face is tracked across the images in which the face is detected. At step 408, the technique includes generating a heat map based on the first position of the detected first face in each image in the image sequence. For example, the position information of the detected faces may be accumulated for the image sequence and used to generate a heat map that describes where the detected faces are located and how often the detected faces appear at specific positions in frames throughout the image sequence. At step 410, the technique includes determining a facial quality score for the detected first face for each of one or more images in an image sequence having a detected first face. For example, an MLM may be used to determine the facial quality scores for the detected faces. The MLM may be trained to generate facial quality scores based on the overall picture quality of the face, rather than on specific features of the face or picture quality. At step 412, the technique includes weighting the facial quality scores for each of the one or more images relative to the detected first face. In some cases, the weights may be based on heatmap values determined based on the location of the detected first face in each image. For example, a first detected face that appears in approximately the same location within the image (i.e., relative to the overall extent of the image) across a greater number of images in the image sequence may be associated with a higher heatmap value than a second detected face that has moved around the frames or appears only occasionally in the image sequence. The facial quality score associated with each respective detected face may be adjusted, for example, via a weighting process that includes multiplying the respective facial quality score by the heatmap value for each image in the image sequence in which the respective detected face appears. In some cases, the facial quality value may alternatively (or additionally) be weighted based on the facial recognition module's determination of whether the detected face is a person known (or otherwise recognized) by the user. At step 414, the technique includes determining a peak facial quality score for the first detected face. For example, the facial quality scores of the detected faces may be aggregated, and a peak within the aggregated facial quality score may be detected. Peak detection may be performed within a sliding window of N images.At step 416, the technique includes selecting a first image in the sequence of images that corresponds to a peak face quality score for the first detected face based at least in part on the face quality scores and the generated heat map. For example, an image corresponding to a detected peak in the face quality score of the detected face may be selected. The face quality score of each image may be adjusted based on the heat map value. The selected image may be saved or used as an image for display to a user. In some cases, the selected image may be saved as the captured image, such as when capturing images in a shutterless mode, in which the image capture device determines when to capture the image, rather than the user of the image capture device. It may be noted that multiple faces may be tracked for a sequence of images, and that this may be done, for example, by detecting a second face in one or more images in the sequence of images, selecting more than a single image from the sequence of images, determining a second position of the detected second face in each of the one or more images having the detected second face, generating a heat map based on the second position of the detected second face in each image in the sequence of images, determining a second facial quality score of the detected second face for each of the one or more images in the sequence of images having the detected second face, determining a second peak facial quality score for the detected second face based at least in part on the facial quality scores and the generated heat map, and selecting a second image corresponding to the second peak facial quality score of the detected second face.
[0036] Figure 5 is a block diagram illustrating an image selection module 500 according to aspects of the present disclosure. Face detection may be performed by a face detection / recognition module 502 for each image in an image sequence to detect and / or identify faces visible within the image. The face detection / recognition module 502 may detect, track, and provide location information for faces within the image. In some cases, the face detection / recognition module 502 may be configured to identify detected faces based on a metadata network associated with, for example, a user's DA library. The image selection module 500 may also include a heat map module 504. This heat map indicates the frequency with which faces within the image sequence are located in different parts of the image frame. The detected faces in each frame may also be evaluated and scored by a face scoring module 506. The face scoring module may include one or more MLM classifiers, such as the aforementioned MLM classifiers trained based on the overall "picture value" of faces. The resulting face score for each face may be tracked across the image sequence by a score tracking module 508 and used in conjunction with the heat map to select one or more images from the image sequence.
[0037] Exemplary Hardware and Software
[0038] Now see Figure 6, shows a simplified functional block diagram of an exemplary programmable electronic device 600 according to one embodiment. The electronic device 600 can be, for example, a mobile phone, a personal media device, a portable camera, or a tablet, laptop, or desktop computer system. As shown, the electronic device 600 may include a processor 605, a display 610, a user interface 615, graphics hardware 620, device sensors 625 (e.g., a proximity sensor / ambient light sensor, an accelerometer, and / or a gyrometer), a microphone 630, an audio codec 635, a speaker 640, a communication circuit 645, an imaging device 650 (e.g., which may include multiple camera units / optical image sensors with different characteristics or capabilities (e.g., high dynamic range (HDR), optical image stabilization (OIS) system, optical zoom and digital zoom, etc.), a video codec 655, a memory 660, a storage device 665, and a communication bus 670.
[0039] The processor 605 may execute instructions necessary to implement or control the operation of various functions performed by the electronic device 600 (e.g., such as selecting an image from a sequence of images, according to various embodiments described herein). The processor 605 may, for example, drive the display 610 and may receive user input from a user interface 615. The user interface 615 may take various forms, such as buttons, a keypad, a dial, a click wheel, a keyboard, a display screen, and / or a touch screen. The user interface 615 may, for example, be a conduit through which a user can view a captured video stream and / or indicate a specific image that the user wishes to capture (e.g., by clicking a physical or virtual button while the desired image is being displayed on the device's display screen). In one embodiment, the display 610 may display a captured video stream as the processor 605 and / or graphics hardware 620 and / or imaging circuitry simultaneously generate and store the video stream in memory 660 and / or storage 665. The processor 605 may be a system-on-chip, such as those found in mobile devices, and may include one or more dedicated graphics processing units (GPUs). The processor 605 may be based on a reduced instruction set computer (RISC) or complex instruction set computer (CISC) architecture, or any other suitable architecture, and may include one or more processing cores. The graphics hardware 620 may be specialized computing hardware for processing graphics and / or assisting the processor 605 in performing computing tasks. In one embodiment, the graphics hardware 620 may include one or more programmable graphics processing units (GPUs).
[0040] For example, according to the present disclosure, imaging device 650 may include one or more camera units configured to capture images that, for example, may be processed to generate images with depth / disparity information for such captured images. Output from imaging device 650 may be processed, at least in part, by a video codec 655 and / or processor 605 and / or graphics hardware 620, and / or a dedicated image processing unit or image signal processor incorporated within imaging device 650. Such captured images may be stored in memory 660 and / or storage 665. Memory 660 may include one or more different types of media used by processor 605, graphics hardware 620, and imaging device 650 to perform device functions. For example, memory 660 may include a memory cache, read-only memory (ROM), and / or random access memory (RAM). Storage 665 may store media (e.g., audio files, image files, and video files), computer program instructions or software, preference information, device profile information, and any other suitable data. The storage device 665 may include one or more non-transitory storage media, including, for example, magnetic disks (fixed hard disks, floppy disks, and removable disks) and tapes, optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such as electrically programmable read-only memories (EPROMs) and electrically erasable programmable read-only memories (EEPROMs). The memory 660 and the storage device 665 may be used to hold computer program instructions or codes organized into one or more modules and written in any desired computer programming language. For example, when executed by the processor 605, such computer program code may implement one or more of the methods or processes described herein.
[0041] It should be understood that the above description is intended to be illustrative and not restrictive. For example, image sequences may be obtained from a variety of imaging devices, including but not limited to still imaging devices, video devices, and invisible light imaging devices. It should be understood that various techniques may be used to detect and locate objects, determine object trajectories, and score the determined trajectories. The determination and aggregation of trajectory scores may also be tailored to address specific scenarios.
[0042] Many other embodiments will be apparent to those of skill in the art upon reviewing the above description.The scope of the invention should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0043] As described above, one aspect of the present technology is the collection and use of data available from various sources to improve the quality of the face captured and the selection of representative images. The present disclosure contemplates that, in some instances, such collected data may include personal information data that uniquely identifies or can be used to contact or locate a specific person. Such personal information data may include facial images, demographic data, location-based data, phone numbers, email addresses, Twitter IDs, home addresses, data or records related to the user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.
[0044] This disclosure recognizes that the use of such personal information data within the present technology can be used to benefit users. For example, personal information data can be used to capture or select images that are of greater interest to the user. Furthermore, this disclosure contemplates other uses of personal information data that can benefit users. For example, health and fitness data can be used to provide insights into the user's overall health or as positive feedback to individuals using technology to pursue health goals.
[0045] This disclosure contemplates that entities responsible for collecting, analyzing, disclosing, transmitting, storing, or otherwise using such personal information will adhere to established privacy policies and / or practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information. Such policies should be easily accessible to users and updated as the collection and / or use of data changes. Personal information collected from users should be used for the entity's legitimate and reasonable purposes and not shared or sold beyond those legitimate uses. Furthermore, such collection / sharing should be conducted with the user's informed consent. Furthermore, such entities should consider taking any necessary steps to safeguard and secure access to such personal information and ensure that others with access to the personal information adhere to their privacy policies and procedures. Furthermore, such entities may subject themselves to third-party assessments to demonstrate compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be tailored to the specific type of personal information collected and / or accessed and to applicable laws and standards, including jurisdictional considerations. For example, in the United States, the collection or access of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); whereas health data in other countries may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy practices should be maintained for different types of personal data in each country.
[0046] Regardless of the foregoing, the present disclosure also contemplates implementation scenarios in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates providing hardware elements and / or software elements to prevent or block access to such personal information data. For example, in the case of facial recognition services or access to a user's DA library, the technology of the present invention can be configured to allow a user to choose to "opt in" or "opt out" to participate in the collection of personal information data at any time during or after registration for the service. For another example, a user can choose not to access facial recognition services or not allow access to the user's DA library. In such cases, the present disclosure contemplates providing certain services, such as facial tracking, that can be used without utilizing services or permissions that have been opted out. In addition to providing "opt-in" and "opt-out" options, the present disclosure contemplates providing notifications related to access or use of personal information. For example, a user can be notified that their personal information data will be accessed when downloading an application, and then reminded again just before the personal information data is accessed by the application.
[0047] Furthermore, it is an object of the present disclosure that personal information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use. Risk can be minimized by limiting data collection and deleting data once it is no longer needed. In addition, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data across users), and / or other methods, where appropriate.
[0048] Thus, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that various embodiments may be implemented without access to such personal information data. That is, various embodiments of the present technology will not be unable to function properly due to the lack of all or a portion of such personal information data. For example, an image may be selected on a tracked face and based on non-personal information data or a minimal amount of personal information (such as content requested by a device associated with the user, other non-personal information available to the imaging service, or publicly available information).
Claims
1. A computer-implemented image selection method, the method comprising: Obtain an image sequence; detecting a first face in one or more images in the sequence of images; determining a first position of the detected first face in each of the one or more images in the sequence of images having the detected first face; generating a heat map based on the first position of the detected first face in each of the images in the sequence of images; determining a face quality score for the detected first face for each of the one or more images in the sequence of images having the detected first face; determining a peak face quality score for the detected first face based at least in part on the face quality score and the generated heat map; as well as A first image in the sequence of images corresponding to the peak face quality score for the detected first face is selected.
2. The method according to claim 1, further comprising: detecting a second face in one or more images in the sequence of images; determining a second position of the detected second face in each of the one or more images in the sequence of images having the detected second face, wherein the heat map is further based on the second position of the detected second face, and wherein the heat map indicates that the second position of the detected second face has changed more than the first position of the detected first face; as well as The detected second face is filtered based on the heat map.
3. The method of claim 2 , wherein the heat map comprises heat map values corresponding to the second position of the detected second face in the image sequence, and The method further comprises comparing the heat map value corresponding to the second position to a threshold heat map value. 4 . The method of claim 1 , wherein the peak facial quality score is determined based on a probability value output by a machine learning model.
5. The method of claim 4, wherein the machine learning model is configured to detect picture values of faces rather than facial objects. The method of claim 1 , wherein determining the peak facial quality score comprises determining whether the facial quality score reaches a peak within a sliding window. The method of claim 1 , wherein selecting the first image in the sequence of images comprises storing the first image.
8. A non-transitory program storage device comprising instructions stored thereon, the instructions causing one or more processors to: Obtain an image sequence; detecting a first face in one or more images in the sequence of images; determining a first position of the detected first face in each of the one or more images in the sequence of images having the detected first face; generating a heat map based on the first position of the detected first face in each of the images in the sequence of images; determining a face quality score for the detected first face for each of the one or more images in the sequence of images having the detected first face; determining a peak face quality score for the detected first face based at least in part on the face quality score and the generated heat map; as well as A first image in the sequence of images corresponding to the peak face quality score for the detected first face is selected.
9. The non-transitory program storage device of claim 8, wherein the instructions for determining the selected set of images further cause the one or more processors to: detecting a second face in one or more images in the sequence of images; determining a second position of the detected second face in each of the one or more images in the sequence of images having the detected second face, wherein the heat map is further based on the second position of the detected second face, and wherein the heat map indicates that the second position of the detected second face has changed more than the first position of the detected first face; and The detected second face is filtered based on the heat map.
10. The non-transitory program storage device of claim 9, wherein the heat map comprises heat map values corresponding to the second position of the detected second face in the sequence of images, and Wherein the instructions for determining the selected image set further cause the one or more processors to compare the heat map value corresponding to the second location to a threshold heat map value.
11. The non-transitory program storage device of claim 8, wherein the peak facial quality score is determined based on a probability value output by a machine learning model.
12. The non-transitory program storage device of claim 11, wherein the machine learning model is configured to detect picture values of faces rather than facial objects. 13 . The non-transitory program storage device of claim 8 , wherein determining the peak facial quality score comprises determining whether the facial quality score peaks within a sliding window.
14. The non-transitory program storage device of claim 8, wherein selecting the first image in the sequence of images comprises storing the first image.
15. An electronic device comprising: Memory; one or more image capture devices; as well as One or more processors operably coupled to the memory, wherein the one or more processors are configured to execute instructions that cause the one or more processors to: Obtain an image sequence; detecting a first face in one or more images in the sequence of images; determining a first position of the detected first face in each of the one or more images in the sequence of images having the detected first face; generating a heat map based on the first position of the detected first face in each of the images in the sequence of images; determining, by a machine learning model, a facial quality score for the detected first face for each of the one or more images having the detected first face, wherein the machine learning model determines the facial quality score based on an overall assessment of facial quality of the detected first face; as well as A first image in the sequence of images is selected based on the generated heat map and the determined face quality score of the detected first face.
16. The apparatus of claim 15, wherein the one or more processors are configured to execute instructions that further cause the one or more processors to: detecting a second face in one or more images in the sequence of images; determining a second position of the detected second face in each of the one or more images in the sequence of images having the detected second face, wherein the heat map is further based on the second position of the detected second face, and wherein the heat map indicates that the second position of the detected second face has changed more than the first position of the detected first face; and The detected second face is filtered based on the heat map.
17. The apparatus of claim 15, wherein the heat map comprises heat map values corresponding to a second position of a detected second face in the sequence of images, and Wherein the one or more processors are configured to execute instructions that further cause the one or more processors to compare the heat map value corresponding to the second location to a threshold heat map value.
18. The apparatus of claim 15, wherein the facial quality score is determined based on a probability value output by the machine learning model.
19. The apparatus of claim 18, wherein the machine learning model is configured to detect picture values of faces rather than facial objects.
20. The apparatus of claim 15, wherein the instructions further cause the one or more processors to: A peak face quality score for the detected first face is determined based at least in part on the face quality score and the generated heat map, wherein determining the peak face quality score comprises determining whether the face quality score peaks within a sliding window.
Citation Information
Patent Citations
Image processing method, device and system, and storage medium
CN108875540A
Apparatus and method for optimization of ultrasound images
US20160242740A1