Aesthetic-based portrait image assessment
By automatically identifying regions of interest in images of people using computing devices and calculating total scores, the problem of traditional devices being unable to assess the aesthetic value of images is solved, achieving efficient assessment and resource saving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2019-07-08
- Publication Date
- 2026-04-24
AI Technical Summary
Traditional image acquisition equipment cannot effectively assess the aesthetic value of images, especially based on the subjective characteristics of people in the images, which makes manual editing time-consuming and detrimental to the original aesthetics.
The system automatically identifies regions of interest in images of people using computing devices, calculates predefined scores for each region based on training data, and generates a total score to evaluate the aesthetic value of the image.
It enables efficient evaluation of image aesthetic value, saves processing resources and storage space, improves image editing efficiency, and ensures high-quality storage.
Smart Images

Figure CN112424792B_ABST
Abstract
Description
[0001] Cross-citation of related applications
[0002] This application claims priority and interest to U.S. non-provisional patent application No. 16 / 034,693, filed July 13, 2018, entitled "Aesthetic-Based Portrait Image Evaluation," and to U.S. non-provisional patent application No. 16 / 131,681, filed September 14, 2018, both of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to image analysis. More specifically, this disclosure generally relates to image analysis on a device based on the aesthetic features of an image. Background Technology
[0004] Traditional devices, such as smartphones, mobile tablets, digital cameras, and camcorders, can be used to capture images and videos. These devices can alter objective features of an image, such as shadows, colors, brightness, and pixel texture. For example, users of these devices can manually edit objective features of an image using filters, which are typically used to change the appearance of an image or a portion of it. However, manually editing images is time-consuming and sometimes detracts from the original aesthetics or value of the image.
[0005] The aesthetic value or beauty of an image refers to the subjective reaction of a user viewing an image or video. Thus, the aesthetic value of an image can be based on both its objective and subjective characteristics. Traditional image acquisition equipment can only determine the objective characteristics of an image, such as the brightness, contrast, saturation, sharpness, hue, and color of pixels. However, the aesthetic value of an image can be based not only on its objective characteristics but also on its subjective characteristics or features related to the people depicted within the image. Summary of the Invention
[0006] According to one aspect of this disclosure, a method implemented by a computing device is provided. The method includes: the computing device determining a plurality of attributes, each attribute describing a region of interest corresponding to a body part of a person displayed in an image; the computing device determining a corresponding score for each of the plurality of attributes; and the computing device calculating a total score based on the corresponding scores of the plurality of attributes.
[0007] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that, in response to receiving a selection of the person displayed in the image, the person is determined, and the image displays the person and a plurality of other people.
[0008] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the person is determined based on the face of the person detected in the image.
[0009] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the corresponding score of each of the plurality of attributes is determined based on training data including a plurality of predefined scores of each of the plurality of attributes.
[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the training data includes multiple mappings, each mapping a predefined score from a plurality of predefined scores to a predefined attribute from a plurality of predefined attributes.
[0011] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the image is one of a plurality of images included in a video, and the method further includes: the computing device determining one or more images from the plurality of images including the person; the computing device combining one or more images from the plurality of images including the person to create a summary video of the person, wherein the plurality of images included in the summary video are selected based on the total score of each of the plurality of images, and the total score is calculated based on the attributes of the person.
[0012] Alternatively, in any of the above aspects, another implementation of the aspect specifies that the total score is calculated based on the person's general attributes and location attributes.
[0013] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that when the image displays the person and multiple other people, the method further includes: the computing device determining a score for the background of the image, wherein the background of the image includes the multiple other people, and further calculating the total score based on the score of the background of the image.
[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the method further includes: the computing device searching for regions of interest corresponding to different body parts of the person depicted in the image based on the probability that the body part is located at a certain position within the image.
[0015] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that searching the region of interest includes searching the region of interest based on training data, wherein the training data includes predefined anchor points pointing to specific parts or points in the image.
[0016] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that searching the region of interest includes searching for a region of interest corresponding to at least one of the person's eyes, the person's nose, or the person's mouth based on the facial position of the person in the image.
[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that multiple predefined scores of multiple attributes are stored respectively, and determining the corresponding score of each of the multiple attributes includes searching in the training data for a predefined score corresponding to the attribute determined for the region of interest.
[0018] Optionally, in any of the above aspects, another implementation of the aspect specifies that the plurality of attributes includes a plurality of location attributes that respectively describe location information corresponding to the region of interest.
[0019] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the plurality of attributes includes a plurality of general attributes of the person depicted in the image, the plurality of general attributes describing the overall quality of the person.
[0020] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that when the image depicts more than one person, the method further includes: the computing device determining a corresponding score for each of a plurality of group attributes, wherein the plurality of group attributes respectively describe at least one of the following: the relationship between the plurality of other people depicted in the image, the space between each of the plurality of other people depicted in the image, the pose of one or more of the plurality of other people depicted in the image, or the arrangement of the plurality of other people depicted in the image, and further calculating the total score based on the corresponding scores of the plurality of group attributes.
[0021] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that weights are associated with each of the plurality of attributes, wherein the weight of the corresponding attribute is applied to the corresponding score of the corresponding attribute to create a weighted score for the corresponding attribute, and the total score is calculated based on the set of each of the weighted scores for each of the corresponding attributes.
[0022] According to one aspect of this disclosure, a method implemented by a computing device is provided. The method includes: the computing device determining one or more images from a plurality of images comprising a person in a video; the computing device combining the one or more images from the plurality of images comprising the person to create a summary video of the person.
[0023] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the method further includes the computing device receiving a selection of the person.
[0024] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the person is determined in response to detecting the face of the person in one or more of the plurality of images.
[0025] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that one or more images of the plurality of images including the person are determined based on the total score of each of the plurality of images.
[0026] Alternatively, in any of the above aspects, another implementation of the aspect specifies that the total score is calculated based on multiple attributes, each attribute describing a region of interest corresponding to a body part of the person.
[0027] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the multiple attributes of the person include multiple location attributes that respectively describe location information corresponding to the region of interest.
[0028] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the multiple attributes of the person include multiple general attributes of the person, wherein the multiple general attributes describe the overall quality of the person.
[0029] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the plurality of images including the person are determined based on a first plurality of attributes and a second plurality of attributes, wherein each of the first plurality of attributes describes a region of interest corresponding to a body part of the person, and each of the second plurality of attributes has a lower weight than the first plurality of attributes.
[0030] Optionally, in another implementation of any of the foregoing aspects, the method includes: the computing device creating a thumbnail representing the summary video of the person, the thumbnail including an image showing the person's face; and the computing device displaying the thumbnail.
[0031] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that one or more images of the plurality of images including the person are combined by adding one or more transition images to one or more images of the plurality of images.
[0032] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the summary video is automatically created as background activity of the computing device.
[0033] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the method further includes: the computing device displaying a notification indicating that the summary video is being created or has been created.
[0034] According to one aspect of this disclosure, an apparatus implemented as a computing device is provided. The apparatus includes a memory and one or more processors, the memory including instructions, the one or more processors communicating with the memory, the one or more processors executing the instructions to: determine a plurality of attributes, each attribute describing a region of interest corresponding to a body part of a person displayed in an image; determine a corresponding score for each of the plurality of attributes; and calculate a total score based on the corresponding scores of the plurality of attributes.
[0035] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that, in response to receiving a selection of the person displayed in the image, the person is determined, and the image displays the person and a plurality of other people.
[0036] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the person is determined based on the face of the person detected in the image.
[0037] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the corresponding score of each of the plurality of attributes is determined based on training data including a plurality of predefined scores of each of the plurality of attributes.
[0038] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the one or more processors further execute the instructions to search for regions of interest corresponding to different body parts of the person depicted in the image, based on the probability that the body part is located at a certain position within the image.
[0039] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that when the image depicts more than one person, the one or more processors further execute the instructions to determine a corresponding score for each of a plurality of group attributes, wherein the plurality of group attributes describe at least one of the following: the relationship between the plurality of other people depicted in the image, the space between each of the plurality of other people depicted in the image, the pose of one or more of the plurality of other people depicted in the image, or the arrangement of the plurality of other people depicted in the image, and further calculate the total score based on the corresponding scores of the plurality of group attributes.
[0040] According to one aspect of this disclosure, an apparatus implemented as a computing device is provided. The apparatus includes a memory and one or more processors, the memory including instructions, the one or more processors communicating with the memory, the one or more processors executing the instructions to: determine one or more images from a plurality of images including a person in a video; and combine the one or more images from the plurality of images including the person to create a summary video of the person.
[0041] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the one or more processors further execute the instructions to receive a selection of the person.
[0042] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the one or more processors further execute the instructions to detect a human face in one or more of the plurality of images, wherein the human is determined by detecting the human face.
[0043] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that one or more images of the plurality of images including the person are determined based on the total score of each of the plurality of images.
[0044] Alternatively, in any of the above aspects, another implementation of the aspect specifies that the total score is calculated based on multiple attributes, each attribute describing a region of interest corresponding to a body part of the person.
[0045] Optionally, in any of the foregoing aspects, another implementation of the aspect specifies that the one or more processors further execute the instructions to create a thumbnail representing the summary video of the person, the thumbnail including an image showing the person's face; and cause a display device to display the thumbnail.
[0046] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the summary video is automatically created as background activity of the computing device.
[0047] Alternatively, in any of the foregoing aspects, another implementation of the aspect specifies that the summary video is automatically created when the computing device is charging.
[0048] The embodiments disclosed herein enable computing devices to automatically determine the aesthetic value of an image based on a total score calculated for that image. Computing devices capable of automatically determining the aesthetic value of an image can utilize processing and storage resources more efficiently. For example, computing devices that calculate the total score of an image do not need to unnecessarily waste processing power or resources on manually edited images. Furthermore, computing devices that calculate the total score of an image can be used to maintain high-quality storage, rather than unnecessarily wasting storage resources on low-quality images.
[0049] For clarity, any of the above embodiments may be combined with any one or more embodiments in the other embodiments described above to create a new embodiment within the scope of this disclosure.
[0050] These and other features will become clearer from the following detailed description in conjunction with the accompanying drawings and claims. Attached Figure Description
[0051] To gain a more thorough understanding of this disclosure, reference is now made to the following brief description in conjunction with the accompanying drawings and specific embodiments, wherein the same reference numerals denote the same parts.
[0052] Figure 1 Figures illustrating a system for implementing portrait image evaluation, provided for various embodiments;
[0053] Figure 2 This is a schematic diagram of an embodiment of a computing device.
[0054] Figure 3 A flowchart of a method for performing portrait image analysis provided for embodiments disclosed herein.
[0055] Figure 4 A graph representing a single-person portrait image segmented based on the region of interest of a person in the image.
[0056] Figure 5A C is a graph of multiple images segmented based on the people depicted in the image.
[0057] Figure 6 A flowchart of a method for determining the properties of an image being analyzed.
[0058] Figure 7A This is a graph of a score tree that can be used to calculate the total score of an image.
[0059] Figure 7B An example of an image rating tree is shown.
[0060] Figures 8A to 8B A diagram illustrating the segmentation and object classification methods provided for various embodiments of this disclosure.
[0061] Figure 9A and Figure 9BThe diagrams provided for various embodiments of this disclosure illustrate how to identify locations in an image that may indicate certain regions of interest.
[0062] Figure 10 A flowchart illustrating a method for performing portrait image analysis provided for various embodiments of this disclosure.
[0063] Figure 11 A schematic diagram of an album including original video and one or more summary videos based on people depicted in the original video, provided for various embodiments of this disclosure.
[0064] Figure 12 A schematic diagram of an information page for a summary video provided for various embodiments of this disclosure.
[0065] Figure 13 A schematic diagram illustrating a method for evaluating an image as a portrait image of multiple people as a single portrait image, provided for various embodiments of this disclosure.
[0066] Figure 14 A flowchart illustrating a method for performing portrait image analysis based on a person depicted in an image, provided for various embodiments of this disclosure.
[0067] Figure 15 A flowchart illustrating a method for creating a summary video based on a selected person, provided for various embodiments of this disclosure. Detailed Implementation
[0068] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques (whether currently known or existing). This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0069] The subjective features of an image can refer to the quality of certain attributes or characteristics of the people depicted in the image. A subjective feature of an image can be the facial expression of the people depicted in the image. Another example of a subjective feature of an image can be the arrangement of multiple people depicted in the image. In some cases, when people are depicted in an image, the aesthetic value of the image can depend largely on the subjective features of the people depicted in the image. An image depicting everyone smiling is more aesthetically pleasing than an image of someone not being ready to be photographed. However, a device may not be able to determine the aesthetic value of an image based on the subjective features that define the attributes of the people depicted in the image. Furthermore, a device may also be unable to determine the aesthetic value of an image based on the subjective features of a specific person in the image while ignoring the subjective features of other people depicted in the image.
[0070] In one embodiment, the computing device may store multiple videos and images, where a video consists of a sequence of multiple images. In one embodiment, an image or video may describe multiple different people as the main feature of the image; such an image is referred to herein as a multi-person portrait image. For a multi-person portrait image, a total score for the image can be calculated based on group attributes of all people depicted in the image and characteristic attributes of each individual depicted in the image. Group attributes may refer to features describing the relationships between multiple people in the image or the spatial arrangement of multiple people in the image, as described further below. Characteristic attributes may refer to the actual characteristics of the people depicted in the image (e.g., a person's emotions or facial expressions), as described further below. However, in some cases, it may be necessary to score a multi-person portrait image based on a specific person in the image, without considering the characteristics of the other people depicted in the image.
[0071] The embodiments disclosed herein aim to perform portrait image evaluation on the main subject (e.g., a person as the central focus) in the image analysis attributes of selected people depicted in a multi-person portrait image to determine the aesthetic value of the image. In embodiments where the image includes multiple different people, a user can select a specific person to rate the image. Similarly, when a video includes multiple different people, a user can select a specific person to rate the video. For example, each image in the video can be rated based on the subjective characteristics of the selected person, without considering the subjective characteristics of other people depicted in each image of the video. Other people depicted in the image can be rated based on objective factors (e.g., color, saturation, blur).
[0072] In one embodiment, a user can also select a specific person to generate a summary video of that specific person using the original video. The summary video can be a video including one or more images from the original video depicting the selected person. Each image included in the summary video of the selected person can be assigned a score based on the selected person and the duration of the summary video. Summary videos of any different person depicted in the original video can be created.
[0073] Figure 1 Figures are provided for various embodiments of a system 100 for performing portrait image evaluation. The system 100 may include a computing device 103, a network 106, and a service provider 109 interconnected via a link 111. It should be understood that the system 100 may include other components. The system 100 may be used to develop, package, and send software components to the computing device 103 for performing portrait image analysis.
[0074] The network 106 is a network infrastructure including multiple network nodes 114 interconnecting the computing device 103 with the service provider 109. The network 106 provided in the embodiments disclosed herein may be a packet network for supporting the transmission of software components and data that can be used to perform portrait image analysis. The network 106 is used to implement network configuration to configure flow paths or virtual connections between the computing device 103 and the service provider 109. The network 106 may be a backbone network connecting the service provider 109 to the computing device 103. The network 106 may also connect the service provider 109 and the computing device 103 to other systems, such as the external Internet, other cloud computing systems, data centers, and any other entities accessing the service provider 109.
[0075] The network node 114 can be a router, bridge, gateway, virtual machine, and / or any other type of node used for packet forwarding. The network nodes 114 can be interconnected using link 116. Link 116 can be a virtual link (a logical path between the network nodes 114) or a physical link. Those skilled in the art will understand that any suitable virtual or physical link can be used to interconnect the network nodes 114. Link 111 can be a wired or wireless link that connects the edge network node 114 located at the edge of the network 106 to the service provider 109 and the computing device 103.
[0076] The computing device 103 may be a user device, such as a mobile phone, mobile tablet, wearable device, Internet of Things (IoT) device, or personal computer. In some embodiments, the computing device 103 may be a device capable of acquiring images or videos using a camera 119 or any other image acquisition device. In some embodiments, the computing device 103 may not include a camera, but may also perform portrait image analysis using images received from other devices or from memory, according to the embodiments disclosed herein.
[0077] The service provider 109 may be one or more devices or servers providing services to the computing device 103 via the network 106. In the system 100, the service provider 109 may be used to create training data 120, and the portrait image analysis module 125 may perform portrait image analysis. In some embodiments, the training data 120 may be data generated based on the analysis of a large number of professional-quality prototype images 123. The service provider 109 may store a collection of hundreds or thousands of professional-quality prototype images 123. The prototype images 123 may be portrait images depicting one or more people. In the portrait images, one or more people may be depicted as the most important feature of the image, rather than landscapes or backgrounds.
[0078] The prototype images 123 can be divided into multiple human-related training sets. Each human-related training set may include images of the same quality taken by a photographer using the same camera. Each image in a human-related training set may depict the same person and may contain the same number of people. In some cases, each image in the human-related training set may contain a single subject or scene. In some cases, each image in the human-related training set may have multiple portrait compositions, where each person depicted displays a variety of different emotions and poses.
[0079] For example, the human-related training set may have a minimum threshold number of images. Each image in the human-related training set may depict the same person performing different actions and expressing different emotions. Similarly, each image in the human-related training set may depict the person from various angles and scales. The prototype images 123 may include thousands of human-related training images. In this way, each image in the human-related training set can be used to determine accurate predefined scores for various attributes of the person, as described further below.
[0080] The training data 120 includes data determined using the prototype image 123 and can subsequently be used by the computing device 103 to perform portrait image analysis on the image 130 currently being analyzed. In some embodiments, the training data 120 may include predefined scores mapped to certain attributes, scoring rules that can be applied to certain attributes, and predefined weights assigned to certain attributes.
[0081] Users or professional photographers can determine the predefined scores based on an analysis of how certain attributes of a person contribute to the aesthetic value of multiple prototype images 123. For example, each of the prototype images 123 (or each image in a different human-related training set) can be examined to determine a predefined score for the attribute shown in each of the images. The attribute of the person depicted in the image refers to the features or regions of interest of the person in the image (e.g., face, mouth, eyes), as described further below. The predefined score can be a value that rates the attributes of an image proportionally (e.g., from 0 to 1 or from 1 to 10), reflecting how the attribute contributes to the overall aesthetic value of the prototype image 123. As an illustrative example, suppose that for the eyes of a person shown in an image (e.g., a region of interest), the attribute can describe whether the person's eyes are open or closed. In this case, the predefined score for the attribute of open eyes could be 1, while the predefined score for the attribute of closed eyes could be 0, where a predefined score of 1 indicates a higher quality attribute than a predefined score of 0.
[0082] In one embodiment, the predefined score of an attribute can be based on multiple predefined scores of the attribute manually determined from multiple different professional photographers or users at the service provider 109. The predefined scores of the attribute for each of the different professional photographers can be averaged together as the predefined score of the attribute stored in the training data 120. For example, suppose the angle of the face of the person shown in the image contributes to the aesthetic value of the image. In this case, the photographer can analyze the faces (e.g., regions of interest) of multiple different prototype images 123 with multiple different facial angles (e.g., attributes) to determine multiple different predefined scores for each facial angle shown in the prototype images 123. These predefined scores can be averaged to create a single predefined score for each facial angle, and then the single predefined score is stored in the training data 120. In one embodiment, predefined scores generated for similar attributes can be averaged to create a single predefined score for the attribute.
[0083] The scoring rule can be a rule used in conjunction with the predefined score to determine the score of a region of interest (ROI) or an attribute of the image being analyzed. In one embodiment, the scoring rule can be a value calculated and considered during the determination of the score for the ROI of the image being analyzed. For example, the scoring rule can be used for attributes or ROIs that may not have a matching predefined score that exactly matches the attribute or ROI being scored. Further details regarding the scoring rule will be described below.
[0084] The people depicted in the various prototype images 123 can be described using multiple different types of attributes. The attributes of each person depicted in the images that can be identified may include feature attributes, location attributes, general attributes, group attributes, action / behavior attributes, and various other types of attributes.
[0085] In one embodiment, a feature attribute can describe the facial expression or emotion of the depicted person. The feature attribute can indicate whether the person depicted in the image is smiling, and this feature attribute can be determined by analyzing the mouth (e.g., region of interest) of the person in each of the prototype images 123. Based on this analysis, a relative and predefined score can be assigned to each of the different types of mouth expressions depicted in the prototype images 123. The mouth portion depicting a smiling person can receive the highest score of 1, and the mouth portion depicting a non-smiling person can receive the lowest score of 0. The training data 120 can store the predefined scores for each different feature attribute.
[0086] In one embodiment, the position attribute describes the location of a person within the image, or the location of various body parts of a person within the image. The position attribute may include the angle of the body, the angle of the face, the proportions of the person's body within the image, or the position of the arms or legs.
[0087] In some embodiments, a predefined score can also be assigned to each of the different location attributes. The predefined score for each location attribute can be determined by analyzing multiple prototype images 123 and then correlating how the location attribute of each of these prototype images 123 affects the aesthetic value of the image. The training data 120 can store the predefined score for each variant of the different variations of the location attribute.
[0088] In one embodiment, the general attributes include a general description of the person depicted in the image, such as gender and age range. Scores may or may not be assigned to the general attributes. However, the general attributes can be used to determine predefined weights for certain attributes, as further described in Figures 7 and 9 below. The predefined weights for the portions and attributes can be determined in a manner similar to determining the predefined scores, for example, by analyzing multiple prototype images 123 to determine the proportion by which certain general attributes influence the aesthetic value of the image.
[0089] In one embodiment, a motion behavior attribute refers to a characteristic description of the action or movement performed by the person depicted in the prototype image 123. A motion behavior attribute can describe whether the person is posing, running, sitting, standing, playing, falling, or jumping. A motion behavior attribute can also be a characteristic description of a specific action performed by a single body part of the person. For example, a motion behavior attribute could refer to whether the person's hand is open or closed.
[0090] Predefined scores can also be assigned to action behavior attributes in a manner similar to assigning predefined scores to the part, attribute, and position attributes. The scores of the action behavior attributes can also be determined by analyzing multiple prototype images 123 and then associating how each action behavior attribute in the action behavior attributes of each prototype image 123 affects the aesthetic value of the image. The training data 120 can store the predefined scores for each different action behavior attribute.
[0091] In some embodiments, such as when an image depicts more than one person, predefined scores for group attributes can describe the relationships between the individuals displayed in the prototype image 123. Group attributes can be the spatial relationships between the individuals depicted in the image, or the arrangement of the individuals in the image. The training data 120 can store predefined scores for each group attribute.
[0092] Thus, the service provider 109 stores different predefined scores for all different attributes of different people that can be depicted in the images, based on the analysis of one or more prototype images 123. The training data 120 includes predefined scores for all different variations of attributes of regions of interest that can be depicted in the prototype images 123. The training data 120 may also store all different predefined weights for each of the general attributes and other weights that can be used to determine the total score of the image 130 currently being analyzed by the computing device. In some embodiments, the training data 120 may include descriptive data for each predefined score, such that the computing device 103 can use the descriptive data to match the attributes of the image 130 identified at the computing device 103 with the descriptive data of the predefined scores of the prototype images 123 in the training data 120. In one embodiment, the descriptive data may include a mapping between the predefined scores and the attributes of the prototype images 123.
[0093] In one embodiment, the computing device 103 can store the training data 120 and implement the portrait image analysis module 125. The following is in conjunction with... Figures 3 to 10Further description: the portrait image analysis module 125 can be used to identify portions (also referred to as detector phases) corresponding to different regions of interest in the currently analyzed image 130. One portion is a rectangle surrounding the region of interest in the image. The following is in conjunction with... Figure 4 Examples of the portions shown and described are provided. Various methods for detecting objects within the image 130 (e.g., region-based convolutional neural networks (R-CNN) or Faster R-CNN) can be used to identify portions within the image 130. Faster R-CNN is further described in the Institute of Electrical and Electronics Engineers (IEEE) document entitled “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks” (hereinafter referred to as the “Faster R-CNN document”), authored by Shaoqing Ren et al. in June 2017, which is incorporated herein by reference in its entirety.
[0094] After identifying the portions of the image 130, the portrait image analysis module 125 can analyze each portion to determine attributes describing the region of interest corresponding to that portion. As described above, the attributes can be human feature attributes, location attributes, group attributes, general attributes, action / behavioral attributes, or any other type of attribute that can be identified by analyzing the image 130. Various methods for detecting objects within the image 130 (e.g., R-CNN or faster R-CNN) can be used to determine the various attributes within the image 130.
[0095] The process of segmenting the image 130 and identifying the attributes within the image 130 can be performed using various different layers in various ways, such as convolutional layers, fully connected layers, and loss layers, each of which is further described in the Faster R-CNN file.
[0096] After determining the attributes, the portrait image analysis module 125 can then score each of the attributes based on the training data 120. In one embodiment, the computing device 103 can scan the descriptive data in the training data 120 to find descriptions corresponding to attributes that match the attributes identified in a portion of the image. When a matching description is found, the computing device 103 can retrieve a predefined score corresponding to the matching description and then determine the score of the portion as the predefined score.
[0097] In one embodiment, for location-based attributes, the computing device 103 can use predefined scores from descriptions used to describe similar location-based attributes, and then use scoring rules to perform regression analysis, etc., to determine a score for each location-based attribute, as described further below. In one embodiment, for the group attributes, the computing device 103 can similarly use the descriptions from the training data 120, which are similar to the group attributes of the identified images.
[0098] The following description, in conjunction with Figure 9, further illustrates that in some embodiments, weights can be assigned to one or more parts, attributes, location attributes, general attributes, or group attributes. The weights can be based on the training data 120, which includes predefined weights or predefined proportions that define the weights to be given to the feature attributes, location attributes, general attributes, or group attributes when calculating the total score of the image 130.
[0099] In one embodiment, the service provider 109 can generate a portrait image analysis module 125, which may include software instructions. The computing device 103 can execute these software instructions to perform portrait image analysis using the training data 120. In this embodiment, the service provider 109 can package the training data 120 and the portrait image analysis module 125 into a data packet, and then transmit the data packet to the computing device 103 via the link 111 and the network 106. The computing device 103 can download the packet and then locally install the portrait image analysis module 125 and the training data 120 onto the computing device 103, enabling the computing device 103 to implement the portrait image analysis mechanism disclosed herein.
[0100] In one embodiment, during the manufacture of the computing device 103, the training data 120 and the portrait image analysis module 125 may have already been installed on the computing device 103. The portrait image analysis module 125 and the training data 120 may be installed as part of the operating system or kernel of the computing device 103.
[0101] As disclosed herein, the computing device 103 is able to perform subjective analysis of the image without user intervention by scoring the various components of the image using predefined scores based on subjective analysis performed by professional photographers. The embodiments disclosed herein enable the computing device 103 to automatically determine whether an image possesses aesthetic value.
[0102] The total score of an image can be used in a variety of different use cases and situations. For example, the computing device 103 can be used to delete images 130 with a total score below a threshold to save memory and disk space on the computing device. In some cases, the total score of each image 130 can help the user determine whether a photo or video should be saved or deleted, thereby also saving memory and disk space on the computing device 103.
[0103] For a video (or image set), the computing device 103 typically uses the first image of the video or a random image of the video as the video's cover image. Similarly, the cover image of a photo album is typically the first image 130 of the album or a random image 130 of the album. However, for a video or photo album, the cover image may be automatically set based on the video image with the highest total score. Therefore, the computing device 103 does not waste processing power on randomly identifying the cover image of the video or photo album.
[0104] In some cases, the computing device 103 can be used to calculate the total score of the image 130 or determine its aesthetic value when the computing device 103 acquires the image 130. For example, when the computing device 103 is using the camera 119, the display of the computing device can show the total score of the image 130 intended to be acquired by the camera 119. Based on the total score displayed on the display, the user can easily determine the aesthetic value of the image 130, which may show people at various angles and with various emotions. This prevents the user from unnecessarily acquiring and storing low-quality images.
[0105] In the case where a user of the computing device 103 captures multiple consecutive images 130 of the same portrait setting within a short time frame (also referred to herein as continuous shooting or burst mode), a total score can be calculated for each of the images 130. Using the total score helps the user easily identify which images have greater aesthetic value, allowing the user to easily delete images 130 that do not have aesthetic value. Once higher-quality images 130 are determined using the total score, the user does not need to manually adjust the objective features of the images 130 or the video to create higher-quality images.
[0106] The computing device 103 is used to create custom videos or slideshows (sometimes called highlight images or videos) based on videos and images 130 stored on the computing device. For example, these custom videos or slideshows are typically small files that are easy to share on social media. In some cases, the computing device 103 can use the total score of the images 130 to create a custom video or slideshow. For example, only images 130 with higher total scores can be included in the custom video or slideshow.
[0107] In some embodiments, the computing device 103 may also use the total score to determine images 130 that are aesthetically similar to each other. For example, the computing device 103 may organize images 130 based on aesthetic similarity and may create folders of images 130 with similar aesthetic quality or total scores. It should be understood that the total score of the images 130 can be used for a variety of different applications, such as video summarization, sorted images in a photo album, sorted frames in a video, etc.
[0108] Figure 2 This is a schematic diagram of an embodiment of computing device 200. The computing device 200 can be used to implement and / or support the portrait image analysis mechanisms and schemes described herein. The computing device 200 can implement the computing device 103 as described above. The computing device 200 can be implemented in a single node, or the functionality of the computing device 200 can be implemented in multiple nodes. Those skilled in the art will recognize that the term "computing device" includes a wide range of devices, and the computing device 200 is merely an example. For example, the computing device can be a general-purpose computer, a mobile device, a tablet computer, a wearable device, or any other type of user device. The inclusion of the computing device 200 is for clarity of discussion and is in no way intended to limit the application of this disclosure to a particular computing device embodiment or a category of computing device embodiments.
[0109] At least some of the features / methods described in this disclosure are implemented in a computing device or component (e.g., the computing device 200). For example, the features / methods in this disclosure can be implemented using hardware, firmware, and / or software installed on and executing on the hardware. Figure 2 As shown, the computing device 200 includes a transceiver (Tx / Rx) 210, which may be a transmitter, a receiver, or a combination of both. The Tx / Rx 210 is coupled to multiple ports 220 for sending and / or receiving data packets from other nodes.
[0110] Processor 205 is coupled to each Tx / Rx 210. Processor 205 may include one or more multi-core processors and / or memory devices 250, which may be used as data storage, buffers, etc. Processor 205 may be implemented as a general-purpose processor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a digital signal processor (DSP).
[0111] In one embodiment, the processor 205 includes internal logic circuitry to implement the portrait image analysis module 125, and may include internal logic circuitry to implement the functional steps in methods 300 and 1000 (as discussed more fully below), and / or any other flowcharts, schemes, and methods discussed herein. Therefore, including the portrait image analysis module 125 and related methods and systems improves the functionality of the computing device 200. In an alternative embodiment, the portrait image analysis module 125 may be implemented as instructions stored in the memory device 250, which the processor 205 may execute to perform the operations of the portrait image analysis module 125. Furthermore, the portrait image analysis module 125 may optionally be omitted from the computing device 200 or 103.
[0112] The storage device 250 may include a cache for temporary storage of content, such as random-access memory (RAM). Additionally, the storage device 250 may include long-term storage for storing content for relatively long periods, such as read-only memory (ROM). For example, the cache and long-term storage may include dynamic RAM (DRAM), a solid-state drive (SSD), a hard disk, or a combination thereof.
[0113] The memory device 250 can be used to store attributes 251, the image 130, the training data 120, the score 280, the total score 285, and the weights 290. The image 130 may include multiple images stored in the memory device 250 of the computing device 200. The camera 119 on the computing device 200 or 103 can capture at least some of the images 130 and subsequently store them in the memory device 250 of the computing device 200 or 103. The image 130 may also be received from another source and stored in the memory device 250 of the computing device 200 or 103. The image 130 in relation to this disclosure is a portrait image that describes one or more people as the most important part of the image, rather than the scenery or landscape around or behind the people.
[0114] The training data 120 includes a predefined score 252 for each attribute 251 identified in a portion. The training data 120 may also include descriptive data 254 that describes the attribute 251 of the predefined score 252. The training data 120 may also include predefined weights 253 for certain regions of interest or attributes 251, which will be discussed below. Figures 7A to 7B Further description. The training data 120 may also include a scoring rule 299, which may be a value calculated and considered during the determination of a score 280 for the region of interest in the image 130 being analyzed. The attributes 251 include feature attributes 255, location attributes 260, group attributes 265, action / behavior attributes 266, and general attributes 270.
[0115] As described above Figure 1 The feature attribute 255 may be a feature of the person depicted in the image 130, determined from portions of the image (e.g., the person's emotion or expression). The position attribute 260 may be a feature describing the person's position and certain regions of interest of the person relative to the overall image 130 (e.g., facial angles or body proportions in the image). The general attribute 270 may be a feature describing the person at a general level, without referring to specific regions of interest of the person. The group attribute 265 may be a feature describing the relationship between multiple people in the image 130 and the arrangement of multiple people in the image 130 (e.g., a family photo, a group photo). The action / behavior attribute 266 may be a pose or action (e.g., a gesture or gesture) performed by the person in the image 130.
[0116] As described above Figure 1The score 280 can be a score assigned to the attributes 251 (e.g., position attribute 260, general attribute 270, group attribute 265, and / or action behavior attribute 266) of the image 130 based on the predefined score 252 and the scoring rule 299. The total score 285 can be a score 280 aggregated for each image 130 associated with (or describing) the aesthetic value of the image 130. The weight 290 can be a proportional weight assigned to each of the scores 280 based on predefined weights and predefined weighting rules when aggregating the scores 280 of the images 130 to create the total score 285.
[0117] It should be understood that by programming and / or loading executable instructions onto the computing device 200, at least one of the processor 205 and / or memory device 250 is modified, thereby partially transforming the computing device 200 into a particular machine or apparatus with novel functionality taught in this disclosure, such as a multi-core forwarding architecture. Fundamentally for electrical and software engineering techniques, functionality achievable by loading executable software into a computer can be translated into a hardware implementation using well-known design rules. The decision between implementing a concept in software or hardware typically depends on considerations of design stability and the number of units to be produced, and is unrelated to any issues involved in the conversion from the software domain to the hardware domain. Generally, designs that still require frequent changes are preferred to be implemented in software because refactoring a hardware implementation is more expensive than refactoring a software design. Generally, stable and mass-producible designs are more suitable for hardware implementation (e.g., ASICs) because mass production running the hardware implementation is cheaper than software implementations. Typically, a design can be developed and tested in software and then translated into an equivalent hardware implementation in an ASIC, which hardwareifies the software instructions, using known design rules. Just as a machine controlled by a new ASIC is a specific machine or device, a computer that has been programmed and / or loaded with executable instructions (e.g., a computer program product stored in a non-transitory medium / memory) can also be considered a specific machine or device.
[0118] In an exemplary embodiment, the computing device 200 includes: an attribute module that determines a plurality of attributes, each attribute describing a region of interest corresponding to a body part of a person displayed in an image; a scoring module that determines a corresponding score for each of the plurality of attributes; and an aggregate scoring module that calculates a total score based on the corresponding scores of the plurality of attributes. In some embodiments, the computer 200 may include other modules or additional modules for performing any step or combination of steps described in the embodiments. Furthermore, any additional or alternative embodiments or aspects of the methods described as shown in any of the figures or in any of the claims may also be contemplated to include similar modules.
[0119] In an exemplary embodiment, the computing device 200 includes: a person determination module that determines one or more images from a plurality of images, the plurality of images including a person in a video; and a summarization module that combines one or more images from the plurality of images including the person to create a summary video of the person. In some embodiments, the computer 200 may include other modules or additional modules for performing any step or combination of steps described in the embodiments. Furthermore, any additional or alternative embodiments or aspects of the methods as shown in any of the figures or recounted in any claim may also be contemplated to include similar modules.
[0120] Figure 3 This is a flowchart of a method 300 for performing portrait image analysis, provided for embodiments disclosed herein. The method 300 can be executed by computing device 200 or 103 after the portrait image analysis module 125 has been installed on computing device 103. The method 300 can be executed by computing device 200 or 103 after acquiring image 130. Image 130 can be acquired via camera 119 after it has been captured. Image 130 can be acquired by retrieving it from memory device 250. Image 130 can be acquired by receiving it from another source device.
[0121] In step 303, image segmentation can be performed on the image 130 to determine multiple portions corresponding to different regions of interest (ROIs) of one or more people depicted in the image 130. The processor 205 can execute the portrait image analysis module 125 to determine the portions of the image 130. Segmentation can be performed according to any object detection method in the image 130 (e.g., R-CNN or Faster R-CNN). For example, in Faster R-CNN, segmentation involves regression, which is a process of fine-tuning the portions surrounding the ROIs such that the portions accurately and completely surround the ROIs of the image 130. In one embodiment, as follows... Figures 8A to 8B As shown, segmentation can be performed using a faster implementation method. Each of these segments can be a rectangular frame that encloses a specific region of interest or body part of the person depicted in the image. The image 130 can be segmented into portions representing the upper body, face, mouth, and eyes of the person depicted in the image 130.
[0122] In step 306, attributes 251 of one or more portions of the image 130 can be determined. The processor 205 can execute the portrait image analysis module 125 to determine the attributes 251 of the image 130. In some cases, the service provider 109 can also execute the portrait image analysis module 125 when the service provider performs portrait image analysis. The feature attributes 255, location attributes 260, general attributes 270, group attributes 265, and action / behavior attributes 266 can be determined according to any object classification method in the image 130 (e.g., R-CNN or Faster R-CNN). In Faster R-CNN, determining attributes 251 in the image 130 can be referred to as classifying the identified portions in the image 130. Classifying the portions involves labeling the portions as certain regions of interest or labeling the portions as having the attribute 251.
[0123] The image 130, segmented into upper body, face, mouth, and eyes, can be analyzed to determine attributes 251. The upper body segment can be used to identify positional attributes 260, such as body angle. The eyes segment can be used to identify feature attributes 255, such as whether the eyes are open or closed, and the mouth segment can be used to identify further feature attributes 255, such as whether the person depicted in the image is smiling.
[0124] In step 309, a general attribute 270 of one or more portions of the image 130 can be determined. The processor 205 can execute the portrait image analysis module 125 to determine the general attribute 270 of the image 130. The general attribute 270 can also be determined based on any object classification method in the image 130 (e.g., R-CNN or Faster R-CNN).
[0125] The image 130, which has been segmented into parts of the upper body, face, mouth, and eyes, can be analyzed to determine the general attributes 270.
[0126] In step 311, group attributes 265 can be determined for one or more people depicted in image 130 based on the portions associated with each person in image 130. The processor 205 can execute the portrait image analysis module 125 to determine the group attributes 265 of image 130. The group attributes 265 can be determined according to any group classification method in image 130 (e.g., R-CNN or Faster R-CNN). Figure 5A An example of identifying group attribute 265 in image 130 is shown.
[0127] When the image depicts more than one person, group attributes 265 can also be determined for one or more portions of the image 130. For example, suppose the image depicts two people and parts of each person's body are identified. In this case, the positions of these parts can be used to determine group attributes 265, such as the spatial relationship between the two people. Figure 5A An example of identifying group attribute 265 in image 130 is shown.
[0128] In step 314, based on the training data 120 stored in the memory device 250 of the computing device 200 or 103, a score 280 is determined for at least one of the attributes 251 (e.g., feature attribute 255, location attribute 260, general attribute 270, group attribute 265, action behavior attribute 266). The processor 205 may execute the portrait image analysis module 125 to determine the score 280 of at least one of the attributes 251 identified in the image 130.
[0129] In one embodiment, the portrait image analysis module 125 may acquire an attribute 251 identified in the image 130, and then compare the attribute 251 with the descriptive data 254 to determine whether the attribute 251 has a predefined score 252 or a similar attribute 251 has a predefined score 252. In one embodiment, when the attribute 251 matches the descriptive data 254 of the predefined score 252, the portrait image analysis module 125 may determine that the predefined score 252 associated with the descriptive data 254 should be a score 280 for the specific attribute 251 being scored.
[0130] For example, suppose the training data 120 includes a predefined score 252 of 0 for descriptive data 254 describing a portion of the eyes that are closed, and a predefined score 252 of 1 for descriptive data 254 describing a portion of the eyes that are open. In this case, the portion of the eyes extracted from the image 130, or the feature attribute 255 that identifies the portion of the eyes as open or closed, can be compared with the descriptive data 254 in the training data 120 to determine the score 280 of the portion of the eyes in the image 130. When the predefined score 252 for open eyes is 1, then the score 280 of the image 130 depicting open eyes is also 1. Similarly, if the predefined score 252 for closed eyes is 0, then the score 280 of the image 130 depicting open eyes is also 1.
[0131] In some cases, the attribute 251 identified based on the region of interest may not perfectly match the value of the descriptive data 254 of the predefined score 252. That is, a region of interest may have multiple different attributes 251, and therefore not all attributes can be scored or rated in the predefined score 252. These regions of interest types with multiple variations of attribute 251 may have a discrete number of attributes 251 or a continuous number of attributes 251 describing the region of interest.
[0132] When a portion corresponding to the region of interest (ROI) can have multiple different attributes 251 describing the characteristics of the ROI, the ROI of the image 130 can have a discrete number of attributes 251. For example, the attributes 251 for mouth portion recognition can include multiple different feature attributes 255, such as neutral mouth expressions, smiling, laughing, yawning, and a talking mouth. Thus, the mouth of a person displayed in the image 130 (e.g., the ROI) can have a discrete number of attributes 251 describing the expression of the mouth. In one embodiment, a predefined score 252 can be stored for each of the discrete attributes 251 describing the ROI, where a higher predefined score 252 indicates a higher aesthetic value for the ROI. However, in some cases, the image 130 being analyzed may define the attributes 251 of the ROI, and the ROI may not have an exact corresponding predefined score 252.
[0133] Similarly, when a portion corresponding to the region of interest can exhibit a consecutive number (or a large number) of variations, the region of interest identified in the portion of the image can have a consecutive number of attributes 251. For example, the location attribute 260 of the position of a body portion within the image 130 can have a large number of variations because the body portion can be located anywhere in the image. For each of these different variations, there may be no predefined score 252 regarding where the body portion might be located. The aesthetic value of the image 130 may be poor when the body portion associated with the body of the person shown in the image 130 is located at the far left or far right of the image 130. In one embodiment, predefined scores 252 of one or more of the consecutive number of attributes 251 describing the region of interest can be stored, where a higher predefined score 252 indicates a higher aesthetic value for the region of interest. However, in some cases, the image 130 being analyzed may define the attributes 251 of the region of interest, and the region of interest may not have an exact corresponding predefined score 252.
[0134] In both cases where the region of interest can be described by a discrete number of attributes 251 or a continuous number of attributes 251, the portrait image analysis module 125 can determine the score 280 of attribute 251 based on multiple different predefined scores 252 of descriptive data 254, which describes similarity feature attributes 255 and scoring rules 299. The portrait image analysis module 125 can identify predefined scores 252 associated with similar attributes 251 as attributes 251 for scoring based on the descriptive data 254. The multiple predefined scores 252 can be normalized and processed using scoring rules 299 and a regression machine learning (ML) model to determine the score 280 of attribute 251 of the image 130. The regression ML model can be a linear regression algorithm for defining the score 280 based on multiple different predefined scores 252.
[0135] As described above, the scoring rule 299 is a value calculated and considered during the determination of a score 280 for a region of interest or attribute 251 of the image 130 being analyzed. In some embodiments, each attribute 251 or region of interest may have a corresponding scoring rule 299 for determining a score 280 for that attribute 251. As an example, for a positional attribute 260 describing the body position of a body part of a person displayed in the image 130, the scoring rule 299 may include a horizontal body position scoring rule 299 and a vertical body position scoring rule 299. The horizontal body position scoring rule 299 may be the horizontal distance between the center of the body part and the center of the image 130. The horizontal distance may be normalized relative to the width of the image 130 to determine a score 280 for the positional attribute 260 describing the body position of the body part. The vertical body position scoring rule 299 may be the vertical distance between the center of the body part and the center of the image 130. The vertical distance may be normalized relative to the width of the image 130 to determine a score 280 for the positional attribute 260 describing the body position of the body part.
[0136] As another example, for the positional attribute 260 describing the body proportions of the body portion of a person displayed in the image 130, the scoring rule 299 may include a body proportion scoring rule 299. The body proportion scoring rule 299 may be the following ratio: (height of body portion)(width of body portion) / (height of image 130)(width of image 130). The body proportion scoring rule 299 is the ratio of the size of the body portion to the size of the image 130.
[0137] As another example, for the positional attribute 260 describing the body proportions of a portion of a person displayed in image 130, scoring rule 299 may include a body proportion scoring rule 299. The body proportion scoring rule 299 may be a ratio of (height of the body portion) / (width of the body portion). The body proportion scoring rule 299 can be used to determine whether image 130 is a half-body portrait or a full-body portrait. The body proportion scoring rule 299 can also be used to determine whether the person shown in the image is sitting or standing.
[0138] In some embodiments, the scoring rule 299 can be converted into a proportion of one-dimensional continuous values that includes all different variations of the scoring rule 299. These scoring rules 299 can be aggregated into a total score for that sub-dimension through a weighted aggregation process obtained during regression model training. The goal of training the regression model is to determine the weights 290 and the aggregation equation that forms the weighted total score 285. In some cases, the training data 120 generates a training ML model that determines that the aesthetic value of the image 130 is as close as possible to that of the training data 120 generated from the prototype image 123.
[0139] Suppose that when scoring location attribute 260, such as scoring the proportion of the body within image 130, the portrait image analysis module 125 may fail to identify a specific predefined score 252 that matches the location attribute 260 of the currently analyzed image 130. When a specific attribute 251 is consecutive, the portrait image analysis module 125 may fail to identify the predefined binary score 252 for that attribute (e.g., attribute 251 has too many different variations, making it difficult to identify an exact match between attribute 251 of the currently analyzed image 130 and the descriptive data 254 of the predefined score 252). In this case, the portrait image analysis module 125 can identify multiple predefined scores 252 that have body proportions similar to those identified in the currently analyzed image 130. These predefined scores 252 and predefined scoring rules 299 can be input into the regression ML model to output a score 280 for the location attribute 260.
[0140] In some embodiments, the score 280 determined for each attribute 251 or feature of the image 130 can be further fine-tuned using an absolute error loss function. The absolute error loss function can be used to minimize the error of the score 280 determined for each attribute 251 or other feature of the image 130.
[0141] In step 317, a total score 285 representing the aesthetic value of the image 130 can be determined based on weights 290 assigned to each of the attributes 251 or features of the already scored image 130. The weights 290 can be determined based on predefined weights 253 in a manner similar to how the score 280 is determined based on predefined scores 252, as described above. The weights 290 correspond to the proportion or percentage of weight that should be given to the score 280 in the total score 285. These proportions can be stored in the predefined weights 253 and used to determine the weights 290 in the same way that the score 280 of the image 130 is determined using multiple different predefined scores 252 for multiple different attributes 251 and features of the image 130.
[0142] The processor 205 may execute the portrait image analysis module 125 to determine the total score 285 based on the score 280 and / or weight 290. If applicable, each score 280 determined for the image 130 may be weighted according to the corresponding weight 290 and then summed to generate the total score 285 for the image 130.
[0143] Figure 4 This is a diagram of a single-person portrait image 400 segmented based on the region of interest (ROI) of a person in an image. The single-person portrait image 400 has been segmented into four parts: a body part 403, a face part 406, an eye part 409, and a mouth part 411. The face part 406 can be a sub-part of the body part 403, such that there is a dependency between the face part 406 and the body part 403. Similarly, the eye part 409 and the mouth part 411 can be sub-parts of the face part 406, such that there is a dependency between the eye part 409 and the face part 406, and a dependency between the mouth part 411 and the face part 406. This dependency refers to the relationship between a part and a sub-part, where the sub-part may be located within the part. The following will combine... Figures 5A to 5C To further describe, the dependency can be used to perform more efficient image 130 and 400 segmentation.
[0144] In some embodiments, each of portions 403, 406, 409, and 411 may be analyzed to determine attribute 251 of the person depicted in the single-person portrait image 400. The body part 403 may be analyzed to determine attribute 251, such as characteristic attributes 255 and positional attributes 260 of the body part depicted in the body part 403. Characteristic attributes 255 (e.g., body size or body posture) and positional attributes 260 (e.g., body angles) may be determined by the body part 403. It should be understood that other types of attributes 251 or other features of the image 400 may be determined using the body part 403.
[0145] The facial portion 406 can be analyzed to determine general attributes 270. The eye portion 409 can be analyzed to determine characteristic attributes 255, such as whether the eyes are open or closed. The mouth portion 411 can be analyzed to determine characteristic attributes 255, such as whether the mouth is smiling, laughing, or frowning.
[0146] Figure 5A This is a diagram of a multi-person portrait image 500 that has been segmented based on the different people depicted in image 500. The multi-person portrait image 500 includes four parts for each person depicted in image 500: a first person part 503, a second person part 506, a third person part 509, and a fourth person part 511. The first person part 503, the second person part 506, the third person part 509, and the fourth person part 511 are rectangular frames that respectively enclose the four different people depicted in image 500.
[0147] In some embodiments, similar to how the single-person portrait image 400 is segmented, each of the person portions 504, 506, 509, and 511 is segmented individually and then analyzed to determine attributes 251, such as feature attributes 255, location attributes 260, general attributes 270, and action / behavior attributes 266. The general attributes of each of the person portions 504, 506, 509, and 511 can indicate the age range and gender of each person shown in the image 500.
[0148] In some embodiments, all person portions 504, 506, 509, and 511 can be analyzed to determine group attributes 265 that define the spatial relationships and arrangement of people within the image 500. Analysis of the person portions 504, 506, 509, and 511 can determine group attributes 265, such as an indication that each person depicted in each of the person portions 504, 506, 509, and 511 is sitting in a row on the floor in an embracing manner. Another set of attributes 265 that can be identified in the image 500 is that each of the person portions 504, 506, 509, and 511 slightly overlaps with each other, indicating that the people depicted in the image 500 are positioned closely together.
[0149] In some embodiments, these group attributes 265 can be combined with the general attribute 270 to identify key features of the multi-person portrait image 500, which can help generate a total score 285 for the multi-person portrait image 500. For example, from an aesthetic point of view, family portraits and family photos are highly valued by most users. The computing device 200 or 103 can easily determine whether the multi-person portrait image 500 is a family photo using the general attribute 270 and the group attribute 265. When the general attribute 270 of two of the person portions 503 and 506 indicates two different genders (male and female) with an age range much higher than the general attribute 270 of the other person portions 509 and 511, the computing device 200 or 103 can determine that the person portions 504, 506, 509, and 511 correspond to family members. Thus, the group attribute 265, such as the close spatial relationship and row arrangement of the family, can indicate that the multi-person portrait image 500 is a family portrait in which family members pose individually.
[0150] In some embodiments, these types of group portrait images 500 may include higher weights 290 for the group attribute 265 and the general attribute 270, which helps to define the group portrait image 500 as a family photograph. Higher weights 290 may also be assigned to other characteristic attributes 255 indicating whether an individual is smiling, since children sometimes do not smile while posing for photos. Objective features may also be included in the total score 285 of the group portrait image 500. The quality of the image background can be considered an attribute of the image and can be scored and weighted as described above. For example, the background of the group portrait image 500 can be analyzed to determine whether there is a single color and strong contrast between the background and the human parts 504, 506, 509, and 511. Other objective features (e.g., brightness, color balance, background blur) may also be considered as background quality of the image.
[0151] Figures 5B to 5COther examples of multi-person images 550 and 560 are shown. Multi-person image 550 shows randomly positioned human figures within a landscape. A group attribute 265 can be defined for multi-person image 550, indicating that the spatial relationships and arrangement of the human figures shown in multi-person image 550 are scattered and unorganized. A low predefined score 252 can be assigned to such group attribute 265, thereby assigning a low score 280 to the currently analyzed image 130.
[0152] The multi-person image 560 illustrates a group setup, where the human figures are arranged in several different layers. In this multi-person image 560, the group attribute 265 can reflect that the human figures are close to each other but located in layers at different depths within the multi-person image 560. Based on whether a professional photographer considers the layered image to increase or decrease the aesthetic value of image 130, a predefined score 252 can be assigned to this group attribute 265.
[0153] Figure 6 This is a flowchart of a method 600 for determining attribute 251 of image 130 being analyzed. Method 600 can be executed by computing device 200 or 103 after the portrait image analysis module 125 has been installed on the computing device 200 or 103. Method 600 can also be executed by computing device 200 or 103 after image 130 has been acquired. Alternatively, method 600 can be executed by service provider 109 after image 130 has been acquired.
[0154] In step 603, image 130 is acquired by capturing image 130 from camera 119 or retrieving image 130 from memory device 250. In some embodiments, image 130 may be one image in a collection of images 130 included in a video. For the video, each image 130 may be analyzed and scored separately, and then aggregated to create a total score 285, which is the sum of the scores 280 of each image 130 in the video.
[0155] In step 605, it is determined whether the image 130 is a portrait image. Certain portions or pixels of the image 130 may be examined to determine whether the majority of the image 130 depicts one or more people. When a certain number of pixels in the image 130 displaying human features exceeds a threshold number, the image 130 can be determined to be a portrait image. In some embodiments, the processor 205 may execute the portrait image analysis module 125 to determine whether the image 130 is a portrait image.
[0156] In step 607, human semantic detection may be performed on the image 130. Human semantic detection may involve determining attributes 251, such as feature attributes 255, location attributes 260, group attributes 265, general attributes 270, action / behavior attributes 266, and other features of the image 130. The processor 205 may execute the portrait image analysis module 125 to perform human semantic detection on the image.
[0157] like Figure 6 As shown, human semantic detection involves the detection of several layers or levels of human-related objects within image 130. In step 609, portrait classification detection may be performed on image 130. The portrait classification detection may involve determining whether image 130 can be classified as a portrait image 611 (e.g., image 400) or a multi-person image 613 (e.g., image 500). A portrait image shows a single person, while a multi-person image shows multiple people. The processor 205 may execute the portrait image analysis module 125 to determine whether image 130 shows a single person or multiple people.
[0158] In some embodiments, when the image 130 depicts multiple people, the image 130 can be segmented to create portions of each person, and each person can also be segmented by region of interest or body part. Each of these portions of each person in the image 130 can be analyzed to determine feature attribute 255A. Each of these portions of each person depicted in the image 130 can also be analyzed to determine group attribute 265, such as spatial relationship 265A and arrangement feature 265B. The spatial relationship 265A refers to the analysis of how much space is between each person depicted in the image 130. The arrangement feature 265B refers to the analysis of how the people in the image 130 are arranged. It should be understood that the spatial relationship 265A and the arrangement feature 265B are merely two examples of group attribute 265, and any number or type of group attribute 265 determined in the image 130 can exist.
[0159] In step 615, feature attributes 255 and positional attributes 260 of the image 130 can be detected. The processor 205 can execute the portrait image analysis module 125 to determine the feature attributes 255 and positional attributes 260 of the image 130. Examples of positional attributes 260 that can be determined include human body position and proportion 260A and facial angle 260B. Human body position and proportion 260A can be determined using the body portion 403, etc., and facial angle 260B can be determined using the face portion 406, etc. Examples of feature attributes 255 that can be determined include facial expression 255C and eye state 255B. Facial expression 255C can be determined using the mouth portion 411, etc., and eye state 255B can be determined using the eye portion 409, etc. It should be understood that the human body position and proportion 260A and facial angle 260B are only two examples of positional attributes 260, and any number or type of positional attributes 260 determined in the image 130 can exist. Similarly, facial expression 255C and eye state 255B are just two examples of feature attributes 255, which can exist in any number or type of feature attributes 255 as determined in the image 130.
[0160] In step 619, general attributes 270 of the people depicted in the image 130 can be detected. The processor 205 can execute the portrait image analysis module 125 to determine the general attributes 270 of the image 130 based on one or more portions of each person. Examples of general attributes 270 that can be determined include gender 270A, age range 270B, etc. It should be understood that gender 270A and age range 270B are merely examples of general attributes 270, and any number or type of general attributes 270 determined in the image 130 can exist.
[0161] In step 621, action behavior attributes 266 of the person depicted in the image 130 can be detected. The action behavior attributes 266 can be used to co-create a video from the image 130. The processor 205 can execute the portrait image analysis module 125 to determine the action behavior attributes 266 of the image 130 based on one or more parts of each person. Examples of determineable action behavior attributes 266 include determining whether the person is standing 266A, sitting 266B, walking 266C, running 266D, posing 266E, or making a gesture 266F. These action behavior attributes 266 can be determined using body parts 403, etc. It should be understood that any number or type of action behavior attributes 266 can be determined in the image 130.
[0162] In some embodiments, once the feature attribute 255, location attribute 260, group attribute 265, and general attribute 270 are determined using the portions identified in the image 130, at least one of these feature attribute 255, location attribute 260, group attribute 265, and general attribute 270 can be scored. As described above, the predefined scores 252 in the training data 120 can be used to determine a score 280 for each of these feature attribute 255, location attribute 260, group attribute 265, and general attribute 270. In some embodiments, each of the scores 280 can be summed to obtain the total score 285. In some embodiments, when determining the total score 285 of the image 130, some of these scores can be weighted according to weights 290.
[0163] Figure 7A This is a graph of a scoring tree 700 that can be used to calculate a total score 285 for image 130. The scoring tree 700 includes multiple regions of interest 703A to 703E of the person depicted in image 130. The scoring tree 700 also includes several location attributes 260A to 260C and feature attributes 255A to 255D as leaf nodes of one or more regions of interest 703A to 703E. Although only leaf nodes of location attribute 260 and feature attribute 255 are shown in the scoring tree 700, it should be understood that scoring trees 700 for other images 130 may include other attributes 251, such as group attribute 265, general attribute 270, and action / behavior attribute 266.
[0164] The scoring tree 700 may include nodes of at least one of regions of interest 703A to 703E, feature attribute 255, location attribute 260, group attribute 265, or general attribute 270, wherein each node is arranged in the scoring tree 700 based on the dependencies between the nodes. The dependency may refer to the relationship between two regions of interest, or the relationship between the regions of interest 703A to 703E and attribute 251. A dependency exists between the two regions of interest 703A to 703E when one of the regions of interest 703A to 703E is located within another region of interest 703A to 703E in the image 130. A dependency may also exist between the regions of interest 703A to 703E and attribute 251 when attribute 251 describes a specific region of interest 703A to 703E.
[0165] In some embodiments, each node representing the regions of interest 703A to 703E may include other regions of interest 703A to 703E or several leaf nodes of attribute 251, such as feature attribute 255, location attribute 260, group attribute 265, general attribute 270, action behavior attribute 266, objective features of the image 130, or other features of the image 130. However, nodes representing attribute 251 (e.g., feature attribute 255, location attribute 260, group attribute 265, or general attribute 270) may not include any leaf nodes.
[0166] As shown in Figure 7, the scoring tree 700 typically has parent nodes representing parent attributes, such as regions of interest 703A. Regions of interest 703A to 703E correspond to various parent attributes, regions of interest, or portions identified in the image 130. In some cases, the parent node representing region of interest 703A corresponds to the body portion 403 of the image 130. Figure 4 As described above, the body portion 403 refers to the rectangular frame surrounding the entire person shown in the image 130 (or, in the case of a multi-person image, one person).
[0167] Starting from the parent node representing the parent attribute of the region of interest 703A, there can be multiple leaf nodes representing other regions of interest 703B that are dependent on the region of interest 703A. As shown in FIG7, the region of interest 703B is a leaf node of the region of interest 703A because the portion corresponding to the region of interest 703B can be located within the portion corresponding to the region of interest 703A.
[0168] Starting from the parent node representing the parent attribute of the region of interest 703A, there can also be multiple leaf nodes representing attributes 251 that have a dependency relationship with the region of interest 703A. As shown in Figure 7, position attributes 260A and 260B are leaf nodes of the region of interest 703A because the position attributes 260A and 260B describe the positional characteristics of the region of interest 703A.
[0169] As shown in Figure 7, the region of interest 703B includes four leaf nodes: one leaf node with location attribute 260C, and leaf nodes for three different regions of interest 703C to 703E. The location attribute 260C may be dependent on the region of interest 703B (or describe a feature of the region of interest 703B). The regions of interest 703C to 703E may be dependent on the region of interest 703B (or located within the region of interest 703B). Similarly, each of the regions of interest 703C to 703E has leaf nodes representing various feature attributes 255A to 255D, which describe the features of the regions of interest 703C to 703E.
[0170] In some embodiments, the rating tree 700 may include weights 290A to 290K for each node within the rating tree 700, except for the topmost parent node. All nodes in the rating tree 700 include weights 290A to 290K, except for the node representing the region of interest 703A.
[0171] The weights 290A to 290K are values between 0 and 1, and these values can be assigned to certain parts or regions of interest 703A to 703E corresponding to said parts. The weights 290A to 290K can also be values between 0 and 1 assigned to attribute 251 (e.g., feature attribute 255, location attribute 260, group attribute 265, general attribute 270, action behavior attribute 266, objective attribute, or any feature of the image 130). The weights 290A to 290K can indicate the proportion of weight given to a certain part, region of interest 703A to 703E, or attribute 251 when calculating the total score 285 of the image 130.
[0172] In some embodiments, the weights 290 can be determined in a manner similar to determining the score 280 of the image. A professional photographer located at the service provider 109 can determine a percentage or proportion of a portion of the prototype image 123 (e.g., a part, regions of interest 703A to 703E, feature attribute 255, location attribute 260, group attribute 265, general attribute 270, action / behavior attribute 266, or other attribute / feature) relative to the overall aesthetic value of the prototype image 123. Each of the prototype images 123 can be analyzed to determine the relative proportion of each of the portion, regions of interest 703A to 703E, feature attribute 255, location attribute 260, group attribute 265, general attribute 270, action / behavior attribute 266, or other attribute / feature to the total aesthetic value of the prototype image 123. Based on this, the service provider 109 can store predefined weights 253, which may correspond to the portion, regions of interest 703A to 703E, feature attributes 255, location attributes 260, group attributes 265, general attributes 270, action behavior attributes, or other attributes / features of the image 130. The service provider 109 can send these predefined weights 253 from the training data 120 to the computing device 200 or 103, so that the computing device 200 or 103 can use the predefined weights to determine the actual weights 290 of the image 130.
[0173] In some cases, the actual weight 290 of an image 130 may not actually match the predefined weight 253 of the training data 120. This is because each image 130 does not include the same parts and features. Therefore, the computing device 103 can calculate the weight 290 of a particular image 130 using the predefined weight 253 relatively based on all parts, regions of interest 703A to 703E, feature attribute 255, position attribute 260, group attribute 265, general attribute 270, action behavior attribute 266, or other attributes / features rated in the image 130.
[0174] like Figure 7A As shown, weights 290A to 290K are assigned to all nodes in the rating tree 700. In some embodiments, all leaf nodes originating from a single node and located at a level of the rating tree 700 should be equal to 1. The set of weights 290A to 290C should be equal to 1 because the nodes of the region of interest 703B, position attribute 260A, and position attribute 260B originate from the nodes of the region of interest 703A. The set of weights 290D to 290E should be equal to 1 because the nodes of position attribute 260C and regions of interest 703C to 703E originate from the nodes of the region of interest 703B.
[0175] The nodes of region of interest 703C have only one leaf node with feature attribute 255A. A weight 290H of 1 can be assigned to feature attribute 255A. Similarly, the nodes of region of interest 703E have only one leaf node with feature attribute 255D. A weight 290K of 1 can also be assigned to feature attribute 255D. The nodes of region of interest 703D have leaf nodes with two feature attributes, 255B and 255C. In this case, the set of weights 290I and 290K can be equal to 1.
[0176] In some embodiments, a total score 285 can be calculated based on the scoring tree 700 using the score 280 of each node in the scoring tree 700 and the weights 290A to 290K of each node in the scoring tree 700. The score 280 can be calculated for each of the regions of interest 703A to 703E, location attributes 260A to 260C, and feature attributes 255A to 255D shown in the scoring tree 700. The total score 285 can be calculated by first multiplying each score 280 by its corresponding weights 290A to 290K to determine the weighted score 280 of each node in the scoring tree 700, and then calculating the set of weighted scores 280.
[0177] In some embodiments, the structure of the scoring tree 700 allows for easy adjustment of the weights 290A to 290K of the features when considering scoring additional features of the image 130 to account for new features. In some embodiments, the weights 290A to 290K can be readjusted based on the predefined weights 253 and the total weights of the levels in the scoring tree 700.
[0178] For example, suppose that when calculating the new total score 285, a location attribute 260E that was not previously considered for the total score 285 needs to be taken into account. Suppose that location attribute 260E defines a feature of the region of interest 703B. In this case, leaf nodes can be added to the region of interest 703B. Similarly, weights 290D to 290G can be recalculated to add another weight 290 to the new location attribute 260E based on the predefined weights 253, while ensuring that the set of weights 290D to 290G and the new weight 290 remains equal to 1. Thus, no adjustments to other weights 290 or scores 280 are needed to calculate the new total score 285.
[0179] Figure 7B An example of a rating tree 750 for image 130 is shown. Figure 7BAs shown, the region of interest 703A of the parent node of the parent attribute of the scoring tree 750 corresponds to the body portion 403 of the image 130. The score 280 of the parent node (parent attribute) corresponding to the body portion 403 is 0.6836. The node representing the region of interest 703A has four leaf nodes: one leaf node for the region of interest 703B (face portion 406), two leaf nodes for position attributes 260A and 260B, and one leaf node for action attribute 266A. The region of interest 703B is dependent on the region of interest 703A because the face portion 406 is located within the body portion 403. The score 280 of the node representing the region of interest 703B is 0.534, and the weight 290 is 0.4.
[0180] The location attributes 260A and 260B describe the positioning of the body part 403 and are therefore dependent on the region of interest 703A. Location attribute 260A describes the position of the person within the image 130, and location attribute 260B describes the size ratio of the person relative to the image 130. Location attribute 260A has a score of 0.9 (280 points) and a weight of 0.3 (290 points). Location attribute 260B has a score of 0.5 (280 points) and a weight of 0.2 (290 points). The action / behavior attribute 266A describes the posture of the body part 403 depicted in the image 130 and is therefore dependent on the region of interest 703A. Action / behavior attribute 266A has a score of 1 (280 points) and a weight of 0.1 (290 points).
[0181] A node representing the region of interest 703B (face portion 406) can have four leaf nodes: one node representing the region of interest 703C (eye portion 409), one node representing the region of interest 703D (mouth portion 411), one node representing the region of interest 703E (skin portion), and one node representing a position attribute 260C. The position attribute 260C describes the angle of the face and is therefore dependent on the region of interest 703B. The position attribute 260C has a score 280 of 0.8 and a weight 290 of 0.1.
[0182] The regions of interest 703C to 703E are dependent on the region of interest 703B because the eye portion 409, the mouth portion 411, and the skin portion can be located within the face portion 406. The region of interest 703C (eye portion 409) has a score of 0.55 and a weight of 0.4. The region of interest 703D (mouth portion 409) has a score of 0.5 and a weight of 0.3. The region of interest 703E (skin portion) has a score of 0.78 and a weight of 0.2.
[0183] The node representing the region of interest 703C (eye portion 409) has two feature attributes, 255A and 255B. Feature attribute 255A indicates whether the eye in eye portion 409 is open or closed, and feature attribute 255B indicates the degree of focus of the eye in eye portion 409 in the image 130. Thus, feature attributes 255A and 255B are dependent on the region of interest 703C because they define the characteristics of the region of interest 703C. Feature attribute 255A has a score 280 of 1 and a weight 290 of 0.5. As mentioned above, a score 280 of 1 for eye portion 409 indicates that the eye shown in eye portion 409 is open. Feature attribute 255B has a score 280 of 0.1 and a weight 290 of 0.5. For example, this low score of 0.1 for eye focus 280 can indicate that the eye is not focused on the camera capturing the image 130, or that the pixels of the eye portion 409 are not focused.
[0184] A leaf node representing the region of interest 703D (mouth portion 411) has a feature attribute 255C, which indicates whether the mouth in the mouth portion 411 is smiling. Thus, feature attribute 255C is dependent on the region of interest 703D because it defines the characteristics of the region of interest 703D. The score 280 of attribute 255C is 0.5, and the weight 290 is 1 (because there are no other leaf nodes originating from the node of the region of interest 703D). A score 280 of 0.5 for the mouth portion 411 can indicate that the person depicted in image 130 is not fully smiling or is indifferent.
[0185] The node representing the region of interest 703E (skin portion) has leaf nodes with two feature attributes, 255D and 255E. Feature attribute 255D represents skin color, and feature attribute 255E represents skin smoothness. Thus, feature attributes 255D and 255E are dependent on the region of interest 703E because they define the characteristics of the region of interest 703E. Feature attribute 255D has a score of 0.5 (280) and a weight of 0.3 (290). Feature attribute 255E has a score of 0.9 (280) and a weight of 0.7 (290).
[0186] As shown in the scoring tree 750, the set of weights 290 of leaf nodes originating from a single node should be equal to 1. The set of leaf nodes originating from the parent node representing the region of interest 703A is 1 (0.4 + 0.3 + 0.2 + 0.1). The set of leaf nodes originating from the node representing the region of interest 703B is 1 (0.1 + 0.4 + 0.3 + 0.2). The set of leaf nodes originating from the node representing the region of interest 703C is 1 (0.5 + 0.5). The set of leaf nodes originating from the node representing the region of interest 703D is also 1, because there is only a single leaf node representing the feature attribute 255C. The set of leaf nodes originating from the node representing the region of interest 703D is 1 (0.3 + 0.7).
[0187] The total score 285 can be calculated by weighting and summing all scores 280 of all nodes (representing parts of the image 130, regions of interest 703A to 703E, feature attribute 255, location attribute 260, group attribute 265, general attribute 270, action behavior attribute 266, or other attributes / features). If the total score 285 includes the set of all scores 280, then the total score 285 of the image 130 represented by the scoring tree 750 is 9.2476 (0.6836+0.534+0.9+0.5+1+0.8+0.55+0.5+0.78+1+0.1+0.5+0.5+0.9). If the total score 285 includes the set of all weighted scores 280 (the scores 280 multiplied by their respective weights 290), then the total score 285 of the image represented by the scoring tree 750 is 3.0832((0.6836)+(0.534×0.4)+(0.9×0.3)+(0.5×0.2)+(1×0.1)+(0.8×0.1)+(0.55×0.4)+(0.5×0.3)+(0.78×0.2)+(1×0.5)+(0.1×0.5)+(0.5×1)+(0.5×0.3)+(0.9×0.7)).
[0188] The scoring tree 700 is merely an example of a data structure that can be used to generate a total score 285 for image 130. It should be understood that any other type of data structure or trained model that can readily factorize the additional attributes 251 or feature factors of image 130 into the total score 285 based on weights 290 can be used to determine the total score of image 130.
[0189] Figures 8A to 8B Figures 800 and 850 illustrate methods for segmentation and object classification provided for various embodiments of this disclosure. Figure 800 shows a conventional method for identifying portions in image 130. One of the initial steps in segmenting image 130 is to use a region proposal network (RPN), which involves searching image 130 using a sliding window from the upper left corner to the lower right corner of image 130 to find possible portions. The sliding window is a sliding rectangle that is resized and scaled to slide across the entire image 130 for multiple iterations. During each iteration, the sliding window moves across the entire image 130 until the portion is identified.
[0190] Figure 8A Figure 800 illustrates a conventional method for identifying portions using a sliding window 803 in implementing an RPN. Anchor point 806 represents the center point of the sliding window 803 as it moves over time. The sliding window 803 is positioned across the entire image 130 at a first scale 809 (the size of the sliding window 803) and a first ratio 811 (the dimensions of the sliding window 803) for the first iteration. After the first iteration, the scale 809 and ratio 811 can be changed, and the sliding window 803 is again moved across the entire image 130 to identify portions. Several iterations are performed by changing the scale 809 and ratio 811 of the sliding window 803 to determine proposals (or proposed portions) that can surround the regions of interest 703A to 703E. Regression can be performed on the proposed portions to correct them and ensure that the portions surround the regions of interest near their edges. The portions can then be classified to label them and determine attribute 251.
[0191] The conventional RPN method using the sliding window 803 is inefficient because the sliding window 803 is often located at the edges and regions of the image 130, where it is unlikely that the regions of interest 703A to 703E can be located. Therefore, the processor 205 spends a significant amount of time attempting to define proposed portions of the image 130 that are unrelated to the portrait image analysis embodiments disclosed herein.
[0192] Figure 8BFigure 850 illustrates a more efficient method for identifying portions using the updated anchor point 853 based on the probability that a region of interest is located near the updated anchor point 853. The updated anchor point 853 can be used in implementing the RPN according to various embodiments disclosed herein. In some embodiments, the training data 120 may include predefined anchor points based on the probability that certain regions of interest 703A to 703E will be located in the image 130. Predefined anchor points defining human body portions (body portion 403) are more likely to be located in the center of the image 130. The updated anchor point 853 of the sliding window 803 used for the first iteration of the RPN can be located based on the predefined anchor points. Figure 8B The updated anchor point 853 shown can be the anchor point of the sliding window 803, which is used to identify the corresponding part of the body.
[0193] In some embodiments, since the likelihood of identifying portions of the regions of interest 703A to 703E is high, the number of iterations using the sliding window 803 can be reduced. Thus, the number of ratios 809 and 811 for the various iterations used to move the sliding window 803 along the image 130 is also reduced.
[0194] Similar predefined anchor points can be included in the training data 120 for the various regions of interest 703A to 703E segmented and analyzed in the portrait image analysis mechanism disclosed herein. The training data 120 may include predefined anchor points pointing to specific parts or points in the image 130 where the person's face is located, the location of the person's eyes, the location of the person's mouth, etc. Using these embodiments for segmentation significantly reduces the number of proposed parts identified and the time required to process the image 130. Therefore, if these segmentation embodiments are utilized, the mechanism for portrait image analysis can be implemented much faster.
[0195] Figure 9A and Figure 9B The diagram provided for various embodiments of this disclosure illustrates how to identify the locations in image 130 that may show certain regions of interest 703A to 703E. Figure 9A A heatmap 903 and a three-dimensional (3D) image 906 are shown, corresponding to the possible locations of the upper body portion within the image 130. These images 903 and 906 can be generated based on analysis of the prototype image 123. Similarly, Figure 9B A heatmap 953 and a 3D diagram 956 are shown, corresponding to the possible locations of the eye (eye portion 409) within the image 130. These diagrams 953 and 956 can also be generated based on analysis of the prototype image 123.
[0196] Figure 10 A flowchart of a method 1000 for performing portrait image analysis provided for various embodiments of the present disclosure. The method 1000 can be executed by the portrait image analysis module 125 after acquiring an image 130 to be analyzed and scored. In step 1003, multiple attributes 251 are determined, each describing a multiple region of interest corresponding to a body part of a person depicted in the image 130. The processor 205 executes the portrait image analysis module 125 to determine the attributes 251 in the image 130 based on portions identified in the image 130. The identified attributes 251 may be feature attributes 255, location attributes 260, general attributes 270, group attributes 265, action / behavior attributes 266, objective features of the image 130, and / or other features describing the image 130.
[0197] In step 1006, a corresponding score 280 for each attribute 251 can be determined based on the training data 120. The processor 205 executes the portrait image analysis module 125 to determine a corresponding score 280 for each attribute in the attributes 251 based on predefined scores 252 stored in the training data 120. The predefined scores 252 are preset scores for various attributes based on the prototype image 123. In some embodiments, each score 280 can be weighted according to a weight 290 assigned to each attribute in the attributes 251 being scored. In one embodiment, the weight 290 for each attribute in the attributes 251 is based on predefined weights 253 included in the training data 120.
[0198] In step 1009, a total score 285 is calculated based on the corresponding scores 280 of attribute 251. The processor 205 executes the portrait image analysis module 125 to calculate the total score 285 based on the corresponding scores 280 of attribute 251. The total score 285 can be a set of scores 280 of attribute 251. Alternatively, the total score 285 can be a set of scores 280 weighted according to the weights 290 of attribute 251.
[0199] Figure 11A schematic diagram 1100 is provided for various embodiments of this disclosure, including an original video 1106 and an album 1103 comprising one or more summary videos 1109 and 1112 based on the people depicted in the original video 1106. The original video 1106 may include a series of one or more images 130, some of which may be portrait images of multiple people. The original video 1106 may be shown in the album 1103 via a cover image 1117A, which may be selected based on a method of performing portrait image analysis as described above. For example, the image 130 of the original video 1106 with the highest total score of 285 may be the cover image 1117A of the original video 1106.
[0200] Each of the summary videos 1109 and 1112 may be a video comprising a series of images 130 (e.g., frames) from the original video 1106 displaying the selected person. In one embodiment, Figure 2 The computing device 200 can be used to generate summary videos 1109 and 1112 from the original video 1106 based on specific people shown in the original video 1106. For example, a user of the computing device 200 can watch the original video 1106 and then access an information page for the original video 1106. The information page for the original video 1106 can display thumbnails of each person depicted in the original video 1106. For example, the information page for the original video 1106 may include thumbnails of each person depicted in at least a threshold number of images 130 in the original video 1106. The following is in conjunction with... Figure 12 A further example of an information page.
[0201] Users accessing the information page of the original video 1106 can select one of the thumbnails corresponding to the person depicted in the original video 1106 to create a summary video 1109 or 1112 of the selected person. For example, the information page of the original video 1106 may include a thumbnail 1115A of the man depicted in the original video 1106 and a thumbnail 1115B of the girl depicted in the original video. Users accessing the information page of the original video 1106 can select these two thumbnails 1115A and 1115B at different times to create a summary video 1109 corresponding to the man of thumbnail 1115A and a summary video 1112 corresponding to the girl of thumbnail 1115B, respectively.
[0202] The summary videos 1109 and 1112 can be created from the original video 1106 by first analyzing each of the images 130 that are part of the original video 1106 to determine the images 130 that include the selected person. For example, the summary video 1109 can be created by first analyzing each of the images 130 that are part of the original video 1106 and include the man shown in the thumbnail 1115A. Similarly, the summary video 1112 can be created by first analyzing each of the images 130 that are part of the original video 1106 and include the girl shown in the thumbnail 1115B.
[0203] Next, the length of the summary video 1109 or 1112 can be determined. For example, the summary video 1109 or 1112 can be any length less than or equal to the length of the original video 1106. In some cases, the summary video 1109 or 1112 may not have a set maximum length. In this case, the summary video 1109 or 1112 can be the same as the original video 1106 when the selected person is included in each image 130 of the original video 1106.
[0204] In some embodiments, the summary video 1109 or 1112 may also include one or more transition images inserted between one or more images 130 included in the summary video 1109 or 1112. The transition images can be used to enrich and smooth the video. For example, the transition images may be other images selected from the video that include the same person, or a set of preset images used solely for transition purposes.
[0205] When a maximum length (e.g., 10 seconds) is set for the summary videos 1109 and 1112, one or more images 130 including the selected person can be combined to create the summary video 1109 or 1112. In some cases, the selected person's time in the original video 1106 may be less than the maximum length set for the summary videos 1109 and 1112. When the combination of all images 130 including the selected person creates a summary video 1109 or 1112 of a length less than or equal to the maximum length, the summary video 1109 or 1112 includes all images 130 depicting the selected person.
[0206] When a summary video 1109 or 1112 greater than the maximum length is created by combining all images 130 of the selected person, the summary video 1109 or 1112 may include a subset of images 130 depicting the selected person. In one embodiment, a subset of images 130 included in the summary video 1109 or 1112 may be randomly selected from all images 130 depicting the selected person in the original video 1106. In one embodiment, a subset of images 130 included in the summary video 1109 or 1112 may be selected based on a total score 285 of images 130 depicting each of the images 130 of the selected person in the original video. In one embodiment, a combination of the above may be used. Figures 3 to 10 The total score of each of the images 130 is calculated in a similar manner to the description of the images 130.
[0207] In another embodiment, the total score 285 can be calculated solely based on the selected person, wherein the image is analyzed as a single-person portrait image 130, without considering the group attributes of the image 130 or the characteristic attributes 225, general attributes 270, location attributes 260, or action / behavioral attributes 266 of any other person in the image. In one embodiment, people in the image 130 other than the selected person can be considered as background features of the image; therefore, analysis can be based solely on objective features. Subjective characteristics of people in the image 130 other than the selected person can be considered in the total score 285 without evaluating them. The following is in conjunction with... Figure 13 A further example is given of calculating the total score of 285 for image 130 based solely on the selected person.
[0208] The thumbnail 1115A or 1115B depicted in the lower left corner of the cover images 1117B to 1117C of the summary videos 1109 or 1112 can indicate the selected person for a particular summary video 1109 or 1112. Figure 11 As shown, the summary video 1109 includes a thumbnail 1115A of the lower left corner of the cover image 1117B of the summary video 1109. Similarly, the summary video 1112 includes a thumbnail 1115B of the lower left corner of the cover image 1117C of the summary video 1112.
[0209] The total score 285 of the respective summary videos 1109 and 1112 can be selected from the images 130 that are part of the respective summary videos 1109 and 1112. Each of the cover images 1117B to 1117C of the summary videos 1109 and 1112 can be selected based on the total score 285. In one embodiment, the total score 285 of the summary video 1109 or 1112 can be calculated using a method similar to that described above for multiple portrait images 130 that are part of a video. In one embodiment, the total score 285 of the summary video 1109 or 1112 can be calculated by analyzing the images 130 in the summary video 1109 or 1112 as single-person portrait images 130 and ignoring all other people in the images 130 of the summary video 1109 or 1112 besides the selected person in the summary video 1109 or 1112.
[0210] The following is combined Figure 13 A further example is described where image 130 from the summary video 1109 or 1112 is analyzed as a single-person portrait image 130. In one embodiment, when a user selects image 130, original video 1106, or summary videos 1109 and 1112 from the album 1103, the user can be redirected to the information page corresponding to the selected video or image 130.
[0211] Figure 12 A schematic diagram of an information page 1200 for a summary video 1109 provided for various embodiments of this disclosure. As described above, the information page 1200 includes several types of details associated with the described video, in this case, the summary video 1109. Figure 12 As shown, the information page 1200 of the summary video 1109 includes a cover image 1117B, a description 1211, thumbnails 1115A to 1115E of the various people depicted in the summary video 1109, links to other summary videos 1206 and 1209, and corresponding descriptions 1217A and 1217B of the other summary videos 1206 and 1209. It should be understood that the information page 1200 may include... Figure 12 Additional information not shown in the image.
[0212] The cover image 1117B is displayed at the top of the information page 1200. As described above, the cover image 1117B can be selected based on image 130, which is part of the summary video 1109 with the highest total score of 285. As described above, the total score of 285 can be assigned to a set of multiple portrait images 130 that are part of the summary video 1109, or the total score of 285 can be assigned to a set of single portrait images 130 that are part of the summary video 1109.
[0213] The description 1211 includes data or information describing the summary video 1109. For example, the description 1211 may include the video's name, the video's length 1214A, the location where the summary video 1109 was recorded or received, a link to the original video 1106 on which the summary video 1109 is based, and / or other data or metadata associated with the summary video 1109. The thumbnails 1115A to 1115E may be portraits of various people shown in the summary video 1109.
[0214] In one embodiment, the avatars used for the thumbnails 1115A to 1115E are cropped images of the corresponding people directly obtained from one of the images 130 in the summary video 1109. In another embodiment, the user can pre-configure the avatars for the thumbnails 1115A to 1115E from other videos or images previously stored in the computing device 200.
[0215] In one embodiment, when there are greater than or equal to a threshold number of images 130 in the summary video 1109 depicting a particular person, thumbnails 1115A to 1115E of that person may be included only in the information page 1200. In one embodiment, there is no such threshold, and thumbnails 1115A to 1115E can be presented for each person in the summary video 1109. Thus, although Figure 12 Only five thumbnails 1115A to 1115E are shown, but it should be understood that the information page 1200 may include thumbnails 1115A to 1115E of any number of people. In one embodiment, a portion of the information page 1200 showing the thumbnails 1115A to 1115E may be used for horizontal left-right scrolling to access all the thumbnails 1115A to 1115E related to the summary video 1109. In another embodiment, a portion of the information page 1200 showing the thumbnails 1115A to 1115E may be used for vertical up-down scrolling to access all the thumbnails 1115A to 1115E related to the summary video 1109.
[0216] In one embodiment, each of these thumbnails 1115A to 1115E may be a link to another summary video 1109 or 1112, which focuses on the person corresponding to the thumbnail 1115A to 1115E. Figure 12 As shown, thumbnail 1115A refers to the male central figure or protagonist of the summary video 1109. Therefore, thumbnail 1115A may not be a link to any other summary video 1109 or 1112. However, thumbnails 1115B to 1115E may be links to other summary videos 1109 or 1112. For example, thumbnail 1115B shows a thumbnail image of the girl included in the summary video 1109. In one embodiment, when a user clicks on thumbnail 1115B, a new summary video 1109 or 1112 can be created. The new summary video 1109 or 1112 may include an image 130 from the original video 1106 that includes the girl shown in thumbnail 1115B, or an image 130 from the summary video 1109 that includes the girl shown in thumbnail 1115B.
[0217] Similarly, thumbnail 1115C shows a thumbnail image of the female student included in the summary video 1109. In one embodiment, when a user clicks on thumbnail 1115C, a new summary video 1109 or 1112 can be created. The new summary video 1109 or 1112 may include an image 130 from the original video 1106 including the female student shown in thumbnail 1115C, or an image 130 from the summary video 1109 including the female student shown in thumbnail 1115C.
[0218] The thumbnail 1115D shows a thumbnail image of the little boy included in the summary video 1109. In one embodiment, when a user clicks the thumbnail 1115D, a new summary video 1109 or 1112 can be created. The new summary video 1109 or 1112 may include an image 130 from the original video 1106 including the little boy shown in the thumbnail 1115D, or an image 130 from the summary video 1109 including the little boy shown in the thumbnail 1115D.
[0219] The thumbnail 1115E shows a thumbnail image of the woman included in the summary video 1109. In one embodiment, when a user clicks on the thumbnail 1115E, a new summary video 1109 or 1112 can be created. The new summary video 1109 or 1112 may include an image 130 from the original video 1106 including the woman shown in the thumbnail 1115E, or an image 130 from the summary video 1109 including the woman shown in the thumbnail 1115E.
[0220] The bottom of the information page 1200 shows links to other summary videos 1206 and 1209. These links to other summary videos 1206 and 1209 can also be associated with a selected person in summary video 1109, but can have different lengths 1214A and 1214B, and are therefore different videos. For example, summary video 1206 has a length 1214A of 10 seconds, and summary video 1209 has a length 1214B of 56 seconds, which is sufficient to include all images 130 from the original video 1106, which includes the man shown in the thumbnail 1115A (e.g., no maximum length was set for the video).
[0221] The summary video 1206 with the maximum length may include images 130 from the original video 1106, the total score 285 of which is greater than a threshold score. The threshold score may be preset by the user and vary periodically, or pre-configured by a computing device. The total score 285 of each of the images 130 in the original video 1106 may be calculated based on whether the image 130 is a multi-person portrait image 130 or a single-person portrait image 130, wherein the single person being evaluated is the selected person. In the case of the summary video 1206, the single person being evaluated is the man shown in the thumbnail 1115A.
[0222] The information page 1200 can be displayed on the monitor of the computing device 200 in various ways. In one embodiment, when a user selects a video from the photo album 1103, the video itself can be displayed on the monitor of the computing device 200. An upward scrolling link can be displayed at the bottom of the screen displaying the video. The user can select the upward scrolling link to display the information page 1200. For example, the user can swipe up on the upward scrolling link to display the information page 1200 for a specific video. Similarly, any type of link associated with a video can exist, and the user can select the link to access the information page 1200 associated with that video.
[0223] Figure 13A schematic diagram of image 1300 provided for various embodiments of the present disclosure, showing a method for evaluating an image 130 of multiple people as a single-person portrait image 130. Figure 13 The image 1300 shown is similar to the cover image 1117B, but is referred to as image 1300, and is used to illustrate how image 1300 can be analyzed as a single-person portrait image 1300.
[0224] like Figure 13 As shown, image 1300 is a portrait image of multiple people, wherein multiple people 1303A to 1303G are the central focus of image 1300, rather than the background or landscape. Based on the above... Figures 3 to 10 The method for evaluating portrait images discussed herein assesses and scores the multi-person portrait image 1300 based on several different attributes 251 of each person shown in the multi-person portrait image. For example, when based on the above regarding Figures 3 to 10 The method for performing portrait image evaluation discussed herein is used to evaluate Figure 13 When the image 1300 shown is displayed, the group attribute 265 of all the people shown in the multi-person portrait image 1300 and the feature attribute 255, general attribute 270, location attribute 260 and action behavior attribute 266 of each person shown in the multi-person portrait image 1300 are used to create a total score 285.
[0225] In one embodiment, when a user selects a single person in image 1300 as the central focus or main subject of image 1300 (which may also be part of a video), image 1300 can be rated as a single-person portrait image 1300, even if image 1300 is a multi-person portrait image 1300. In this case, the evaluation of image 1300 may only involve the analysis of the selected person's characteristic attributes 255, general attributes 270, positional attributes 260, and action / behavior attributes 266. It is not necessary to consider the characteristic attributes 255, general attributes 270, positional attributes 260, and action / behavior attributes 266 of other people in image 1300 for rating. It is also not necessary to consider the group attribute 265 for rating. Instead, portions of image 1300 showing people other than the selected person can be considered as the background of image 1300. As described above, objective characteristics of the background of the image, such as brightness, contrast, saturation, sharpness, hue, and color, can be evaluated. The subjective characteristics of other people in the image 1300 are irrelevant to the total score 285 calculated for the image 1300.
[0226] Reference Figure 13Image 1300 shown is clearly a portrait image of seven people 1303A to 1303G. However, when a user selects one of these people as the central focus or protagonist of image 1300, image 1300 can be evaluated and rated as a single-person portrait image based on the selected person. The user can select person 1303A to 1303G as the protagonist of image 1300, thus selecting them as the focus of evaluation for image 1300. That is, the user only cares about the quality of the selected person in the image. The quality of the other people in image 1300 may be irrelevant.
[0227] For example, suppose the user selects person 1303A (corresponding to the man shown in thumbnail 1115A) as the focus of image 1300 to calculate the total score 285 of image 1300. In this case, the total score 285 can be a score 280 calculated based on the characteristic attributes 255, general attributes 270, location attributes 260, and / or action behavior attributes 266 of person 1303A selected for the user. As mentioned above... Figures 3 to 10 As described, the score 280 can be based on the region of interest of the selected person 1303A and on the training data 120.
[0228] In this case, no analysis or scoring is performed on the other individuals 1303B to 1303G, thus not contributing to the total score 285. The group attribute 265, which defines the arrangement and spatial relationships between individuals 1303A to 1303G, is also not considered. Instead, the portion of image 1300 showing other individuals 1303B to 1303G is considered the background of image 1300, and objective factors of the background of image 1300 are scored based on the training data 120 and factorized into the total score 285 of image 1300.
[0229] The total score 285 of image 1300 with the selected person 1303A as the center focus can be higher than the total score 285 of image 1300 with person 1303G as the center focus. Person 1303A shown in image 1300 has better feature attributes 255 and position attributes 260 than person 1303G. For example, person 1303A is focused on the camera and making eye contact with it, which can be associated with a high score 280 for the region of interest corresponding to person 1303A's face. Similarly, person 1303A's body is also facing the camera, which can also be associated with a high score 280 for the region of interest corresponding to person 1303A's body. The portrait image analysis module 125 can use the training data 120 to determine these high scores 280.
[0230] Conversely, the person 1303G is not facing the camera and is not focused on the camera, which can be associated with a lower score 280 corresponding to the region of interest of the person 1303G's face. The portrait image analysis module 125 can also use the training data 120 to determine the low score.
[0231] Thus, based on the selected individuals 1303A to 1303G, the same image 1300 can have different total scores 285. These different total scores 285 can be applied differently when creating summary videos 1109 and 1112 for the different selected individuals 1303A to 1303G. As mentioned above, the summary video 1109 with the longest length can be limited to images 1300 with a total score 285 higher than a threshold. Therefore, the summary video 1109 or 1112 with the longest length for the selected individual 1303A can include the image 1300 because the image 1300 with the center focus of individual 1303A has a higher total score 285. However, the summary video 1109 or 1112 with the longest length for the selected individual 1303G can exclude (e.g., exclude) the image 1300 because the image 1300 with the center focus of individual 1303G has a lower total score 285.
[0232] Figure 14 A flowchart of a method 1400 for performing portrait image analysis based on people depicted in an image, provided for various embodiments of this disclosure. The method 1400 may be executed by the portrait image analysis module 125 after obtaining an image 130 or 1300, which will be analyzed and scored based on a specific person who will become the central focus or protagonist of the image 130. In one embodiment, a selection of a person 1303A displayed in the image 130 or 1300 may be received. In one embodiment, the image 130 or 1300 may be a portrait image of multiple people. The processor 205 receives the selection of a person 1303A displayed in the image 130 or 1300. In step 1406, multiple attributes 251 are determined, each describing a multiple region of interest corresponding to body parts of the people displayed in the image 130 or 1300. The processor 205 executes the portrait image analysis module 125 to determine the attributes 251 in the image 130 or 1300 based on portions identified in the image 130. The identified attribute 251 may be a feature attribute 255, a location attribute 260, a general attribute 270, a group attribute 265, an action behavior attribute 266, an objective feature of the image 130, and / or other features describing the image 130.
[0233] In step 1409, a corresponding score 280 for each of the attributes 251 can be determined. The processor 205 executes the portrait image analysis module 125 to determine the corresponding score 280 for each of the attributes 251 based on predefined scores 252 stored in the training data 120. The predefined scores 252 are preset scores for various attributes based on the prototype image 123. In some embodiments, each of these scores 280 can be weighted according to a weight 290 assigned to each of the attributes 251 being scored. In one embodiment, the weight 290 for each of the attributes 251 is based on predefined weights 253 included in the training data 120.
[0234] In step 1412, a total score 285 is calculated based on the corresponding score 280 of attribute 251. The processor 205 executes the portrait image analysis module 125 to calculate the total score 285 based on the corresponding score 280 of attribute 251 and the background of the image including the plurality of other people 1303B to 1303G. The total score 285 may be a set of scores 280 of attribute 251 for the selected person 1303A. The total score 285 may also be a set of scores 280 weighted according to the weights 290 of attribute 251. The total score 285 may also be based on scores 280 assigned to the background of image 130 based on the training data 120, wherein the background of image 130 includes the plurality of other people 1303B to 1303G.
[0235] Figure 15 A flowchart of a method 1500 for creating summary videos 1109 or 112 based on selected persons, provided for various embodiments of this disclosure. The method 1400 may be executed by the portrait image analysis module 125 after obtaining image 130 or 1300, which will be analyzed and scored based on a specific person who will become the central focus or protagonist of image 130.
[0236] The summary videos 1109 and 1112 can be created at any time the computing device 200 is turned on. In one embodiment, the summary videos 1109 and 1112 can be created based on a user's selection of a person 1303A displayed in the images 130 or 1300. In one embodiment, a selection of a person 1303A displayed in the images 130 or 1300 can be received; the image can be one of a plurality of images 130 included in the video. In this embodiment, the images 130 or 1300 can be portrait images of multiple people. The processor 205 receives the selection of a person 1303A displayed in the images 130 or 1300.
[0237] In one embodiment, the summary videos 1109 and 1112 can be automatically created as background activity for the computing device 200 without the user needing to select the person 1303A displayed in the image. In this embodiment, the processor 205 can be used to create the summary video 1109 in the background when the computing device 200 is charging or in an idle state (e.g., the screen is off and the user is not using the computing device 200). When the screen or display of the computing device 200 is on, a notification can be displayed on the display of the computing device 200 to indicate that the summary videos 1109 and 112 are being created or have been created.
[0238] In step 1503, one or more images from a plurality of images 130 or 1300 that include the person 1303A are determined from the video. One or more images from the plurality of images 130 or 1300 are determined based on a total score 285 for each of the plurality of images 130 or 1300. The total score 285 may be calculated based solely on attribute 251 of the selected person 1303A. The processor 205 may be used to determine one or more images from the plurality of images 130 or 1300 that include the person 1303A selected based on the total score 285 for each of the plurality of images 130 or 1300.
[0239] In step 1509, one or more images from the plurality of images 130 or 1300 may be combined to create a summary video 1109 or 1112 of the selected person 1303A. For example, the processor 205 may combine one or more images from the plurality of images 130 or 1300 to create a summary video 1109 or 1112 of the selected person 1303A. While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in a variety of other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be regarded as illustrative rather than restrictive and are not intended to be limited to the details given herein. Various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0240] In one embodiment, the apparatus includes: means for determining a plurality of attributes, each attribute describing a region of interest corresponding to a body part of a person shown in an image; means for determining a corresponding score for each of the plurality of attributes; and means for calculating a total score based on the corresponding scores of the plurality of attributes.
[0241] In one embodiment, the apparatus includes: means for determining one or more images from a plurality of images including a person in a video; and means for combining one or more images from the plurality of images including the person to create a summary video of the person.
[0242] Furthermore, without departing from the scope of this disclosure, the technologies, systems, subsystems, and methods described and illustrated as independent or separate in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or discussed as coupled may be directly coupled or indirectly coupled or communicated through interfaces, devices, or intermediate components in an electrical, mechanical, or other manner. Other examples of changes, substitutions, and modifications can be determined by those skilled in the art and may be exemplified without departing from the spirit and scope of this disclosure.
Claims
1. A method implemented using a computing device, characterized in that, The method includes: The computing device determines multiple attributes corresponding to multiple regions of interest (ROIs) of human body parts displayed in the image, with each ROI having a corresponding attribute; wherein, the multiple ROIs include a first ROI and a second ROI with a parent-child relationship, and the attributes include a location attribute, which describes the location of the first ROI and the second ROI; The computing device determines a corresponding score for each of the plurality of attributes; wherein the corresponding score includes scores for the location attribute of the first region of interest and the location attribute of the second region of interest; The computing device calculates the total score based on the corresponding scores of the multiple attributes; The image is one of a plurality of images included in the video, and the method further includes: The computing device determines one or more images from the plurality of images that include the person; The computing device combination includes one or more images from the plurality of images of the person to create a summary video of the person, wherein the plurality of images included in the summary video are selected based on the total score of each of the plurality of images, and the total score is calculated based on the attributes of the person.
2. The method according to claim 1, characterized in that, In response to receiving a selection of the person displayed in the image, the person is determined, and the image displays the person and several other people.
3. The method according to claim 1 or 2, characterized in that, The person is identified based on the face detected in the image.
4. The method according to claim 1 or 2, characterized in that, The corresponding score for each of the plurality of attributes is determined based on training data that includes multiple predefined scores for each of the plurality of attributes.
5. The method according to claim 4, characterized in that, The training data includes multiple mappings, which respectively map one of the multiple predefined scores to one of the multiple predefined attributes.
6. The method according to claim 1, characterized in that, The total score is calculated based on the person's general attributes and location attributes.
7. The method according to claim 1 or 2, characterized in that, When the image displays the person and multiple other people, the method further includes: the computing device determining a score for the background of the image, wherein the background of the image includes the multiple other people, and further calculating the total score based on the score of the background of the image.
8. The method according to claim 1 or 2, characterized in that, The method further includes: the computing device searching for regions of interest corresponding to different body parts of the person depicted in the image, based on the probability that the body part is located at a certain position within the image.
9. The method according to claim 8, characterized in that, Searching the region of interest includes searching the region of interest based on training data, wherein the training data includes predefined anchor points pointing to specific parts or points in the image.
10. The method according to claim 8, characterized in that, Searching for the region of interest includes searching for a region of interest corresponding to at least one of the person's eyes, nose, or mouth, based on the person's facial position in the image.
11. The method according to claim 1 or 2, characterized in that, Multiple predefined scores are stored for each of the multiple attributes. Determining the corresponding score for each of the multiple attributes includes searching in the training data for the predefined score corresponding to the attribute determined for the region of interest.
12. The method according to claim 1 or 2, characterized in that, The plurality of attributes includes a plurality of general attributes of the person depicted in the image, which describe the overall quality of the person.
13. The method according to claim 1 or 2, characterized in that, When the image depicts more than one person, the method further includes: the computing device determining a corresponding score for each of a plurality of group attributes, wherein the plurality of group attributes respectively describe at least one of the following: the relationship between the plurality of other people depicted in the image, the space between each of the plurality of other people depicted in the image, the pose of one or more of the plurality of other people depicted in the image, or the arrangement of the plurality of other people depicted in the image, and further calculating the total score based on the corresponding scores of the plurality of group attributes.
14. The method according to claim 1 or 2, characterized in that, The weights are associated with each of the plurality of attributes, wherein the weight of a corresponding attribute is applied to the corresponding score of the corresponding attribute to create a weighted score for the corresponding attribute, and the total score is calculated based on the set of each of the weighted scores for each of the corresponding attributes.
15. A method implemented using a computing device, characterized in that, The method includes: The computing device determines multiple attributes corresponding to multiple regions of interest (ROIs) in an image of a person in a video. Each ROI has a corresponding attribute. The multiple ROIs include a first ROI and a second ROI with a parent-child relationship. The attributes include a location attribute, which describes the location of the first ROI and the second ROI. The computing device determines a corresponding score for each of the plurality of attributes; wherein the corresponding score includes scores for the location attribute of the first region of interest and the location attribute of the second region of interest; The computing device calculates the total score of the image based on the corresponding scores of the multiple attributes; The computing device determines one or more images from a plurality of images including a person from the video based on the total score of the images; The computing device combination includes one or more images from the plurality of images of the person to create a summary video of the person, wherein the plurality of images included in the summary video are selected based on the total score of each of the plurality of images, and the total score is calculated based on the attributes of the person.
16. The method according to claim 15, characterized in that, The method also includes the computing device receiving a selection of the person.
17. The method according to claim 15 or 16, characterized in that, In response to detecting the person's face in one or more of the plurality of images, the person is identified.
18. The method according to claim 15, characterized in that, The multiple attributes of the person include multiple general attributes of the person, which describe the overall quality of the person.
19. The method according to claim 15 or 16, characterized in that, The plurality of images including the person are determined based on a first plurality of attributes and a second plurality of attributes, wherein each of the first plurality of attributes describes a region of interest corresponding to a body part of the person, and each of the second plurality of attributes has a lower weight than the first plurality of attributes.
20. The method according to claim 15 or 16, characterized in that, The method further includes: The computing device creates a thumbnail representing the summary video of the person, the thumbnail including an image showing the person's face; The computing device displays the thumbnail.
21. The method according to claim 15 or 16, characterized in that, One or more images of the person are combined by adding one or more transition images to one or more of the plurality of images.
22. The method according to claim 15 or 16, characterized in that, The summary video is automatically created as background activity on the computing device.
23. The method according to claim 15 or 16, characterized in that, The method includes: the computing device displaying a notification indicating that the summary video is being created or has been created.
24. An apparatus, characterized in that, The device includes: Memory, the memory including instructions; One or more processors, which communicate with the memory, execute the instructions to: Based on multiple regions of interest (ROIs) of human body parts displayed in the image, multiple attributes corresponding to the multiple ROIs are determined, with each ROI having a corresponding attribute; wherein, the multiple ROIs include a first ROI and a second ROI with a parent-child relationship, and the attributes include a location attribute, which is used to describe the location of the first ROI and the second ROI; Determine the corresponding score for each of the plurality of attributes; wherein the corresponding score includes the scores for the location attribute of the first region of interest and the location attribute of the second region of interest; The total score is calculated based on the corresponding scores of the multiple attributes; The image is one of a plurality of images included in the video; The computing device determines one or more images from the plurality of images that include the person; The computing device combination includes one or more images from the plurality of images of the person to create a summary video of the person, wherein the plurality of images included in the summary video are selected based on the total score of each of the plurality of images, and the total score is calculated based on the attributes of the person.
25. The apparatus according to claim 24, characterized in that, In response to receiving a selection of the person displayed in the image, the person is determined, and the image displays the person and several other people.
26. The apparatus according to claim 24 or 25, characterized in that, The person is identified based on the face detected in the image.
27. The apparatus according to claim 24 or 25, characterized in that, The corresponding score for each of the plurality of attributes is determined based on training data that includes multiple predefined scores for each of the plurality of attributes.
28. The apparatus according to claim 24 or 25, characterized in that, The one or more processors also execute the instructions to search for regions of interest corresponding to different body parts of the person depicted in the image, based on the probability that the body part is located at a certain position within the image.
29. The apparatus according to claim 24 or 25, characterized in that, When the image depicts more than one person, the one or more processors further execute the instructions to determine a corresponding score for each of a plurality of group attributes, wherein the plurality of group attributes describe at least one of the following: the relationship between the plurality of other people depicted in the image, the space between each of the plurality of other people depicted in the image, the pose of one or more of the plurality of other people depicted in the image, or the arrangement of the plurality of other people depicted in the image, and further calculate the total score based on the corresponding scores of the plurality of group attributes.
30. An apparatus, characterized in that, The device includes: Memory, the memory including instructions; One or more processors, which communicate with the memory, execute the instructions to: Based on multiple regions of interest (ROIs) in the video containing images of people, multiple attributes corresponding to the multiple ROIs are determined, with each ROI having a corresponding attribute; wherein, the multiple ROIs include a first ROI and a second ROI with a parent-child relationship, and the attributes include a location attribute, which is used to describe the location of the first ROI and the second ROI; Determine the corresponding score for each of the plurality of attributes; wherein the corresponding score includes the scores for the location attribute of the first region of interest and the location attribute of the second region of interest; The total score is calculated based on the corresponding scores of the aforementioned multiple attributes; Based on the total score of the images, one or more images including a person are determined from the video. The images are combined to create a summary video of the person, wherein the images included in the summary video are selected based on the total score of each of the multiple images, and the total score is calculated based on the person's attributes.
31. The apparatus according to claim 30, characterized in that, The one or more processors also execute the instructions to receive a selection of the person.
32. The apparatus according to claim 30 or 31, characterized in that, The one or more processors also execute the instructions to detect the face of the person in one or more of the plurality of images, and to determine the person in response to detecting the face of the person.
33. The apparatus according to claim 30 or 31, characterized in that, The one or more processors also execute the instructions to create a thumbnail representing the summary video of the person, the thumbnail including an image showing the person's face; and cause a display device to display the thumbnail.
34. The apparatus according to claim 30 or 31, characterized in that, The summary video is automatically created as background activity on the computing device.
35. The apparatus according to claim 30 or 31, characterized in that, The summary video is automatically created when the computing device is charging.
Citation Information
Patent Citations
Image processing apparatus, image management apparatus and image management method, and computer program
CN101790047A
Video character extraction method and device
CN107644213A