Shooting method and device and electronic equipment
By using skin tone reference images in multi-person photo scenes to determine the reference portrait shooting parameters, the problems of slow imaging and high power consumption caused by multi-person skin tone recognition are solved, and a faster and energy-saving imaging process is achieved.
Patent Information
- Application Number
- CN202510250710.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-10
AI Technical Summary
In multi-person photo scenes, the prior art has increased power consumption and time-consuming for multi-person skin color recognition, resulting in slow imaging and high imaging power consumption.
By determining the skin color reference image in the shooting preview interface, determining the reference portrait shooting parameters of multiple faces based on the image, and shooting according to these parameters, the portrait image is output.
It reduces the duration and power consumption of skin tone recognition, improves the imaging speed in multi-person group photo scenes, and reduces imaging power consumption.
Smart Images

Figure CN120128811A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of imaging technology, and particularly relates to a shooting method, device, and electronic device. Background Art
[0002] With the wide application of electronic devices such as smart phones and tablet computers, the shooting functions of electronic devices are becoming increasingly rich. For example, when shooting face images of people with different skin colors, in order to make the face images of people with different skin colors meet their respective best effects, skin color recognition technology is used to identify different skin color information, corresponding shooting parameters are selected according to different skin color information, and face images are shot based on the selected shooting parameters.
[0003] In related technologies, for the scenario of group photos of multiple people, it is necessary to perform skin color recognition on each person, identify the skin color information of each person, comprehensively select shooting parameters based on the identified skin color information of all people, and shoot face images of multiple people based on the selected shooting parameters. However, since skin color recognition of multiple people increases exponentially in terms of power consumption and time consumption compared to skin color recognition of a single person. For example, if the time consumption of single-person skin color recognition is 1 second, the time consumption of three-person skin color recognition is 3 seconds. Therefore, in the scenario of group photos of multiple people, related technologies have problems of slow imaging and high imaging power consumption. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a shooting method, device, and electronic device, which can improve the imaging speed and reduce the imaging power consumption in the scenario of group photos of multiple people.
[0005] In a first aspect, the embodiments of this application provide a shooting method, and the method includes:
[0006] Determine a skin color reference image according to N faces in the shooting preview interface; wherein, at least two of the N faces have different skin color levels, the skin color reference image includes a reference face, and the skin color level of the reference face is determined based on the skin color levels of the N faces, N≥2;
[0007] Determine the reference portrait shooting parameters for the N faces according to the skin color reference image;
[0008] Control the camera to shoot according to the reference portrait shooting parameters, and output a portrait image; wherein, the portrait image includes the N faces.
[0009] In a second aspect, the embodiments of this application provide a shooting device, and the device includes:
[0010] A processing module is configured to determine a skin color reference image based on N human faces in a shooting preview interface; wherein, at least two of the N human faces have different skin color levels, the skin color reference image includes a reference human face, and the skin color level of the reference human face is determined based on the skin color levels of the N human faces, N≥2; determine reference portrait shooting parameters for the N human faces according to the skin color reference image; control a camera to take a picture according to the reference portrait shooting parameters, and output a portrait image; wherein, the portrait image includes the N human faces.
[0011] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the shooting method described in the first aspect are implemented.
[0012] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the shooting method described in the first aspect are implemented.
[0013] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the shooting method described in the first aspect.
[0014] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the shooting method described in the first aspect.
[0015] In the embodiment of the present application, a skin color reference image is determined based on N human faces in a shooting preview interface; wherein, at least two of the N human faces have different skin color levels, the skin color reference image includes a reference human face, and the skin color level of the reference human face is determined based on the skin color levels of the N human faces, N≥2; reference portrait shooting parameters for the N human faces are determined according to the skin color reference image; a camera is controlled to take a picture according to the reference portrait shooting parameters, and a portrait image is output; wherein, the portrait image includes the N human faces. It can be seen that in the embodiment of the present application, for the scenario of a group photo, the N human faces in the shooting preview interface can be mapped to a single human face image, and skin color recognition is performed on the single human face image. Since there is only one human face in the above single human face image, only one skin color recognition of the human face in the single human face image is required to determine the portrait shooting parameters in the scenario of a group photo, reducing the duration of skin color recognition and the power consumption of skin color recognition, thereby reducing the duration of the entire shooting process and the power consumption of the entire shooting process, and being able to improve the imaging speed in the scenario of a group photo and reduce the imaging power consumption in the scenario of a group photo. Brief Description of the Drawings
[0016] Figure 1 is a flowchart of a shooting method provided by some embodiments of the present application;
[0017] Figure 2A is an example diagram of a shooting preview interface provided by some embodiments of the present application;
[0018] Figure 2B is an example diagram of a skin color reference image provided by some embodiments of the present application;
[0019] Figure 3A is an example diagram of a shooting preview interface provided by some embodiments of the present application;
[0020] Figure 3B is an example diagram of a skin color reference image provided by some embodiments of the present application;
[0021] Figure 4A is a structural example diagram of an image generation model provided by some embodiments of the present application;
[0022] Figure 4B is a flowchart of an implementation manner of step 1013 provided by some embodiments of the present application;
[0023] Figure 5A is a flowchart of the generation process of the image encoder of the image generation model provided by some embodiments of the present application;
[0024] Figure 5B is a structural example diagram of the first backbone network provided by some embodiments of the present application;
[0025] Figure 6 is a flowchart of the generation process of the denoiser of the image generation model provided by some embodiments of the present application;
[0026] Figure 7A is a flowchart of a shooting method provided by some embodiments of the present application;
[0027] Figure 7B is an example diagram of a shooting preview interface provided by some embodiments of the present application;
[0028] Figure 7C is an example diagram of a shooting preview interface provided by some embodiments of the present application;
[0029] Figure 8A is a flowchart of a shooting method provided by some embodiments of the present application;
[0030] Figure 8B is an example diagram of a shooting preview interface provided by some embodiments of the present application;
[0031] Figure 8C It is an example diagram of a shooting preview interface provided by some embodiments of the present application;
[0032] Figure 9 It is a structural block diagram of a shooting device provided by some embodiments of the present application;
[0033] Figure 10 It is a schematic structural diagram of an electronic device provided by some embodiments of the present application;
[0034] Figure 11 It is a schematic hardware structure diagram of an electronic device provided by some embodiments of the present application for implementing the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0036] The terms "first", "second", etc. in the specification of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0037] Next, in conjunction with the accompanying drawings, the shooting method provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0038] It should be noted that the shooting method provided by the embodiments of the present application is applicable to electronic devices. In practical applications, the electronic devices include but are not limited to: mobile terminals such as smart phones, tablet computers, or personal digital assistants. The embodiments of the present application do not make any limitations in this regard.
[0039] The photographing method provided by the embodiments of the present application can be applied to the scenario of group photos. In the scenario of group photos, the skin colors of at least two people are different. For example, a specific application scenario is that three people take a group photo together using the front camera of a mobile phone. Among them, one person has fair skin color, one person has off-white skin color, and one person has dark brown skin color. Another specific application scenario is that four people take a group photo together using the front camera of a mobile phone. Among them, one person has off-white skin color, one person has neutral light skin color, one person has neutral dark skin color, and one person has black skin color.
[0040] Figure 1 is a flowchart of a photographing method provided by some embodiments of the present application. As Figure 1 shown, the method at least includes the following steps: Step 101, Step 102, and Step 103.
[0041] In Step 101, according to N human faces in the shooting preview interface, a skin color reference image is determined; among the N human faces, at least two human faces have different skin color levels. The skin color reference image includes a reference human face, and the skin color level of the reference human face is determined based on the skin color levels of the N human faces, where N≥2.
[0042] In some embodiments of the present application, for more refined management, based on the international skin color standard, the skin color levels can be divided into six categories, namely: skin color level 0, skin color level 1, skin color level 2, skin color level 3, skin color level 4, and skin color level 5, a total of six different skin color levels, representing the change of skin color from light to dark. The classification standard does not depend on specific skin color values, but is a subjective judgment of the human eye on which level a person's skin color belongs to. Among them, skin color level 0 corresponds to fair skin color, skin color level 1 corresponds to off-white skin color, skin color level 2 corresponds to neutral light skin color, skin color level 3 corresponds to neutral dark skin color, skin color level 4 corresponds to dark brown skin color, and skin color level 5 corresponds to black skin color.
[0043] In some embodiments of the present application, for the international skin color standard, the skin color level is measured according to the Individual Type Angle (ITA) value to classify which category the skin color belongs to. Among them, the ITA value is calculated from the skin image in the CIELAB space. When the ITA value ≥ 50, it is classified as skin color level 0; when 25 ≤ ITA value < 50, it is classified as skin color level 1; when 0 ≤ ITA value < 25, it is classified as skin color level 2; when -25 ≤ ITA value < 0, it is classified as skin color level 3; when -50 ≤ ITA value < -25, it is classified as skin color level 4; when the ITA value < -50, it is classified as skin color level 5. The calculation formula of the ITA value is as follows formula (1):
[0044]
[0045] Among them, L and b in formula (1) are the numerical values of L and b of the skin color value in the CCIELAB color space, respectively.
[0046] In some embodiments of the present application, the skin color levels can be divided based on different rules or standards. For example, in a group photo scenario with a high requirement for skin color effect, the skin color levels are divided more finely; in a group photo scenario with a general requirement for skin color effect, the skin color levels are divided simply, with high flexibility and a wide range of applicable scenarios.
[0047] In some embodiments of the present application, the skin color reference image is a single-face image that averages the skin color information of N faces, and the skin color level of the reference face in the skin color reference image is also within the standard skin color range.
[0048] In some embodiments of the present application, the process of determining the skin color reference image according to the N faces in the shooting preview interface can be executed locally on the electronic device or on the cloud. If it is executed on the cloud, the shooting time and power consumption of the electronic device can be further saved.
[0049] Exemplarily, in a group photo scenario, three people use the front camera of the mobile phone to take a group photo together. As Figure 2A shown, three faces are displayed in the shooting preview interface 21 of the mobile phone 20, namely the face 221 with fair skin color, the face 222 with slightly fair skin color, and the face 223 with dark brown skin color. Among them, Figure 2A only the appearances of the above three faces are shown, and the specific skin colors are not shown. As Figure 2B shown, according to Figure 2A the three faces in Figure 2B a skin color reference image 23 is determined. The skin color reference image 23 contains a reference face 231, and the skin color of the reference face 231 is neutral and slightly dark. Among them,
[0050] Figure 3A only the appearance of the reference face 231 is shown, and the specific skin color is not shown. Exemplarily, in a group photo scenario, four people use the front camera of the mobile phone to take a group photo together. As Figure 3A shown, four faces are displayed in the shooting preview interface 31 of the mobile phone 30, namely the face 321 with slightly fair skin color, the face 322 with neutral and slightly light skin color, the face 323 with neutral and slightly dark skin color, and the face 324 with black skin color. Among them, Figure 3B Figure 3A Figure 3AAmong the four skin color reference images 33 of the human faces, a reference human face 331 is included in the skin color reference image 33, and the skin color of the reference human face 331 is dark brown skin color. Among them, Figure 3B only the appearance of the reference human face 331 is shown, and the specific skin color is not shown.
[0051] In step 102, according to the skin color reference image, the reference portrait shooting parameters of N human faces are determined.
[0052] In the embodiments of the present application, since the skin color level of the reference human face in the skin color reference image is determined based on the skin color levels of N human faces, and the skin color level of the reference human face synthesizes the information of the skin color levels of N human faces, therefore, according to the skin color reference image, the reference portrait shooting parameters suitable for the skin color levels of N human faces can be accurately determined.
[0053] In some embodiments of the present application, a traditional skin color recognition method can be used to perform skin color recognition on the reference human face in the skin color reference image, and according to the reference skin color level of the reference human face, the reference portrait shooting parameters of N human faces are determined; among them, different skin color levels correspond to different portrait shooting parameters.
[0054] In some embodiments of the present application, the skin color reference image can be input into a skin color recognition network for skin color recognition to obtain a reference skin color level; according to the reference skin color level, the reference portrait shooting parameters of N human faces are determined; among them, different skin color levels correspond to different portrait shooting parameters.
[0055] In some embodiments of the present application, the training process of the skin color recognition network may include the following steps: constructing a training set, which includes multiple groups of training data, and each group of training data includes a skin color sample image and a corresponding true skin color level label; creating an AI classification backbone network; inputting the skin color sample image in each group of training data into the AI classification backbone network for processing, and outputting a predicted skin color level label; calculating each loss value according to the true skin color level label and the predicted skin color level label of each skin color sample image; according to each loss value, updating the network parameters of the AI classification backbone network in a backpropagation manner until the AI classification backbone network converges; determining the converged AI classification backbone network as the skin color recognition network.
[0056] In some embodiments of the present application, on the one hand, since the skin color recognition network is trained based on a large number of skin color sample images and can learn the prior knowledge of the skin color level information of skin color images, it can accurately identify the reference skin color level of the reference face in the skin color reference image. By selecting the reference portrait shooting parameters according to the accurate reference skin color level, the best reference portrait shooting parameters for N faces can be selected, and a group photo image of N faces with the best shooting effect can be obtained based on the best reference portrait shooting parameters. On the other hand, during the entire shooting process, the skin color recognition network is only used for skin color recognition once, which greatly saves the shooting time and power consumption.
[0057] In some embodiments of the present application, the portrait shooting parameters are shooting parameters related to the skin color level of the face, and the portrait shooting parameters may include at least one of the following: exposure parameters, white balance parameters, saturation, brightness, contrast, color temperature, and color.
[0058] In some embodiments of the present application, the exposure parameter is a crucial element in photography, which determines the exposure effect, that is, the brightness and clarity. Among them, the exposure parameters mainly include: aperture, shutter speed, and ISO.
[0059] In some embodiments of the present application, the white balance parameter is a setting used in photography and videography to correct the color temperature deviation of the image. The purpose is to ensure that white looks real and neutral in the captured image, so that other colors present accurate hues, thereby ensuring the color balance of the entire image.
[0060] In some embodiments of the present application, saturation refers to the vividness of color, also known as purity. It reflects the proportion of gray components in the color. The higher the saturation, the more vivid the color; the lower the saturation, the closer the color is to gray. In photography and image processing, by adjusting the saturation, the vividness of the image can be changed, thereby affecting the overall visual effect of the image.
[0061] In some embodiments of the present application, brightness refers to the ratio of the luminous intensity of a light-emitting body to the area of the light source, defined as the brightness of the unit of this light source, that is, the luminous intensity per unit projected area. It represents the brightness and darkness of the color. The higher the brightness, the brighter the image; the lower the brightness, the darker the image. In photography, videography, and image processing, brightness is one of the key parameters for adjusting the brightness and darkness of the image. By adjusting the brightness, the exposure effect of the image can be improved, making the image clearer and brighter.
[0062] In some embodiments of the present application, contrast refers to the measurement of different brightness levels between the brightest white and the darkest black in the light and dark areas of an image. It reflects the degree of difference between the light and dark parts of the image. The larger the difference range, the higher the contrast; the smaller the difference range, the lower the contrast. In photography and image processing, contrast is one of the important factors affecting image clarity and visual effects. By adjusting the contrast, the sense of hierarchy of the image can be enhanced, making the image more vivid and three-dimensional.
[0063] In some embodiments of the present application, color temperature is a physical quantity used in lighting optics to define the color of a light source. It represents the heating temperature of a black body when the color of the light emitted by the light source is the same as the color of the light emitted by the black body heated to a certain temperature. The lower the color temperature, the warmer the light (red, orange, yellow); the higher the color temperature, the colder the light (blue, purple). In photography, videography, and image processing, color temperature is one of the important factors affecting image color. By adjusting the color temperature, the color temperature deviation in the image can be corrected, making the image color more realistic and natural.
[0064] In some embodiments of the present application, color is the visual psychological feeling caused by light acting on the human eye, and it is an important attribute for describing the image of things. Color consists of three basic attributes: hue, lightness (brightness), and saturation. In the field of videography, by skillfully using color matching and contrast, the expressiveness and appeal of the work can be enhanced.
[0065] In some embodiments of the present application, a mapping relationship between skin color levels and shooting parameters can be pre-constructed. For example, when the skin color levels are divided into six categories, 6 different skin color levels correspond to 6 groups of different portrait shooting parameters. Then, according to this mapping relationship and the reference skin color level, the reference portrait shooting parameters are determined to make the skin color shooting effect of group photos of multiple people with different skin color levels better.
[0066] Exemplarily, taking the automatic exposure parameter as an example, from fair skin color to dark skin color, the value of the automatic exposure parameter gradually decreases. For example, the value of the automatic exposure parameter corresponding to fair skin color is 45, the value of the automatic exposure parameter corresponding to slightly fair skin color is 45, the value of the automatic exposure parameter corresponding to medium-light skin color is 38, the value of the automatic exposure parameter corresponding to medium-dark skin color is 38, the value of the automatic exposure parameter corresponding to dark brown skin color is 30, and the value of the automatic exposure parameter corresponding to black skin color is 30.
[0067] Exemplarily, taking the white balance parameter as an example, from fair skin color to dark skin color, the value of the white balance parameter gradually decreases. For example, for fair skin color, the hue value of the automatic white balance parameter is 1.50 and the face weight value is 0.7; for slightly fair skin color, the hue value of the automatic white balance parameter is 1.50 and the face weight value is 0.7; for neutral light skin color, the hue value of the automatic white balance parameter is 1.45 and the face weight value is 0.6; for neutral dark skin color, the hue value of the automatic white balance parameter is 1.45 and the face weight value is 0.6; for dark brown skin color, the hue value of the automatic white balance parameter is 1.40 and the face weight value is 0.4; for black skin color, the hue value of the automatic white balance parameter is 1.40 and the face weight value is 0.4.
[0068] In some embodiments of the present application, the portrait shooting parameters may include shooting parameters in multiple dimensions. Different shooting parameters control the shooting effect of the portrait skin color in different dimensions, making the overall skin color shooting effect better.
[0069] Exemplarily, in a group photo scenario of multiple people, 3 people use the front camera of the mobile phone to take a group photo together. As Figure 2A shown, in the shooting preview interface 21 of the mobile phone 20, 3 human faces are displayed, namely the human face 221 with fair skin color, the human face 222 with slightly fair skin color, and the human face 223 with dark brown skin color. Among them, Figure 2A only shows the appearances of the above 3 human faces and does not show the specific skin color. As Figure 2B shown, according to Figure 2A the 3 human faces in it, a skin color reference image 23 is determined. The skin color reference image 23 contains a reference human face 231. Among them, Figure 2B only shows the appearance of the reference human face 231 and does not show the specific skin color. The skin color reference image 23 is input into the skin color recognition network for skin color recognition, and a neutral dark skin color is obtained. The portrait shooting parameters corresponding to the neutral dark skin color are determined as Figure 2A the reference portrait shooting parameters of the 3 human faces in it.
[0070] Exemplarily, in a group photo scenario of multiple people, 4 people use the front camera of the mobile phone to take a group photo together. As Figure 3A shown, in the shooting preview interface 31 of the mobile phone 30, 4 human faces are displayed, namely the human face 321 with slightly fair skin color, the human face 322 with neutral light skin color, the human face 323 with neutral dark skin color, and the human face 324 with black skin color. Among them, Figure 3A only shows the appearances of the above 4 human faces and does not show the specific skin color. As Figure 3B shown, according to Figure 3A the 4 human faces in it, a skin color reference image 33 is determined. The skin color reference image 33 contains a reference human face 331. Among them, Figure 3BOnly the appearance of the reference face 331 is shown, and the specific skin color is not shown. The skin color reference image 33 is input into the skin color recognition network for skin color recognition, and a dark brown skin color is obtained. The portrait shooting parameters corresponding to the dark brown skin color are determined as Figure 3A the reference portrait shooting parameters of the 4 faces in
[0071] In step 103, according to the reference portrait shooting parameters, control the camera to take pictures and output portrait images; among them, the portrait images include N faces.
[0072] Exemplarily, in a group photo scenario of multiple people, 3 people use the front camera of the mobile phone to take a group photo together, as Figure 2A shown, 3 faces are displayed in the shooting preview interface 21 of the mobile phone 20, namely the face 221 with fair skin color, the face 222 with slightly fair skin color, and the face 223 with dark brown skin color. Among them, Figure 2A only the appearances of the above 3 faces are shown, and the specific skin colors are not shown. As Figure 2B shown, according to Figure 2A the 3 faces in Figure 2B a skin color reference image 23 is determined. The skin color reference image 23 contains a reference face 231. Among them, Figure 2A only the appearance of the reference face 231 is shown, and the specific skin color is not shown. The skin color reference image 23 is input into the skin color recognition network for skin color recognition, and a neutral dark skin color is obtained. The portrait shooting parameters corresponding to the neutral dark skin color are determined as
[0073] the reference portrait shooting parameters of the 3 faces in Figure 3A Finally, according to the portrait shooting parameters corresponding to the neutral dark skin color, control the camera to take pictures and output the portrait images of the above 3 faces. Figure 3A Exemplarily, in a group photo scenario of multiple people, 4 people use the front camera of the mobile phone to take a group photo together, as Figure 3B shown, 4 faces are displayed in the shooting preview interface 31 of the mobile phone 30, namely the face 321 with slightly fair skin color, the face 322 with neutral light skin color, the face 323 with neutral dark skin color, and the face 324 with black skin color. Among them, Figure 3A only the appearances of the above 4 faces are shown, and the specific skin colors are not shown. As Figure 3B shown, according to Figure 3AThe reference portrait shooting parameters of the 4 human faces in []. Finally, according to the portrait shooting parameters corresponding to the dark brown skin color, control the camera to take pictures and output the portrait images of the above 4 human faces.
[0074] In some embodiments of the present application, according to N human faces in the shooting preview interface, a skin color reference image is determined; wherein, at least two of the N human faces have different skin color levels, the skin color reference image includes a reference human face, and the skin color level of the reference human face is determined based on the skin color levels of the N human faces, N≥2; according to the skin color reference image, the reference portrait shooting parameters of the N human faces are determined; according to the reference portrait shooting parameters, control the camera to take pictures and output portrait images; wherein, the portrait images include N human faces. It can be seen that in the embodiments of the present application, for the scenario of a group photo of multiple people, the N human faces in the shooting preview interface can be mapped to a single human face image, and skin color recognition is performed on the single human face image. Since only one human face is included in the above single human face image, it is only necessary to perform skin color recognition on the human face in the single human face image once to determine the portrait shooting parameters in the scenario of a group photo of multiple people, reducing the duration of skin color recognition and the power consumption of skin color recognition, thereby reducing the duration of the entire shooting process and the power consumption of the entire shooting process, and being able to improve the imaging speed in the scenario of a group photo of multiple people and reduce the imaging power consumption in the scenario of a group photo of multiple people.
[0075] In some embodiments of the present application, image comparison technology can be used to map the N human faces in the shooting preview interface to a skin color reference image. Correspondingly, step 101 can include the following steps: step 1011 and step 1012.
[0076] In step 1011, the original group photo image of the N human faces in the shooting preview interface and each skin color template image in the image library are respectively input into the image comparison model for skin color comparison, and the skin color comparison result is output; wherein, the original group photo image is the image taken by the camera according to the original portrait shooting parameters; the skin color comparison result is used to represent the matching degree between the skin color levels of the N human faces and the skin color levels of the human faces in each skin color template image; the image library includes at least two skin color template images.
[0077] Exemplarily, in a scenario of a group photo of multiple people, 3 people take a group photo together using the front camera of a mobile phone. As Figure 2A shown, 3 human faces are displayed in the shooting preview interface 21 of the mobile phone 20, namely the human face 221 with fair skin color, the human face 222 with slightly white skin color, and the human face 223 with dark brown skin color. Among them, Figure 2A only the appearances of the above 3 human faces are shown, and the specific skin colors are not shown. The camera takes the original group photo image of the above 3 human faces according to the original portrait shooting parameters.
[0078] Exemplarily, in a group photo scenario, four people use the front camera of their mobile phones to take a group photo together. As Figure 3A shown, four human faces are displayed in the shooting preview interface 31 of the mobile phone 30, namely the human face 321 with a fair skin tone, the human face 322 with a neutral light skin tone, the human face 323 with a neutral dark skin tone, and the human face 324 with a black skin tone. Among them, Figure 3A only the appearances of the above four human faces are shown, and the specific skin tones are not shown. The camera captures the original group photo image of the above four human faces according to the original portrait shooting parameters.
[0079] In some embodiments of the present application, the original portrait shooting parameters are the portrait shooting parameters that have not been adjusted, or the portrait shooting parameters defaulted by the system.
[0080] In some embodiments of the present application, an image library can be pre-constructed. The image library includes at least two skin tone template images, and each skin tone template image contains only one human face. In practical applications, the appearance of the human face in the skin tone template image can be related to at least one of the N human faces in the shooting preview interface, or can be completely unrelated.
[0081] In some embodiments of the present application, the image comparison model can extract the skin tone feature information of two images, compare the skin tone feature information of the two images, and output a skin tone comparison result.
[0082] In some embodiments of the present application, the training process of the image comparison model may include the following steps: constructing a training set, which includes multiple groups of training data. Each group of training data includes two face sample images and a true label for characterizing whether the skin tone levels of the two face sample images match; creating an AI comparison backbone network; inputting the two face sample images in each group of training data into the AI comparison backbone network for processing, and outputting a predicted label for each group of training data; calculating each loss value according to the true label and the predicted label of each group of training data; updating the network parameters of the AI comparison backbone network in a backpropagation manner according to each loss value until the AI comparison backbone network converges; and determining the converged AI comparison backbone network as the image comparison model.
[0083] In step 1012, according to the skin tone comparison result, the skin tone template image with the highest matching degree is determined as the skin tone reference image.
[0084] In some embodiments of the present application, since the higher the matching degree indicated by the skin tone comparison result, the more matching the skin tone levels of the two images are. Therefore, from the skin tone comparison results output by the image comparison model, the skin tone template image with the highest matching degree is selected and determined as the skin tone reference image.
[0085] In some embodiments of the present application, since the image comparison model is trained based on a large number of image pairs and can learn the prior information of skin color differences between two face images, when the original group photo image of N faces in the shooting preview interface and each skin color template image in the image library are respectively input into the image comparison model for skin color comparison, accurate skin color comparison results can be output. Based on the accurate skin color comparison results, the skin color reference image can be accurately determined from the image library.
[0086] In some embodiments provided by the present application, with the help of artificial intelligence technology, the N faces in the shooting preview interface can be mapped to skin color reference images. Correspondingly, step 101 may include the following steps: step 1013.
[0087] In step 1013, the original group photo image of N faces in the shooting preview interface is input into the image generation model to output a skin color reference image; wherein, the original group photo image is an image captured by the camera according to the original portrait shooting parameters; the image generation model is used to generate a single-person image with a single skin color level based on a group photo image including at least two skin color levels.
[0088] Exemplarily, in a group photo scenario, 3 people take a group photo together using the front camera of a mobile phone, as Figure 2A shown, in the shooting preview interface 21 of the mobile phone 20, 3 faces are displayed, namely the face 221 with fair skin color, the face 222 with slightly fair skin color, and the face 223 with dark brown skin color. Among them, Figure 2A only the appearances of the above 3 faces are shown, and the specific skin colors are not shown. The camera captures the original group photo image of the above 3 faces according to the original portrait shooting parameters. The original group photo image is input into the image generation model for processing, and the skin color reference image 23 as Figure 2B shown is output. The skin color reference image 23 includes a reference face 231, and the skin color of the reference face 231 is neutral and slightly dark. Among them, Figure 2B only the appearance of the reference face 231 is shown, and the specific skin color is not shown.
[0089] Exemplarily, in a group photo scenario, 4 people take a group photo together using the front camera of a mobile phone, as Figure 3A shown, in the shooting preview interface 31 of the mobile phone 30, 4 faces are displayed, namely the face 321 with slightly fair skin color, the face 322 with neutral and slightly light skin color, the face 323 with neutral and slightly dark skin color, and the face 324 with black skin color. Among them, Figure 3A only the appearances of the above 4 faces are shown, and the specific skin colors are not shown. The camera captures the original group photo image of the above 4 faces according to the original portrait shooting parameters. The original group photo image is input into the image generation model for processing, and output asFigure 3B The skin color reference image 33 shown, wherein the skin color reference image 33 contains a reference human face 331, and the skin color of the reference human face 331 is dark brown skin color. Among them, Figure 3B only the appearance of the reference human face 331 is shown, and the specific skin color is not shown.
[0090] In some embodiments of the present application, the image generation model is trained based on a large number of image pairs composed of group photos of multiple people - single human face images, and can fully learn the prior knowledge of converting multiple human face images into single human face images. Therefore, it can accurately generate the skin color reference images corresponding to N human faces in the shooting preview interface.
[0091] In some embodiments provided by the present application, as Figure 4A shown, the image generation model 40 may include: an image encoder 41, a text encoder 42, a denoiser 43, and an image decoder 44. Correspondingly, as Figure 4B shown, the above step 1013 may include the following steps: step 10131, step 10132, step 10133, and step 10134.
[0092] In step 10131, the original group photo image is input into the image encoder of the image generation model for encoding processing, and the original multi-person skin color feature vector is output; wherein, the original multi-person skin color feature vector is used to represent the skin color levels of N human faces in the original group photo image.
[0093] In some embodiments of the present application, the purpose of this step is to compress the image data into a low-dimensional code through the image encoder while retaining the key features of the image, thereby reducing the amount of data for image processing and improving the processing efficiency.
[0094] In step 10132, the reference prompt is input into the text encoder of the image generation model for encoding processing, and the text feature vector is output.
[0095] In some embodiments of the present application, in the fields of computer science and programming, especially in the context related to artificial intelligence and machine learning, a prompt refers to the input or instruction provided to an AI model to guide the model to generate a specific output. For example, when generating text, images, or audio, the user will provide a prompt to clarify the type and style of content that the model is expected to generate.
[0096] In some embodiments of the present application, the reference prompt is used to describe the number of human faces in the skin color reference image, and the reference prompt is usually a simple text, such as one, single, single person, one person, etc.
[0097] In some embodiments of the present application, since the denoiser can only process data in vector dimensions, it is necessary to input the reference prompt into the text encoder of the image generation model. Through this text encoder, the reference prompt is converted from the text dimension to the feature vector dimension, and the text feature vector of the reference prompt is used to guide the denoiser to generate a skin color reference image containing only one face.
[0098] In some embodiments of the present application, the purpose of this step is to convert the text into embeddings through the text encoder, and these embeddings are used to control the image generation process.
[0099] In step 10133, the original multi-person skin color feature vector, the text feature vector, and random noise are input into the denoiser of the image generation model for denoising processing, and the original single-person skin color feature vector is output; wherein, the original single-person skin color feature vector is used to represent the skin color level of the face in the skin color reference image.
[0100] In some embodiments of the present application, the original multi-person skin color feature vector and the text feature vector are used as the control conditions of the denoiser to guide the denoiser to gradually denoise the random noise according to the original multi-person group photo image and the reference prompt, and the original single-person skin color feature vector after denoising processing is obtained.
[0101] In step 10134, the original single-person skin color feature vector is input into the image decoder of the image generation model for decoding processing, and the skin color reference image is output.
[0102] In some embodiments of the present application, since the original single-person skin color feature vector after denoising processing is in the dimension of image feature vector and cannot be directly displayed as an image, it is necessary to use the image decoder to perform decoding processing on the original single-person skin color feature vector to convert the data in vector dimension into an image, that is, the skin color reference image.
[0103] In some embodiments of the present application, since the image generation model includes a denoiser, and the denoiser can generate ultra-high-quality images through the process of gradually denoising, through the cooperation of each network module in the image generation model, it is possible to accurately map the multi-person group photo image to a single-person image.
[0104] In some embodiments provided by the present application, as Figure 5A shown, the generation process of the image encoder of the image generation model at least includes the following steps: step 501, step 502, step 503, step 504, step 505, and step 506.
[0105] In step 501, a first backbone network is constructed; wherein, the first backbone network includes: two CLIP image encoders.
[0106] Contrastive Language–Image Pre-training (CLIP) is a method for multi-modal vision and text learning. The core idea of this method is to enable the model to understand and associate images with their related text descriptions through contrastive learning. In CLIP, images and texts are respectively encoded into vector forms, and then the cosine similarity between these vectors is calculated to measure the matching degree between them. During the training process, the model learns how to adjust the representations of these vectors so that the vectors related to the same image or text description are closer in the vector space, while the vectors unrelated to them are farther away. The model built based on CLIP has strong generalization ability and can perform effective matching and retrieval on unseen images and texts, which makes CLIP have broad application prospects in many fields such as cross-modal retrieval, image description generation, and visual question answering.
[0107] Considering the above characteristics of CLIP, in some embodiments of the present application, the CLIP image encoder is used to replace the CLIP text encoder to construct a first backbone network, and through the first backbone network, the multi-face group photo image features can be mapped to single-face images.
[0108] Exemplarily, as Figure 5B shown, the first backbone network 50 may include: a first CLIP image encoder 51 and a second CLIP image encoder 52. Among them, the first CLIP image encoder 51 is used to encode the multi-face group photo image, and the second CLIP image encoder 52 is used to encode the single-face image.
[0109] In step 502, a data set is constructed; where the data set includes at least two groups of image samples, and each group of image samples includes: at least two single-face sample images and at least two multi-face sample group photos taken by the camera according to the portrait shooting parameters corresponding to a skin color level, and the skin color levels of different groups of image samples are different.
[0110] In some embodiments of the present application, the skin color levels of at least two faces in each multi-face sample group photo image are the same.
[0111] Exemplarily, there are a total of 6 skin color levels, namely fair skin color, off-white skin color, neutral light skin color, neutral dark skin color, dark brown skin color, and black skin color. The above 6 skin color levels respectively correspond to 6 different groups of portrait shooting parameters. Taking the automatic exposure parameter as an example of the portrait shooting parameter, the value of the automatic exposure parameter corresponding to the fair skin color is 45, the value of the automatic exposure parameter corresponding to the off-white skin color is 45, the value of the automatic exposure parameter corresponding to the neutral light skin color is 38, the value of the automatic exposure parameter corresponding to the neutral dark skin color is 38, the value of the automatic exposure parameter corresponding to the dark brown skin color is 30, and the value of the automatic exposure parameter corresponding to the black skin color is 30. When the value of the automatic exposure parameter is 45, 5 single-face sample images and 5 multi-face sample group photos are taken. The skin color of the face in each single-face sample image is fair skin color, and the skin color of at least two faces in each multi-face sample group photo is fair skin color. When the value of the automatic exposure parameter is 45, 5 single-face sample images and 5 multi-face sample group photos are taken. The skin color of the face in each single-face sample image is off-white skin color, and the skin color of at least two faces in each multi-face sample group photo is off-white skin color. When the value of the automatic exposure parameter is 38, 5 single-face sample images and 5 multi-face sample group photos are taken. The skin color of the face in each single-face sample image is neutral light skin color, and the skin color of at least two faces in each multi-face sample group photo is neutral light skin color. When the value of the automatic exposure parameter is 38, 5 single-face sample images and 5 multi-face sample group photos are taken. The skin color of the face in each single-face sample image is neutral dark skin color, and the skin color of at least two faces in each multi-face sample group photo is neutral dark skin color. When the value of the automatic exposure parameter is 30, 5 single-face sample images and 5 multi-face sample group photos are taken. The skin color of the face in each single-face sample image is dark brown skin color, and the skin color of at least two faces in each multi-face sample group photo is dark brown skin color. When the value of the automatic exposure parameter is 30, 5 single-face sample images and 5 multi-face sample group photos are taken. The skin color of the face in each single-face sample image is black skin color, and the skin color of at least two faces in each multi-face sample group photo is black skin color.
[0112] In some embodiments of the present application, based on the data pairs of multiple multi-face sample group photos and single-face sample images under the shooting parameters of different skin color levels, the first backbone network is trained to form a mapping between the multi-face group photo image and the single-face image.
[0113] In step 503, each group photo image of multiple face samples in each set of image samples is input into a CLIP image encoder of a first backbone network for encoding processing, and a multi-person skin color sample feature vector of each group photo image of multiple face samples is output; wherein, the multi-person skin color sample feature vector is used to characterize the skin color levels of all faces in the group photo image of multiple face samples.
[0114] Exemplarily, as Figure 5B shown, the first backbone network 50 may include: a first CLIP image encoder 51 and a second CLIP image encoder 52. Each time, a group photo image of multiple face samples is selected and input into the first CLIP image encoder 51 for encoding processing to obtain the corresponding multi-person skin color sample feature vector.
[0115] In step 504, each single-face sample image in each set of image samples is input into another CLIP image encoder of the first backbone network for encoding processing, and a single-person skin color sample feature vector of each single-face sample image is output; wherein, the single-person skin color sample feature vector is used to characterize the skin color level of the face in the single-face sample image.
[0116] Exemplarily, as Figure 5B shown, the first backbone network 50 may include: a first CLIP image encoder 51 and a second CLIP image encoder 52. Each time, a group photo image of multiple face samples is selected and input into the first CLIP image encoder 51 for encoding processing to obtain the corresponding multi-person skin color sample feature vector. At the same time, a single-face sample image is selected and input into the second CLIP image encoder 52 for encoding processing to obtain the corresponding single-person skin color sample feature vector.
[0117] In step 505, according to the multi-person skin color sample feature vectors of each group photo image of multiple face samples and the single-person skin color sample feature vectors of each single-face sample image, the network parameters of the first backbone network are adjusted to obtain a converged first backbone network.
[0118] In some embodiments of the present application, according to each multi-person skin color sample feature vector and the corresponding single-person skin color sample feature vector, each loss value is calculated, and the network parameters of the first backbone network are updated according to each loss value until the first backbone network meets the convergence condition.
[0119] In some embodiments of the present application, the cosine similarity between the multi-person skin color sample feature vector and the corresponding single-person skin color sample feature vector can be calculated, and the cosine similarity is determined as the loss value. Through the backpropagation algorithm, the first backbone network is fine-tuned. By repeatedly inputting different image pairs of group photo images of multiple face samples - single-face sample images, a converged first backbone network is obtained.
[0120] In step 506, the CLIP image encoder in the first backbone network after convergence for processing the group photo image of multiple face samples is determined as the image encoder of the image generation model.
[0121] In some embodiments of the present application, during the training process of the first backbone network, the generated multi-person skin color sample feature vectors can be saved and used as the training data for the denoiser of the subsequent image generation model.
[0122] In some embodiments of the present application, since the first backbone network where the image encoder of the image generation model is located is trained based on a large number of image pairs of group photo images of multiple face samples and single face sample images, and can learn the prior knowledge of the mapping relationship between rich multi-person skin color sample feature vectors and corresponding single-person skin color sample feature vectors, the image encoder of the trained image generation model can accurately encode the original multi-person skin color feature vectors of the original group photo image.
[0123] In some embodiments of the present application, when the image encoder of the image generation model is a CLIP image encoder, the text encoder of the image generation model is a CLIP text encoder. Since the same type of image encoder and text encoder have a high degree of adaptability, the features extracted have better effects during the fusion process.
[0124] In some embodiments provided by the present application, as Figure 6 shown, the generation process of the denoiser of the image generation model may include the following steps: step 601, step 602, step 603, step 604, step 605, and step 606.
[0125] In step 601, a second backbone network is constructed; wherein, the second backbone network includes: a forward noise addition algorithm and a denoiser.
[0126] In some embodiments of the present application, the second backbone network is essentially a Diffusion Probabilistic Models (DPMs). DPMs are a type of generative model that defines a Markov chain of diffusion steps to slowly add random noise to the data and then learn the reverse diffusion process to gradually transform the Gaussian noise distribution into the target data. The application of the diffusion model is usually divided into two stages: the forward diffusion process (represented by a dotted line) and the reverse denoising process (represented by a solid line). The forward diffusion process of the diffusion model is to add Gaussian noise to an image x t |x t-1 ) step by step for each time step {1, 2, …, t - 1, t, …, T}, and successively obtain the noisy data x 0 , x 1 , x 2 , …, xt-1 , x t , …, x T , where the noisy data x T is the finally obtained random noise data. The reverse denoising process of the diffusion model is to denoise the noisy data x T step by step for each time step {T, T - 1, T - 2, …, 2, 1} through a denoiser, gradually restoring the original image x 0 , that is to say, the reverse denoising process is to gradually remove noise starting from a random noise until an image is generated.
[0127] In some embodiments of the present application, the forward diffusion process of the diffusion model is a process of gradually adding noise. The process of adding noise refers to gradually adding Gaussian noise to a real image, which conforms to the Markov assumption. The noise at each step is Gaussian noise. Therefore, the forward diffusion process belongs to a parameter-free model and does not require learning. The reverse denoising process of the diffusion model is to gradually denoise the image with added noise, thereby restoring the real image. The denoising process uses a neural network model, namely a denoiser, so learning is required.
[0128] In some embodiments of the present application, the denoiser is also called a denoising autoencoder and is a key component in the diffusion model. It can take damaged data as input and, through training, predict the undamaged original data as output. This process is similar to the reverse process of the diffusion model, that is, recovering the original data from the data with added noise. The process of learning to recover the original data from the noisy data is one of the key mechanisms for the diffusion model to successfully generate clear images from noise.
[0129] In step 602, each single-face sample image in each group of sample images is input into the VAE image encoder for encoding processing, and the latent single-person skin color feature vector of each single-face sample image is output; wherein, the latent single-person skin color feature vector is used to characterize the skin color level of the face in the single-face sample image.
[0130] The Variational AutoEncoder (VAE) is a generative model based on a probabilistic graphical model that combines the advantages of an autoencoder and a probabilistic generative model. The VAE model maps the input data to a low-dimensional latent space through an encoder and imposes a certain probability distribution, such as a Gaussian distribution, on this space. Then, the original data is reconstructed by sampling from this latent space through a decoder. In this way, the VAE can learn the latent representation of the data and generate new data that is similar to the original data but has a certain degree of diversity.
[0131] In some embodiments of the present application, the VAE image encoder is a deep learning model for image processing. Its main purpose is to compress image data into a low-dimensional code while retaining the key features of the image. Its working principle involves multiple layers, including convolutional layers, pooling layers, and fully connected layers. Among them, the convolutional layer extracts features by sliding the convolutional kernel over the input data, the pooling layer reduces the computational amount by downsampling the feature map, and the fully connected layer maps the extracted features to the final output categories. The characteristics of this model lie in its hierarchical structure, local perception, and weight sharing, making it particularly suitable for processing image data. The image encoder can be applied to tasks such as data dimensionality reduction and feature extraction. By training, the reconstructed data is made as close as possible to the original input data, thereby learning the important feature representations of the input data.
[0132] In step 603, through the forward noise addition algorithm of the second backbone network, noise is added to the latent single-person skin color feature vector of each single-face sample image to obtain the noisy feature vector of each latent single-person skin color feature vector.
[0133] In some embodiments of the present application, Gaussian random noise is added to the latent single-person skin color feature vector of each single-face sample image through the forward noise addition algorithm of the second backbone network to obtain the noisy feature vector of each latent single-person skin color feature vector.
[0134] In step 604, each noisy feature vector, the multi-person skin color sample feature vectors of the same skin color level corresponding to each noisy feature vector, and the text feature vector are input into the denoiser of the second backbone network for denoising processing, and the predicted noise corresponding to each noisy feature vector is output.
[0135] In some embodiments of the present application, the multi-person skin color sample feature vectors of the same skin color level corresponding to each noisy feature vector and the text feature vector are used as the denoising control conditions of the denoiser of the second backbone network to control the denoiser to gradually denoise the noisy feature vector according to the multi-face sample group photo image and the reference prompt words, and the predicted noise corresponding to the noisy feature vector is output.
[0136] In step 605, according to the predicted noise corresponding to each noisy feature vector and the real noise added by the forward noise addition algorithm, the network parameters of the denoiser of the second backbone network are adjusted to obtain the converged second backbone network.
[0137] In some embodiments of the present application, each loss value is calculated according to the predicted noise corresponding to each noisy feature vector and the real noise added by the forward noise addition algorithm, and the network parameters of the denoiser of the second backbone network are updated according to each loss value until the second backbone network meets the convergence condition.
[0138] In some embodiments of the present application, a relevant loss function of the diffusion model can be adopted to calculate the loss value between the predicted noise corresponding to each noisy feature vector and the real noise added by the forward noise addition algorithm. Through the backpropagation algorithm, the denoiser of the second backbone network is fine-tuned. By repeatedly inputting different noisy feature vectors, the multi-person skin color sample feature vectors and text feature vectors corresponding to the same skin color level of the noisy feature vector, the converged second backbone network is obtained.
[0139] In step 606, the denoiser of the converged second backbone network is determined as the denoiser of the image generation model.
[0140] In some embodiments of the present application, the denoiser of the second backbone network is trained based on each noisy feature vector, the multi-person skin color sample feature vectors corresponding to the same skin color level of each noisy feature vector, and the text feature vector, realizing the mapping of the combination of the feature encoding results of the multi-person face group photo image and the text feature encoding results of the reference prompt words to the single-face image, so that an accurate single-face image corresponding to the skin color level can be obtained by inputting the feature encoding results of the multi-person face group photo image, the text feature encoding results of the reference prompt words, and random noise into the denoiser of the image generation model for processing.
[0141] In some embodiments provided by the present application, as Figure 7A shown, for the provided shooting method, before step 103 of the embodiment shown in Figure 1 it, the following steps may further be included: step 701, step 702, step 703, step 704, and step 705.
[0142] In step 701, on the shooting preview interface, a skin color adjustment window is displayed; wherein, the skin color adjustment window includes a skin color reference image.
[0143] In some embodiments of the present application, the size of the skin color adjustment window is usually smaller than the size of the shooting preview interface, and the skin color adjustment window is usually located at the edge position of the shooting preview interface. For example, the skin color adjustment window is displayed at the upper left corner position of the shooting preview interface to avoid blocking the main picture content on the shooting preview interface.
[0144] In some embodiments of the present application, the display position of the skin color adjustment window can be customized and changed by the user. For example, the user long-presses the skin color adjustment window with a finger, and then can drag the skin color adjustment window to any position on the shooting preview interface, and the position where the finger is released is the final display position of the skin color adjustment window.
[0145] In some embodiments of the present application, the display size of the skin tone adjustment window can be customized and changed by the user. For example, the user can drag the lower right corner boundary of the skin tone adjustment window with a finger. Dragging the lower right corner boundary inward can reduce the display size of the skin tone adjustment window, and dragging the lower right corner boundary outward can increase the display size of the skin tone adjustment window.
[0146] In some embodiments of the present application, displaying a skin tone reference image within the skin tone adjustment window can facilitate the user to view the skin tone shooting effects of different portrait shooting parameters on the skin tone reference image.
[0147] In step 702, according to the reference portrait shooting parameters, update the skin tone of the reference face in the skin tone reference image in the skin tone adjustment window.
[0148] Exemplarily, in a group photo scenario of multiple people, three people use the front camera of the mobile phone to take a group photo together. As Figure 2A shown, three faces are displayed in the shooting preview interface 21 of the mobile phone 20, namely the face 221 with fair skin tone, the face 222 with slightly fair skin tone, and the face 223 with dark brown skin tone. Among them, Figure 2A only shows the appearances of the above three faces, and does not show the specific skin tone. As Figure 2B shown, according to Figure 2A the three faces in it, determine the skin tone reference image 23. The skin tone reference image 23 contains a reference face 231. Among them, Figure 2B only shows the appearance of the reference face 231, and does not show the specific skin tone. Input the skin tone reference image 23 into the skin tone recognition network for skin tone recognition, and obtain a neutral to dark skin tone. Determine the portrait shooting parameters corresponding to the neutral to dark skin tone as Figure 2A the reference portrait shooting parameters of the three faces in it. As Figure 7B shown, a skin tone adjustment window 25 is also displayed on the shooting preview interface 21. The skin tone reference image 23 is displayed within the skin tone adjustment window 25. The skin tone display effect of the skin tone reference image 23 is a neutral to dark skin tone. Among them, Figure 7B only the appearance of the face is shown, and the specific skin tone is not shown.
[0149] In some embodiments of the present application, before the camera actually takes a photo according to the reference portrait shooting parameters, the skin tone effect of the skin tone reference image under the reference portrait shooting parameters can be displayed within the skin tone adjustment window, so that the user can intuitively understand the skin tone shooting effect corresponding to the reference portrait shooting parameters.
[0150] In step 703, when the skin tone adjustment window includes a portrait shooting parameter adjustment control, receive a control adjustment input for the portrait shooting parameter adjustment control.
[0151] In some embodiments of the present application, the portrait shooting parameter adjustment control is used to adjust the values of portrait shooting parameters.
[0152] In some embodiments of the present application, the reference portrait shooting parameters are portrait shooting parameters automatically recommended by the electronic device based on the skin tone levels of N faces in the shooting preview interface. The user can agree to use the reference portrait shooting parameters for shooting by the camera, or can customize the portrait shooting parameters. If the user wants to customize the portrait shooting parameters, the user can make a control adjustment input to the portrait shooting parameter adjustment control within the skin tone adjustment window.
[0153] In some embodiments of the present application, the control adjustment input is used to operate the portrait shooting parameter adjustment control.
[0154] In some embodiments of the present application, the above control adjustment input can be a first operation. Exemplarily, the above control adjustment input includes but is not limited to: a sliding input, a click input by the user on the portrait shooting parameter adjustment control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be specifically determined according to actual usage requirements and are not limited in the embodiments of the present application. The above specific gesture can be any one of a click gesture, a sliding gesture, a dragging gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, a double click gesture. The above click input can be any one of a single click input, a double click input, a long press input. For example, the above control adjustment input can be: a left sliding input or a right sliding input by the user on the portrait shooting parameter adjustment control.
[0155] Exemplarily, in a group photo scenario of multiple people, 3 people use the front camera of the mobile phone to take a group photo together. As Figure 2A shown, 3 faces are displayed in the shooting preview interface 21 of the mobile phone 20, which are the face 221 with fair skin tone, the face 222 with slightly fair skin tone, and the face 223 with dark brown skin tone. Among them, Figure 2A only the appearances of the above 3 faces are shown, and the specific skin tones are not shown. As Figure 2B shown, according to Figure 2A the 3 faces in Figure 2B a skin tone reference image 23 is determined. The skin tone reference image 23 contains a reference face 231. Among them, Figure 2A only the appearance of the reference face 231 is shown, and the specific skin tone is not shown. The skin tone reference image 23 is input into the skin tone recognition network for skin tone recognition, and a neutral dark skin tone is obtained. The portrait shooting parameters corresponding to the neutral dark skin tone are determined as Figure 7CAs shown, a skin tone adjustment window 25 is displayed on the shooting preview interface 21. A skin tone reference image 23 and a portrait shooting parameter adjustment control 26 are displayed within the skin tone adjustment window 25. Among them, the portrait shooting parameter adjustment control 26 is a control with a triangular slider. For each pixel distance the slider moves to the left, the selected portrait shooting parameter changes from the portrait shooting parameter corresponding to the current skin tone level to the portrait shooting parameter corresponding to the previous skin tone level. For each pixel distance the slider moves to the right, the selected portrait shooting parameter changes from the portrait shooting parameter corresponding to the current skin tone level to the portrait shooting parameter corresponding to the next skin tone level. Among them, Figure 7C only the appearance of the face is shown, and the specific skin tone is not shown. The user can customize the portrait shooting parameters by swiping the slider on the portrait shooting parameter adjustment control 26 left and right with a finger. For example, when the user swipes the slider on the portrait shooting parameter adjustment control 26 to the left with a finger, it means the user selects the portrait shooting parameter corresponding to a light skin tone. If this portrait shooting parameter is adopted, the skin tone of the skin tone reference image 23 within the skin tone adjustment window 25 will become whiter. When the user swipes the slider on the portrait shooting parameter adjustment control 26 to the right with a finger, it means the user selects the portrait shooting parameter corresponding to a dark skin tone. If this portrait shooting parameter is adopted, the skin tone of the skin tone reference image 23 within the skin tone adjustment window 25 will become darker.
[0156] In step 704, in response to the control adjustment input, a custom portrait shooting parameter is determined.
[0157] In some embodiments of the present application, according to the stop position of the portrait shooting parameter adjustment control in response to the control adjustment input, a custom portrait shooting parameter is determined, where different positions on the portrait shooting parameter adjustment control correspond to portrait shooting parameters of different skin tone levels.
[0158] Exemplarily, Figure 7C the current portrait shooting parameter of the skin tone reference image 23 within the skin tone adjustment window 25 in is the portrait shooting parameter corresponding to a neutral to dark skin tone. The user swipes the slider on the portrait shooting parameter adjustment control 26 to the right with a finger by one pixel distance, and the selected portrait shooting parameter is the portrait shooting parameter corresponding to a dark brown skin tone. The portrait shooting parameter corresponding to the dark brown skin tone is determined as the custom shooting parameter.
[0159] In step 705, according to the custom portrait shooting parameter, the skin tones of N faces in the shooting preview interface are updated.
[0160] Exemplarily, Figure 7CThe current portrait shooting parameters of the skin color reference image 23 in the skin color adjustment window 25 are the portrait shooting parameters corresponding to a neutral to dark skin color. The user slides the slider on the portrait shooting parameter adjustment control 26 one pixel distance to the right with a finger, and the selected portrait shooting parameters are the portrait shooting parameters corresponding to a dark brown skin color. The portrait shooting parameters corresponding to the dark brown skin color are determined as the custom shooting parameters. Then Figure 7C the skin colors of the human faces 221, 222, and 223 on the shooting preview interface 21 in
[0161] In some embodiments of the present application, before the camera actually takes a picture, a skin color adjustment window can be displayed on the shooting preview interface. The user can operate the portrait shooting parameter adjustment control in the skin color adjustment window to view the skin color effects under the reference portrait shooting parameters recommended by the electronic device and the skin color effects under the custom portrait shooting parameters, so that the user can intuitively understand the skin color shooting effects corresponding to different portrait shooting parameters, enabling the user to select the portrait shooting parameters of the skin color effect style that they like more, and meeting the user's personalized shooting needs for different skin color effects.
[0162] In some embodiments provided by the present application, as Figure 8A shown, for the provided shooting method, on the basis of the embodiments shown in Figure 8A the following steps can also be added: Step 801, Step 802, and Step 803.
[0163] In Step 801, a shooting parameter recommendation window is displayed; wherein, the shooting parameter recommendation window includes reference portrait shooting parameters and custom portrait shooting parameters.
[0164] In some embodiments of the present application, a shooting parameter recommendation window can also be displayed on the shooting preview interface, enabling the user to intuitively understand the reference portrait shooting parameters recommended by the electronic device and the custom portrait shooting parameters.
[0165] Exemplarily, in a group photo scenario of multiple people, 3 people use the front camera of the mobile phone to take a group photo together. As Figure 2A shown, three human faces are displayed in the shooting preview interface 21 of the mobile phone 20, namely the human face 221 with fair skin color, the human face 222 with slightly fair skin color, and the human face 223 with dark brown skin color. Among them, Figure 2A only the appearances of the above three human faces are shown, and the specific skin colors are not shown. As Figure 2B shown, according to Figure 2A the three human faces in Figure 2BOnly the appearance of the reference face 231 is shown, and the specific skin color is not shown. The skin color reference image 23 is input into the skin color recognition network for skin color recognition, and a medium-dark skin color is obtained. The portrait shooting parameters corresponding to the medium-dark skin color are determined as Figure 2A the reference portrait shooting parameters of the three faces in Figure 7C In [reference numeral], the current portrait shooting parameters of the skin color reference image 23 in the skin color adjustment window 25 are the portrait shooting parameters corresponding to the medium-dark skin color. The user slides the slider on the portrait shooting parameter adjustment control 26 one pixel distance to the right with a finger, and the selected portrait shooting parameters are the portrait shooting parameters corresponding to the dark brown skin color. The portrait shooting parameters corresponding to the dark brown skin color are determined as the custom shooting parameters. As Figure 8B shown, a shooting parameter recommendation window 27 is also displayed on the shooting preview interface 21. The reference portrait shooting parameters and the custom portrait shooting parameters are displayed in the shooting parameter recommendation window 27. Among them, the reference portrait shooting parameters include: the automatic exposure parameter value is 38, the hue value of the automatic white balance parameter is 1.45, and the face weight value is 0.6; the custom portrait shooting parameters include: the automatic exposure parameter value is 30, the hue value of the automatic white balance parameter is 1.40, and the face weight value is 0.4.
[0166] In step 802, a selection input for the reference portrait shooting parameters in the shooting parameter recommendation window is received.
[0167] In some embodiments of the present application, the selection input is used to select the reference portrait shooting parameters in the shooting parameter recommendation window.
[0168] In some embodiments of the present application, the above selection input may be a second operation. Exemplarily, the above selection input includes but is not limited to: a click input by the user on the reference portrait shooting parameters in the shooting parameter recommendation window using a finger or a stylus and other touch devices, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be specifically determined according to actual usage requirements, and the embodiments of the present application do not make limitations. The above specific gesture may be any one of a click gesture, a slide gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture. The above click input may be any one of a single click input, a double click input, and a long press input. For example, the above selection input may be: a single click input by the user on the reference portrait shooting parameters in the shooting parameter recommendation window.
[0169] In step 803, in response to the selection input, the skin colors of the N faces in the shooting preview interface are updated to the skin colors corresponding to the reference portrait shooting parameters.
[0170] Exemplarily, the reference portrait shooting parameters correspond to a medium-dark skin color, and the user clicks with a finger Figure 8BAfter the reference portrait shooting parameters in the shooting parameter recommendation window 27 in, the mobile phone 20 updates the skin colors of the faces 221, 222, and 223 in the shooting preview interface 21 to a neutral to dark skin color corresponding to the reference portrait shooting parameters.
[0171] In some embodiments of the present application, a shooting parameter recommendation window may be displayed on the shooting preview interface, enabling the user to intuitively understand the reference portrait shooting parameters recommended by the electronic device and the customized portrait shooting parameters. In addition, the user can also operate the reference portrait shooting parameters in the shooting parameter recommendation window to quickly update the skin colors of N faces in the shooting preview interface to the skin colors corresponding to the reference portrait shooting parameters, improving the convenience of shooting operations.
[0172] In some embodiments provided by the present application, some auxiliary controls may also be displayed to help the user quickly set portrait shooting parameters.
[0173] Exemplarily, as Figure 8C shown, an auxiliary control area 28 is displayed at a position slightly below the middle of the shooting preview interface 21 of the electronic device 20. Among them, the auxiliary control area 28 includes: a custom control, a hide skin color frame control, a display skin color frame control, a reset parameter control, and an Exit control. The custom control is used to trigger the display of the portrait shooting parameter adjustment control 26. The hide skin color frame control is used to trigger the skin color adjustment window 25. The display skin color frame control is used to trigger the display of the skin color adjustment window 25. The reset parameter control is used to trigger the restoration of the currently modified portrait shooting parameters to the original portrait shooting parameters. The Exit control is used to save the currently set portrait shooting parameters and exit this interface. Subsequently, the portrait shooting parameters of the subsequent camera will be based on the current settings; if the original portrait shooting parameters are to be used, the reset parameter control can be used to select to reset the portrait shooting parameters, and then it can be restored to the original portrait shooting parameters.
[0174] In some embodiments provided by the present application, in a single-person shooting scenario, a skin color adjustment window may also be displayed on the shooting preview interface. A thumbnail of the single-person face image is displayed in the skin color adjustment window. The user can view the adaptation relationship between the personal skin color level and the portrait shooting parameters through the thumbnail in the skin color adjustment window, and then adjust the skin color of the thumbnail in the skin color adjustment window according to personal preferences to meet the personalized skin color shooting requirements.
[0175] For the shooting method provided by the embodiments of the present application, the execution subject may be a shooting device. In the embodiments of the present application, taking the shooting device executing the shooting method as an example, the shooting device provided by the embodiments of the present application is described.
[0176] Figure 9 is a structural block diagram of a shooting device provided by some embodiments of the present application, as Figure 9As shown, the photographing device 900 may include: a processing module 901.
[0177] The processing module 901 is configured to determine a skin color reference image according to N human faces in a photographing preview interface; wherein, skin color levels of at least two of the N human faces are different, the skin color reference image includes a reference human face, and a skin color level of the reference human face is determined based on skin color levels of the N human faces, N≥2; determine reference portrait photographing parameters for the N human faces according to the skin color reference image; control a camera to perform photographing according to the reference portrait photographing parameters, and output a portrait image; wherein, the portrait image includes the N human faces.
[0178] In some embodiments of the present application, for a scenario of a group photo, N human faces in a photographing preview interface may be mapped to a single human face image, and skin color recognition is performed on the single human face image. Since only one human face is included in the above single human face image, it is only necessary to perform skin color recognition on the human face in the single human face image once, so that portrait photographing parameters in a group photo scenario can be determined, the duration of skin color recognition is reduced, the power consumption of skin color recognition is reduced, and further the duration of the entire photographing process is reduced, the power consumption of the entire photographing process is reduced, the imaging speed in a group photo scenario can be improved, and the imaging power consumption in a group photo scenario can be reduced.
[0179] Optionally, in some embodiments of the present application, the processing module 901 may specifically be configured to input the skin color reference image into a skin color recognition network for skin color recognition to obtain a reference skin color level; determine reference portrait photographing parameters for the N human faces according to the reference skin color level; wherein, different skin color levels correspond to different portrait photographing parameters.
[0180] Optionally, in some embodiments of the present application, the portrait photographing parameters may include at least one of the following: an exposure parameter, a white balance parameter, a saturation, a brightness, a contrast, a color temperature, and a color.
[0181] Optionally, in some embodiments of the present application, the processing module 901 may specifically be configured to input an original group photo image of N human faces in a photographing preview interface and each skin color template image in an image library into an image comparison model respectively for skin color comparison, and output a skin color comparison result; wherein, the original group photo image is an image photographed by the camera according to original portrait photographing parameters; the skin color comparison result is used to represent a matching degree between skin color levels of the N human faces and skin color levels of human faces in each skin color template image; the image library includes at least two skin color template images; and determine the skin color template image with the highest matching degree as the skin color reference image according to the skin color comparison result.
[0182] Optionally, in some embodiments of the present application, the processing module 901 may specifically be configured to input the original group photo image of N faces in the shooting preview interface into an image generation model to output a skin color reference image; wherein, the original group photo image is an image captured by the camera according to the original portrait shooting parameters; the image generation model is used to generate a single-person image with a single skin color level based on a group photo image including at least two skin color levels.
[0183] Optionally, in some embodiments of the present application, the image generation model may include: an image encoder, a text encoder, a denoiser, and an image decoder;
[0184] The processing module 901 may specifically be configured to input the original group photo image into the image encoder of the image generation model for encoding processing to output an original multi-person skin color feature vector; wherein, the original multi-person skin color feature vector is used to characterize the skin color levels of N faces in the original group photo image; input a reference prompt into the text encoder of the image generation model for encoding processing to output a text feature vector, wherein the reference prompt is used to describe the number of faces in the skin color reference image; input the original multi-person skin color feature vector, the text feature vector, and random noise into the denoiser of the image generation model for denoising processing to output an original single-person skin color feature vector; wherein, the original single-person skin color feature vector is used to characterize the skin color level of the face in the skin color reference image; input the original single-person skin color feature vector into the image decoder of the image generation model for decoding processing to output a skin color reference image.
[0185] Optionally, in some embodiments of the present application, the processing module 901 may further be configured to construct a first backbone network; wherein, the first backbone network includes: two contrastive language-image pre-trained CLIP image encoders; construct a dataset; wherein, the dataset includes at least two sets of image samples, and each set of image samples includes: at least two single-face sample images and at least two multi-face sample group photos taken by a camera according to the portrait shooting parameters corresponding to a skin color level, and the skin color levels of different sets of image samples are different; input each multi-face sample group photo image in each set of image samples into one CLIP image encoder of the first backbone network for encoding processing, and output a multi-person skin color sample feature vector of each multi-face sample group photo image; wherein, the multi-person skin color sample feature vector is used to characterize the skin color levels of all the faces in the multi-face sample group photo image; input each single-face sample image in each set of image samples into the other CLIP image encoder of the first backbone network for encoding processing, and output a single-person skin color sample feature vector of each single-face sample image; wherein, the single-person skin color sample feature vector is used to characterize the skin color level of the face in the single-face sample image; adjust the network parameters of the first backbone network according to the multi-person skin color sample feature vector of each multi-face sample group photo image and the single-person skin color sample feature vector of each single-face sample image to obtain a converged first backbone network; determine the CLIP image encoder in the converged first backbone network for processing the multi-face sample group photo image as the image encoder of the image generation model.
[0186] Optionally, in some embodiments of the present application, the text encoder of the image generation model is a CLIP text encoder.
[0187] Optionally, in some embodiments of the present application, the processing module 901 may further be configured to construct a second backbone network; wherein, the second backbone network includes: a forward noise addition algorithm and a denoiser; input each single-face sample image in each group of sample images into a variational autoencoder (VAE) image encoder for encoding processing, and output a latent single-person skin color feature vector of each single-face sample image; wherein, the latent single-person skin color feature vector is used to characterize the skin color level of the face in the single-face sample image; perform noise addition processing on the latent single-person skin color feature vector of each single-face sample image through the forward noise addition algorithm of the second backbone network to obtain a noisy feature vector of each latent single-person skin color feature vector; input each noisy feature vector, the multi-person skin color sample feature vectors corresponding to the same skin color level as each noisy feature vector, and the text feature vector into the denoiser of the second backbone network for denoising processing, and output the predicted noise corresponding to each noisy feature vector; adjust the network parameters of the denoiser of the second backbone network according to the predicted noise corresponding to each noisy feature vector and the real noise added by the forward noise addition algorithm to obtain a converged second backbone network; determine the denoiser of the converged second backbone network as the denoiser of the image generation model.
[0188] Optionally, in some embodiments of the present application, the photographing device 900 may further include:
[0189] a display module, configured to display a skin color adjustment window on the photographing preview interface; wherein, the skin color adjustment window includes the skin color reference image;
[0190] The processing module 901 may further be configured to update the skin color of the reference face of the skin color reference image in the skin color adjustment window according to the reference portrait photographing parameters;
[0191] a receiving module, configured to receive a control adjustment input for the portrait photographing parameter adjustment control when the skin color adjustment window includes the portrait photographing parameter adjustment control;
[0192] The processing module 901 may further be configured to determine custom portrait photographing parameters in response to the control adjustment input;
[0193] The display module is further configured to update the skin colors of N faces in the photographing preview interface according to the custom portrait photographing parameters.
[0194] Optionally, in some embodiments of the present application, the display module is further configured to display a photographing parameter recommendation window; wherein, the photographing parameter recommendation window includes the reference portrait photographing parameters and the custom portrait photographing parameters;
[0195] The receiving module is further configured to receive a selection input for the reference portrait shooting parameters in the shooting parameter recommendation window;
[0196] The display module is further configured to update the skin colors of N faces in the shooting preview interface to the skin colors corresponding to the reference portrait shooting parameters in response to the selection input.
[0197] The shooting device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an Augmented Reality (AR) / Virtual Reality (VR) device, a robot, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a Personal Computer (PC), a Television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0198] The shooting device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0199] The shooting device provided by the embodiments of the present application can implement Figure 1 , Figure 4B , Figure 5A , Figure 6 , Figure 7A and Figure 8A each process implemented by any of the method embodiments shown, achieving the same technical effects. To avoid repetition, it will not be elaborated here.
[0200] Optionally, as Figure 10As shown in the figure, an embodiment of the present application further provides an electronic device 1000, including a processor 1001 and a memory 1002. A program or instruction that can run on the processor 1001 is stored on the memory 1002. When the program or instruction is executed by the processor 1001, it implements each step of the above-mentioned shooting method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0201] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0202] Figure 11 It is a schematic diagram of the hardware structure of an electronic device provided by some embodiments of the present application.
[0203] The electronic device 1100 includes, but is not limited to: a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109, and a processor 1110 and other components.
[0204] Those skilled in the art can understand that the electronic device 1100 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The electronic device structure shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0205] In some embodiments of the present application, the processor 1110 is configured to determine a skin color reference image according to N human faces in the shooting preview interface; wherein, at least two of the N human faces have different skin color levels, the skin color reference image includes a reference human face, and the skin color level of the reference human face is determined based on the skin color levels of the N human faces, N≥2; determine reference portrait shooting parameters for the N human faces according to the skin color reference image; control the camera to take a picture according to the reference portrait shooting parameters, and output a portrait image; wherein, the portrait image includes the N human faces.
[0206] In some embodiments of the present application, for the scenario of group photos, N human faces in the shooting preview interface can be mapped into a single human face image, and skin color recognition is performed on the single human face image. Since there is only one human face in the above single human face image, only one skin color recognition needs to be performed on the human face in the single human face image to determine the portrait shooting parameters in the group photo scenario, reducing the duration of skin color recognition and the power consumption of skin color recognition, thereby reducing the duration of the entire shooting process and the power consumption of the entire shooting process, and being able to improve the imaging speed in the group photo scenario and reduce the imaging power consumption in the group photo scenario.
[0207] Optionally, in some embodiments of the present application, the processor 1110 is specifically configured to input the skin color reference image into a skin color recognition network for skin color recognition to obtain a reference skin color level; and determine the reference portrait shooting parameters of the N human faces according to the reference skin color level; wherein different skin color levels correspond to different portrait shooting parameters.
[0208] Optionally, in some embodiments of the present application, the portrait shooting parameters include at least one of the following: exposure parameter, white balance parameter, saturation, brightness, contrast, color temperature, color.
[0209] Optionally, in some embodiments of the present application, the processor 1110 is specifically configured to input the original group photo image of N human faces in the shooting preview interface and each skin color template image in the image library into an image comparison model respectively for skin color comparison, and output a skin color comparison result; wherein the original group photo image is an image captured by the camera according to the original portrait shooting parameters; the skin color comparison result is used to characterize the matching degree between the skin color levels of the N human faces and the skin color levels of the human faces in each skin color template image; the image library includes at least two skin color template images; and determine the skin color template image with the highest matching degree as the skin color reference image according to the skin color comparison result.
[0210] Optionally, in some embodiments of the present application, the processor 1110 is specifically configured to input the original group photo image of N human faces in the shooting preview interface into an image generation model, and output a skin color reference image; wherein the original group photo image is an image captured by the camera according to the original portrait shooting parameters; the image generation model is used to generate a single image with a single skin color level based on a group photo image including at least two skin color levels.
[0211] Optionally, in some embodiments of the present application, the image generation model includes: an image encoder, a text encoder, a denoiser, and an image decoder;
[0212] The processor 1110 is specifically configured to input the original group photo image into the image encoder of the image generation model for encoding processing, and output an original multi-person skin color feature vector; wherein, the original multi-person skin color feature vector is used to characterize the skin color levels of N faces in the original group photo image; input the reference prompt into the text encoder of the image generation model for encoding processing, and output a text feature vector, wherein the reference prompt is used to describe the number of faces in the skin color reference image; input the original multi-person skin color feature vector, the text feature vector, and random noise into the denoiser of the image generation model for denoising processing, and output an original single-person skin color feature vector; wherein, the original single-person skin color feature vector is used to characterize the skin color level of the face in the skin color reference image; input the original single-person skin color feature vector into the image decoder of the image generation model for decoding processing, and output a skin color reference image.
[0213] Optionally, in some embodiments of the present application, the processor 1110 is further configured to construct a first backbone network; wherein, the first backbone network includes: two contrastive language-image pre-training CLIP image encoders; construct a dataset; wherein, the dataset includes at least two groups of image samples, and each group of image samples includes: at least two single-face sample images and at least two multi-face sample group photos taken by the camera according to the portrait shooting parameters corresponding to a skin color level, and the skin color levels of different groups of image samples are different; input each multi-face sample group photo image in each group of image samples into one CLIP image encoder of the first backbone network for encoding processing, and output a multi-person skin color sample feature vector of each multi-face sample group photo image; wherein, the multi-person skin color sample feature vector is used to characterize the skin color levels of all faces in the multi-face sample group photo image; input each single-face sample image in each group of image samples into the other CLIP image encoder of the first backbone network for encoding processing, and output a single-person skin color sample feature vector of each single-face sample image; wherein, the single-person skin color sample feature vector is used to characterize the skin color level of the face in the single-face sample image; adjust the network parameters of the first backbone network according to the multi-person skin color sample feature vector of each multi-face sample group photo image and the single-person skin color sample feature vector of each single-face sample image to obtain a converged first backbone network; determine the CLIP image encoder for processing the multi-face sample group photo image in the converged first backbone network as the image encoder of the image generation model.
[0214] Optionally, in some embodiments of the present application, the text encoder of the image generation model is a CLIP text encoder.
[0215] Optionally, in some embodiments of the present application, the processor 1110 is further configured to construct a second backbone network; wherein, the second backbone network includes: a forward noise addition algorithm and a denoiser; input each single-face sample image in each group of sample images into a variational autoencoder (VAE) image encoder for encoding processing, and output a latent single-person skin color feature vector of each single-face sample image; wherein, the latent single-person skin color feature vector is used to characterize the skin color level of the face in the single-face sample image; perform noise addition processing on the latent single-person skin color feature vector of each single-face sample image through the forward noise addition algorithm of the second backbone network to obtain a noisy feature vector of each latent single-person skin color feature vector; input each noisy feature vector, the multi-person skin color sample feature vectors of the same skin color level corresponding to each noisy feature vector, and the text feature vector into the denoiser of the second backbone network for denoising processing, and output the predicted noise corresponding to each noisy feature vector; adjust the network parameters of the denoiser of the second backbone network according to the predicted noise corresponding to each noisy feature vector and the real noise added by the forward noise addition algorithm to obtain a converged second backbone network; determine the denoiser of the converged second backbone network as the denoiser of the image generation model.
[0216] Optionally, in some embodiments of the present application, the display unit 1106 is configured to display a skin color adjustment window on the shooting preview interface; wherein, the skin color adjustment window includes the skin color reference image; update the skin color of the reference face in the skin color reference image in the skin color adjustment window according to the reference portrait shooting parameters.
[0217] The user input unit 1307 is configured to receive a control adjustment input to the portrait shooting parameter adjustment control when the skin color adjustment window includes the portrait shooting parameter adjustment control.
[0218] The processor 1110 is further configured to determine custom portrait shooting parameters in response to the control adjustment input.
[0219] The display unit 1106 is further configured to update the skin colors of N faces in the shooting preview interface according to the custom portrait shooting parameters.
[0220] Optionally, in some embodiments of the present application, the display unit 1106 is further configured to display a shooting parameter recommendation window; wherein, the shooting parameter recommendation window includes the reference portrait shooting parameters and the custom portrait shooting parameters.
[0221] The user input unit 1107 is further configured to receive a selection input to the reference portrait shooting parameters in the shooting parameter recommendation window.
[0222] The display unit 1106 is further configured to update the skin colors of N faces in the captured preview interface to the skin colors corresponding to the reference portrait capture parameters in response to the selection input.
[0223] It should be understood that, in the embodiments of the present application, the input unit 1104 may include a Graphics Processing Unit (GPU) 11041 and a microphone 11042. The GPU 11041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also referred to as a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. The other input devices 11072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated herein.
[0224] The memory 1109 can be used to store software programs and various data. The memory 1109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1109 may include a volatile memory or a non-volatile memory, or the memory 1109 may include both a volatile and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchlink dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0225] The processor 1110 may include one or more processing units; in some embodiments, the processor 1110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1110 either.
[0226] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the shooting method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0227] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.
[0228] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned shooting method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0229] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0230] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned shooting method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0231] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0232] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium such as ROM / RAM, magnetic disk, or optical disc, and includes several instructions to enable a terminal such as a mobile phone, computer, server, or network device to execute the methods described in various embodiments of the present application.
[0233] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A shooting method, characterized in that: The method comprises: Determine a skin color reference image according to N faces in the shooting preview interface; wherein at least two of the N faces have different skin color levels, the skin color reference image includes a reference face, and the skin color level of the reference face is determined based on the skin color levels of the N faces, and N≥2; Determining reference portrait shooting parameters of the N faces according to the skin color reference image; According to the reference portrait shooting parameters, the camera is controlled to shoot and output a portrait image; wherein the portrait image includes the N faces.
2. The method according to claim 1, characterized in that The step of determining reference portrait shooting parameters of the N faces according to the skin color reference image includes: Inputting the skin color reference image into a skin color recognition network to perform skin color recognition to obtain a reference skin color grade; Determining reference portrait shooting parameters of the N faces according to the reference skin color level; Among them, different skin color levels correspond to different portrait shooting parameters.
3. The method according to claim 1 or 2, characterized in that: The portrait shooting parameters include at least one of the following: exposure parameter, white balance parameter, saturation, brightness, contrast, color temperature, and color.
4. The method according to claim 3, characterized in that The step of determining a skin color reference image according to N faces of persons in the shooting preview interface includes: Inputting the original multi-person photo image of N faces in the shooting preview interface and each skin color template image in the image library into the image comparison model respectively, performing skin color comparison, and outputting the skin color comparison result; wherein the original multi-person photo image is an image taken by the camera according to the original portrait shooting parameters; the skin color comparison result is used to characterize the matching degree between the skin color level of the N faces and the skin color level of the faces in each skin color template image; the image library includes at least two skin color template images; According to the skin color comparison result, the skin color template image with the highest matching degree is determined as the skin color reference image.
5. The method according to claim 3, characterized in that: The step of determining a skin color reference image according to N faces of persons in the shooting preview interface includes: Input the original group photo of N faces in the shooting preview interface into the image generation model, and output a skin color reference image; Among them, the original multi-person group photo image is an image taken by a camera according to original portrait shooting parameters; the image generation model is used to generate a single person image containing a single skin color level based on a multi-person group photo image containing at least two skin color levels.
6. The method according to claim 5, characterized in that The image generation model includes: an image encoder, a text encoder, a denoiser and an image decoder; The method of inputting the original group photo of N faces in the shooting preview interface into the image generation model and outputting the skin color reference image includes: Input the original multi-person group photo image into the image encoder of the image generation model, perform encoding processing, and output an original multi-person skin color feature vector; wherein the original multi-person skin color feature vector is used to characterize the skin color levels of N faces in the original multi-person group photo image; Inputting the reference prompt words into the text encoder of the image generation model, performing encoding processing, and outputting a text feature vector, wherein the reference prompt words are used to describe the number of faces in the skin color reference image; Inputting the original multi-person skin color feature vector, the text feature vector and random noise into the denoiser of the image generation model, performing denoising processing, and outputting the original single-person skin color feature vector; wherein the original single-person skin color feature vector is used to characterize the skin color level of the face in the skin color reference image; The original single-person skin color feature vector is input into the image decoder of the image generation model for decoding, and a skin color reference image is output.
7. The method according to claim 6, characterized in that Before the step of inputting the original group photo image of N faces in the shooting preview interface into the image generation model and outputting the skin color reference image, the method further includes: Constructing a first backbone network; wherein the first backbone network comprises: two contrasting language-image pre-trained CLIP image encoders; Constructing a data set; wherein the data set includes at least two groups of image samples, each group of image samples includes: at least two single face sample images and at least two multi-face sample group images taken by a camera according to a portrait shooting parameter corresponding to a skin color level, and the skin color levels of image samples in different groups are different; Input each multi-face sample group photo image in each group of image samples into a CLIP image encoder of the first backbone network for encoding processing, and output a multi-face skin color sample feature vector of each multi-face sample group photo image; wherein the multi-face skin color sample feature vector is used to characterize the skin color levels of all faces in the multi-face sample group photo image; Input each single face sample image in each group of image samples into another CLIP image encoder of the first backbone network for encoding processing, and output a single skin color sample feature vector of each single face sample image; wherein the single skin color sample feature vector is used to characterize the skin color level of the face in the single face sample image; According to the multi-person skin color sample feature vector of each multi-face sample group photo image and the single-person skin color sample feature vector of each single-person face sample image, the network parameters of the first backbone network are adjusted to obtain a converged first backbone network; The CLIP image encoder in the converged first backbone network for processing the group-photographed image of multiple face samples is determined as the image encoder of the image generation model.
8. The method according to claim 7, characterized in that The text encoder of the image generation model is the CLIP text encoder.
9. The method according to claim 7, characterized in that: After the step of determining the CLIP image encoder in the converged first backbone network for processing the multi-face sample group photo image as the image encoder of the image generation model, the method further includes: Constructing a second backbone network; wherein the second backbone network includes: a forward noise addition algorithm and a denoiser; Input each single face sample image in each group of sample images into a variational autoencoder (VAE) image encoder for encoding, and output a potential single person skin color feature vector for each single face sample image; wherein the potential single person skin color feature vector is used to characterize the skin color level of the face in the single face sample image; By using the forward noise adding algorithm of the second backbone network, noise is added to the potential single person skin color feature vector of each single face sample image to obtain a noisy feature vector of each potential single person skin color feature vector; Input each noisy feature vector, the feature vectors of multiple skin color samples at the same skin color level corresponding to each noisy feature vector, and the text feature vector into the denoiser of the second backbone network for denoising, and output the predicted noise corresponding to each noisy feature vector; According to the predicted noise corresponding to each of the noisy feature vectors and the real noise added by the forward noise addition algorithm, the network parameters of the denoiser of the second backbone network are adjusted to obtain a converged second backbone network; The denoiser of the converged second backbone network is determined as the denoiser of the image generation model.
10. The method according to claim 1, characterized in that Before the step of controlling the camera to shoot and outputting the portrait image according to the reference portrait shooting parameters, the method further includes: On the shooting preview interface, a skin color adjustment window is displayed; wherein the skin color adjustment window includes the skin color reference image; According to the reference portrait shooting parameters, updating the skin color of the reference face of the skin color reference image in the skin color adjustment window; In a case where the skin color adjustment window includes a portrait shooting parameter adjustment control, receiving a control adjustment input for the portrait shooting parameter adjustment control; In response to the control adjustment input, determining a custom portrait shooting parameter; According to the customized portrait shooting parameters, the skin colors of N faces in the shooting preview interface are updated.
11. The method according to claim 10, characterized in that The method further comprises: Displaying a shooting parameter recommendation window; wherein the shooting parameter recommendation window includes the reference portrait shooting parameters and the custom portrait shooting parameters; receiving a selection input of a reference portrait shooting parameter in the shooting parameter recommendation window; In response to the selection input, the skin colors of the N faces in the shooting preview interface are updated to skin colors corresponding to the reference portrait shooting parameters.
12. A photographing device, characterized in that: The device comprises: A processing module is used to determine a skin color reference image based on N human faces in a shooting preview interface; wherein at least two of the N human faces have different skin color levels, the skin color reference image includes a reference face, and the skin color level of the reference face is determined based on the skin color levels of the N human faces, N≥2; according to the skin color reference image, determine reference portrait shooting parameters of the N human faces; according to the reference portrait shooting parameters, control the camera to shoot and output a portrait image; wherein the portrait image includes the N human faces.
13. The device according to claim 12, characterized in that The processing module is specifically used to input the skin color reference image into a skin color recognition network for skin color recognition to obtain a reference skin color level; according to the reference skin color level, determine the reference portrait shooting parameters of the N faces; wherein different skin color levels correspond to different portrait shooting parameters; the portrait shooting parameters include at least one of the following: exposure parameter, white balance parameter, saturation, brightness, contrast, color temperature, and color.
14. The device according to claim 13, characterized in that The processing module is specifically used to input an original multi-person photo image of N faces in a shooting preview interface into an image generation model, and output a skin color reference image; wherein the original multi-person photo image is an image taken by a camera according to original portrait shooting parameters; and the image generation model is used to generate a single person image containing a single skin color level based on a multi-person photo image containing at least two skin color levels.
15. The device according to claim 14, characterized in that The image generation model includes: an image encoder, a text encoder, a denoiser and an image decoder; The processing module is specifically used to input the original multi-person photo image into the image encoder of the image generation model, perform encoding processing, and output an original multi-person skin color feature vector; wherein the original multi-person skin color feature vector is used to characterize the skin color levels of N faces in the original multi-person photo image; input a reference prompt word into the text encoder of the image generation model, perform encoding processing, and output a text feature vector, wherein the reference prompt word is used to describe the number of faces in the skin color reference image; input the original multi-person skin color feature vector, the text feature vector and random noise into the denoiser of the image generation model, perform denoising processing, and output an original single-person skin color feature vector; wherein the original single-person skin color feature vector is used to characterize the skin color level of the face in the skin color reference image; input the original single-person skin color feature vector into the image decoder of the image generation model, perform decoding processing, and output a skin color reference image.
16. The device according to claim 15, characterized in that The processing module is also used to construct a first backbone network; wherein the first backbone network includes: two contrast language-image pre-trained CLIP image encoders; construct a data set; wherein the data set includes at least two groups of image samples, each group of image samples includes: at least two single face sample images and at least two multi-face sample group photo images taken by a camera according to a portrait shooting parameter corresponding to a skin color level, and the skin color levels of image samples in different groups are different; each multi-face sample group photo image in each group of image samples is input into a CLIP image encoder of the first backbone network, and encoding processing is performed, and a multi-face skin color sample feature vector of each multi-face sample group photo image is output; wherein the multi-face skin color sample feature vector is used to characterize the multi-face sample group photo image. The invention relates to a method for generating a plurality of face samples of a plurality of faces in a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples are grouped together to obtain a plurality of face samples of a plurality of groups of images, wherein the plurality of face samples of a plurality of groups of images ...
17. The device according to claim 16, characterized in that The processing module is also used to construct a second backbone network; wherein the second backbone network includes: a forward denoising algorithm and a denoiser; each single-face sample image in each group of sample images is input into a variational autoencoder VAE image encoder for encoding processing, and a potential single-person skin colour feature vector of each single-face sample image is output; wherein the potential single-person skin colour feature vector is used to characterize the skin colour level of the face in the single-face sample image; through the forward denoising algorithm of the second backbone network, noise is added to the potential single-person skin colour feature vector of each single-face sample image to obtain each potential single-person skin colour The invention relates to a noisy feature vector of a feature vector; inputting each noisy feature vector, the feature vector of multiple skin color samples at the same skin color level corresponding to each noisy feature vector, and the text feature vector into the denoiser of the second backbone network for denoising, and outputting the predicted noise corresponding to each noisy feature vector; adjusting the network parameters of the denoiser of the second backbone network according to the predicted noise corresponding to each noisy feature vector and the real noise added by the forward denoising algorithm to obtain a converged second backbone network; determining the denoiser of the converged second backbone network as the denoiser of the image generation model.
18. The device according to claim 12, characterized in that The device also includes: A display module, used to display a skin color adjustment window on the shooting preview interface; wherein the skin color adjustment window includes the skin color reference image; The processing module is further used to update the skin color of the reference face of the skin color reference image in the skin color adjustment window according to the reference portrait shooting parameters; a receiving module, configured to receive a control adjustment input of the portrait shooting parameter adjustment control when the skin color adjustment window includes the portrait shooting parameter adjustment control; The processing module is further used to determine the custom portrait shooting parameters in response to the control adjustment input; The display module is further used to update the skin colors of N faces in the shooting preview interface according to the customized portrait shooting parameters.
19. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the shooting method according to any one of claims 1 to 11 are implemented.