Exposure parameter adjustment method and device and storage medium
By segmenting the real face area in the preview image and using an AI model to generate a mask image, the problem of obstructions affecting the accuracy of exposure parameters is solved, achieving more accurate exposure control and higher imaging quality.
Patent Information
- Application Number
- CN202411207691.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-08-29
AI Technical Summary
In the prior art, when a person's face is covered by obstructions such as a mask, glasses, beard or hair, the exposure parameters are not accurate enough, affecting the imaging quality.
By segmenting the real face part from the preview image, using the AI model to generate a mask image, eliminating the influence of obstructions and background, and calculating the exposure parameters based on the brightness of the real face area.
The accuracy of exposure parameters is improved, overexposure or underexposure is avoided, and imaging quality and user experience are enhanced.
Smart Images

Figure CN120751268A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an exposure parameter adjustment method, device, and storage medium. Background Art
[0002] With the development of smart terminals, the camera function has become a common function of electronic devices.
[0003] Currently, most electronic devices are equipped with a feature that automatically adjusts exposure parameters. After the camera is turned on, if there is a portrait in the preview video stream captured by the camera, the electronic device can calculate the overall brightness of the face area and then adjust the camera's exposure parameters based on this overall brightness to capture an image that meets the brightness requirements.
[0004] However, when the face of a person in the preview video stream is blocked by a mask, glasses, beard or hair, if the color of the obstruction is significantly different from the facial skin, it may affect the overall brightness of the facial area, resulting in lower accuracy of the exposure parameters calculated based on the overall brightness of the facial area, causing overexposure or underexposure, affecting the image quality. Summary of the Invention
[0005] The present application provides an exposure parameter adjustment method, device and storage medium, which solves the technical problem of inaccurate exposure parameters caused by the face being covered by obstructions such as masks, glasses, beards or hair.
[0006] To achieve the above objectives, this application adopts the following technical solutions:
[0007] In a first aspect, embodiments of the present application provide a method for adjusting exposure parameters. The method may include:
[0008] Displaying a first image based on a first data stream on a preview interface; cropping a first facial image from a second image based on the first data stream; obtaining a first mask image corresponding to the first facial image, the first mask image including at least data of a facial region not covered by an obstruction; determining the first facial region in a third image based on the data of the facial region not covered by an obstruction; obtaining a first exposure parameter based on the brightness of the first facial region; and controlling a camera to capture a second data stream based on the first exposure parameter. The first image is an RGB image converted from a RAW image captured at a first moment, the second image is a YUV image converted from a RAW image captured at a second moment, and the third image is a YUV image converted from a RAW image captured at a first moment. The first data stream is the data stream captured by the camera before adjustment to the first exposure parameter, and the second data stream is the data stream captured by the camera after adjustment to the first exposure parameter.
[0009] In the above solution, when displaying a preview image based on the preview stream, if a face image is detected in the preview image, the actual face portion can be segmented from the face image. Exposure parameters are then calculated based on the brightness of the actual face portion, and the camera's exposure parameters are then adjusted. By segmenting the actual face portion from the face image, the impact of facial obstructions such as masks, glasses, beards, or hair, as well as the background image, on the overall brightness calculation is reduced. This makes the exposure parameters calculated based on the average brightness of the actual face portion more accurate, resulting in a brightness that better meets user needs and avoids the problem of constant flickering in the captured image, improving the quality of the resulting image and enhancing the user's photography experience.
[0010] In one possible implementation, RGB and YUV are two standard color encoding methods. In RGB, each pixel has three colors: red, green, and blue. In YUV, Y represents brightness, while U and V represent chrominance. Brightness and chrominance are used to specify the color of a pixel. Display screens typically display images using the RGB model, but the YUV model is often used when processing image data because it saves more storage space.
[0011] In one possible implementation, cropping a first facial image from a second image based on a first data stream includes: obtaining facial frame information using a face detection algorithm, the facial frame information including the number of faces in the second image and the coordinates of each facial frame, where each facial frame corresponds to a single face; determining a second facial region based on the facial frame information, the second facial region being the facial frame region in the second image having a facial frame area greater than a preset size and having the largest facial frame area; and cropping the first facial image from the second image when the size of the second facial region is greater than or equal to the preset size, the first facial image being an image of the second facial region. The preset size may be a preset area, or may be a preset height and width. It will be appreciated that the size of each YUV image frame is typically fixed. If the face region with the largest face frame area in the i-th YUV image is smaller than the preset size, it means that the face occupies a small proportion in this YUV image and has little impact on the brightness of the entire image. It is not the shooting focus, so the brightness of this face region should not be used as the basis for adjusting the exposure parameters. If the face region with the largest face frame area in a YUV image is greater than or equal to the preset size, it means that the face occupies a large proportion in this YUV image and may be the shooting focus.
[0012] In one possible implementation, after cropping the first face image from the second image, the method further includes: converting the first face image into RGB format; and adjusting the first face image converted into RGB format to a preset size. It is understandable that since the AI model usually obtains the MSAK image based on the image in RGB format, the image of the face area needs to be converted from the YUV domain to the RGB domain before inputting it into the AI model. In addition, taking the preset size of 64 pixels * 64 pixels as an example, since the size of the face area with the largest face frame area may far exceed 64 pixels * 64 pixels, if it is directly input into the AI model, it may cause the AI model to have a large amount of computation. For example, in order to reduce the computation of the AI model, the MSAK module can first adjust the image of the face area to 64 pixels * 64 pixels, and then input it into the AI model.
[0013] In one possible implementation, obtaining a first mask image corresponding to a first face image includes: inputting the first face image into an AI model to obtain a first mask (MASK) image. The MASK image is a matrix with the same dimension as the original image (i.e., the YUV image of the i-th frame), and the elements in the MASK image determine whether to retain, modify, or block the image pixels at the corresponding position in the YUV image of the i-th frame. The MASK image can be used to select the real face area and filter the non-face area (such as the background image and the area covered by the occluded object). As an example, in the MSAK image of a real face, a pixel value of 0 indicates that the pixel is in the real face area, and a pixel value of 1 indicates that the pixel is in the non-face area. As another example, a pixel value of 1 indicates that the pixel is in the real face area, and a pixel value of 0 indicates that the pixel is in the non-face area. Understandably, if the color of the background image or the area covered by an occluded object differs significantly from the color of the actual face area, the calculated overall brightness of the face area may not be accurate, making the exposure parameters calculated based on the overall brightness of the face area also inaccurate. Based on this, by calling the AI model, it is possible to obtain MSAK images containing real face data to ignore or delete the image of non-face areas.
[0014] In one possible implementation, inputting a first facial image into an AI model to obtain a first mask image includes: arranging the RGB domain data of the first facial image and the depth information data of a first depth map to obtain a first facial depth map, where the first depth map is a depth map collected by a depth estimation device at a second moment; and inputting the first facial depth map into the AI model to obtain a first mask image. It is understood that in certain scenarios, when the preview stream includes a distant image, and the distant image is a large facial image (such as a facial image on a light box), the electronic device may determine that the facial image on the light box has the largest facial frame area and obtain exposure parameters based on the facial image on the light box. Although the facial image on the light box is larger, users are generally more concerned with the close-up image than the distant image, so the exposure parameters obtained based on the facial image on the light box may also be inaccurate. By arranging the RGB domain data of the first facial image and the depth information data of the first depth map, the RGB domain data and the depth information data can be fused into the facial depth map.
[0015] In one possible implementation, the AI model stores multiple frames of historical portraits, each of which is a facial depth map obtained based on a RAW image captured before the second moment. Inputting the first facial depth map into the AI model to obtain a first mask image includes: inputting the first facial depth map into the AI model; calling the AI model to obtain a first mask image corresponding to the third image based on the first facial depth map and the multiple frames of historical portraits. It can be understood that by inputting the facial depth map that fuses RGB domain data and depth information data into the AI model, the original algorithm can be enhanced to determine the overall depth of the portrait, making the MASK image ultimately output by the AI model more accurate.
[0016] In one possible implementation, before displaying the first image based on the first data stream on the preview interface, the method further includes: acquiring a first depth map using a depth estimation device. The depth map carries depth information representing the distance from the photographed object to the camera. Typically, the depth map consists of an array of m*n pixels. The larger the depth value of a pixel, the farther the photographed object is from the depth estimation device. Because the depth estimation device only roughly estimates the distance from the photographed object to the camera and does not require complex computational steps such as portrait detection, the resolution of the depth map acquired by the depth estimation device is lower than the resolution of the RAW image acquired by the camera. For example, m = 30 and n = 40. It should be noted that the frames in the data stream acquired by the camera have a one-to-one mapping with the frames in the data stream acquired by the depth estimation device. For example, each frame in the data stream acquired by the camera and the data stream acquired by the depth estimation device is marked with a timestamp. Two frames marked with the same timestamp represent data acquired at the same time in two different formats, representing the same captured content.
[0017] In one possible implementation, determining the second facial region based on the facial frame information includes determining the second facial region based on the facial frame information and the first depth map, where the second facial region is specifically a facial frame region in the second image having a depth less than a preset depth, a facial frame area greater than a preset size, and the largest facial frame area. It will be appreciated that by adding depth assessment for the portrait, facial regions in distant images can be excluded.
[0018] In one possible implementation, determining a first facial region in a third image based on data of a facial region not covered by an occluder includes: determining the first facial region in the third image based on the coordinates of pixels in the facial region not covered by the occluder, wherein the coordinates of the pixels in the first facial region correspond to the coordinates of the pixels in the facial region not covered by the occluder. It will be understood that since the position of the facial region with the largest face frame area in the YUV image is known and the position of the true facial region in the MASK is known, the 3A module can determine the position of the true facial region in the YUV image and obtain the brightness value of each pixel in the true facial region, thereby calculating the brightness of the true face based on the brightness value of each pixel.
[0019] In one possible implementation, the second moment is before the first moment, and the second and third images are different YUV images. For example, when the preview frame currently displayed on the display screen is the i+kth RGB image, a face detection algorithm can be used to obtain face frame information of the i-th YUV image. Then, based on the face frame information, a face image can be cropped from the i-th YUV image. Subsequently, based on the face image of the i-th YUV image, a MSAK image of the i-th YUV image can be obtained. Then, based on the MSAK image of the i-th YUV image, the brightness of the face region of the i+kth YUV image can be obtained. Alternatively, the second moment and the first moment are the same moment, and the second and third images are the same YUV image. For example, the electronic device can obtain the MSAK image of the i-th YUV image based on the face frame information of the i-th YUV image, and then, based on the MSAK image of the i-th YUV image, the brightness of the actual face in the i-th YUV image can be obtained.
[0020] In one possible implementation, the first mask image also includes at least one of data representing the facial region covered by an occluder and data representing the background image. It is understood that when an occluder obstructs the face, overexposure or underexposure may occur. Furthermore, when the subject and the electronic device remain relatively still but the background image is constantly changing, overall exposure instability may also occur. By obtaining the mask image, the facial region covered by the occluder and the background image can be ignored and deleted, resulting in the true facial portion.
[0021] In a second aspect, the present application provides an apparatus comprising units for executing the method described in the first aspect. The apparatus may be configured to execute the exposure parameter adjustment method described in the first aspect. For a description of the units in the apparatus, please refer to the description of the first aspect above and will not be repeated here for the sake of brevity.
[0022] In a third aspect, the present application provides an electronic device comprising a memory and one or more processors. The memory is configured to store computer program code, which includes computer instructions. When the computer instructions are invoked by the processor, the electronic device executes any of the exposure parameter adjustment methods provided in the first aspect.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the exposure parameter adjustment method provided in the first aspect and any possible implementation thereof.
[0024] In a fifth aspect, the present application provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the exposure parameter adjustment method provided in the first aspect and any possible implementation thereof.
[0025] In a sixth aspect, the present application provides a chip. The chip includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuitry. The chip can be used in an electronic device including a communication module and a memory. The interface circuit is configured to receive a signal from the memory of the electronic device and transmit the received signal to the processor. The signal includes computer instructions stored in the memory. When the processor invokes the computer instructions, the electronic device can execute the exposure parameter adjustment method provided in the first aspect and any possible implementation thereof.
[0026] It can be understood that the beneficial effects that can be achieved by the above-mentioned device of the second aspect, the electronic device of the third aspect, the computer-readable storage medium of the fourth aspect, the computer program product of the fifth aspect and the chip of the sixth aspect can be referred to as the beneficial effects in the first aspect and any possible implementation thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A schematic diagram of an operation in a portrait shooting scenario provided by an embodiment of the present application;
[0028] Figure 2 A schematic diagram of capturing image frames based on exposure parameters provided in an embodiment of the present application;
[0029] Figure 3 A schematic diagram of an application scenario of an exposure parameter adjustment method provided in an embodiment of the present application;
[0030] Figure 4 A schematic diagram of an application scenario of another exposure parameter adjustment method provided in an embodiment of the present application;
[0031] Figure 5 A schematic diagram of an application scenario of another exposure parameter adjustment method provided in an embodiment of the present application;
[0032] Figure 6 A schematic diagram of an application scenario of another exposure parameter adjustment method provided in an embodiment of the present application;
[0033] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0034] Figure 8 A schematic diagram of the architecture of an electronic device provided in an embodiment of the present application;
[0035] Figure 9A module interaction diagram for a portrait shooting scenario provided by an embodiment of the present application;
[0036] Figure 10 For Figure 9 A schematic diagram of a specific method flow chart corresponding to the module interaction diagram;
[0037] Figure 11 A schematic diagram of obtaining an MSAK image carrying real face data provided in an embodiment of the present application;
[0038] Figure 12 A schematic diagram of an embodiment of the present application showing a sudden exposure change when image content differs;
[0039] Figure 13 Another module interaction diagram for a portrait shooting scenario provided by an embodiment of the present application;
[0040] Figure 14 A schematic diagram of a depth map collected by a depth estimation device provided in an embodiment of the present application;
[0041] Figure 15 For Figure 13 A schematic diagram of a specific method flow chart corresponding to the module interaction diagram;
[0042] Figure 16 A schematic diagram of obtaining a face depth map provided in an embodiment of the present application;
[0043] Figure 17 A schematic diagram of obtaining a MASK image based on multiple face depth maps provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0045] In the description of this application, unless otherwise specified, " / " means or. For example, A / B can mean A or B. In the description of this application, "and / or" is simply a way to describe the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0046] In the specification and claims of this application, the terms "first" and "second" are used to distinguish different objects or to distinguish different processing of the same object, rather than to describe a specific order of objects. For example, the terms "first operation" and "second operation" are used to distinguish different operations, rather than to describe a specific order of operations. In the embodiments of this application, "plurality" refers to two or more.
[0047] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in some embodiments" or "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0048] With the development of terminal devices, cameras can provide a wider variety of shooting modes, such as photo, portrait, video, night scene, aperture, panorama, slow motion, etc. Users can select a shooting mode from these according to their needs, such as photo mode. Typically, after the user selects a shooting mode, the terminal device can display a preview interface. If the user is satisfied with the shooting effect on the preview interface, the user can click the photo control, and the terminal device will respond to the user's click operation on the photo control and take a photo or video.
[0049] Currently, most electronic devices are equipped with automatic exposure (AE) and automatic white balance (AWB) functions. AE refers to the camera's automatic adjustment of exposure parameters, such as shutter and gain, so that the photosensitive device (such as the image sensor) captures an image at an appropriate exposure and close to the target brightness, thus completing the process of automatic exposure control. Exposure is related to exposure time: the longer the shutter is open, the more light enters, resulting in a brighter image. Gain is related to ISO (International Organization for Standardization), which describes the sensitivity of a photosensitive element to light. A higher ISO indicates a higher sensitivity. In bright conditions, ISO is set to a relatively low value; in low light conditions, the ISO value can be appropriately increased. Exposure is the light intensity multiplied by the time it takes for light to reach the photosensitive device. Exposure is typically expressed in lux (E). For the AWB function, if the white in the original image captured by the image sensor is not processed by AWB, the white image may appear different colors at different color temperatures. For example, it will appear bluish at low color temperatures (such as cloudy days) and yellowish at high color temperatures (such as sunny days). The AWB function uses a white balance algorithm to restore the white imaged under ambient light of different color temperatures to true white (usually the white observed by the human eye under natural daylight).
[0050] In a scene where there is a portrait in the preview video stream captured by the camera, the electronic device can calculate the overall brightness of the face area, and then perform AE processing and / or AWB processing based on the overall brightness to capture an image that meets the brightness requirements.
[0051] Taking AE processing as an example, Figure 1 A schematic diagram of an operation in a portrait shooting scene is shown.
[0052] like Figure 1As shown in (a), the mobile phone displays icons of applications such as the camera on the desktop. When the user wants to take a picture, the user can click on the camera icon 01. In response to the user's click operation on the camera icon 01, the mobile phone starts the camera application. After the camera application completes initialization, the camera application notifies the camera to capture images through the camera framework and camera driver. The camera passes the data stream collected by the sensor to the image signal processor. The image signal processor performs a first format conversion on the data stream collected by the camera to obtain a first data stream (also called a tiny stream, preview stream or preview video stream). For example, the data stream collected by the camera is in RAW format, and the tiny stream is in RGB format. The resolution of the tiny stream is relatively small, such as a few hundred pixels by a few hundred pixels. The image signal processor sends the tiny stream to the screen for display, so that it is presented in the preview box in the form of a preview image with the tiny stream, for example, in the following example. Figure 1 The preview image displayed in the preview interface of (b) is generated based on the tiny stream. Figure 1 Taking the preview interface shown in (b) of FIG as an example, the preview interface may include a preview box 02, an aperture mode selection option, a night scene mode selection option, a portrait mode selection option, a photo mode selection option, a video mode selection option, a professional mode selection option, a photo preview control 04, a photo control 05, and other controls. Preview box 02 is used to display the preview video stream captured by the camera. For example, the preview video stream may consist of multiple frames of the user's selfie portrait. When the user selects the photo mode selection option, a triangle control 03 is displayed in the area below the photo mode selection option to prompt the user that the photo mode has been successfully selected. Photo preview control 04 is used to display the user's most recently taken photo or video in thumbnail format. Photo control 05 is used by the user to confirm triggering the phone to take a photo or video.
[0053] During the display of the preview video stream, the mobile phone can adjust the exposure parameters of the camera in real time based on the brightness of the face area in the preview video stream to adjust the brightness of the preview video stream subsequently captured. Specifically, the mobile phone can obtain the overall brightness of the face area in the captured i-th frame image, and then adjust the exposure parameters of the camera based on the overall brightness, so as to capture the i+j-th frame image based on the adjusted exposure parameters, where i and j are positive integers. Figure 2As shown, the mobile phone can crop face image 1 from image frame 1 captured at time t1, and determine exposure parameter 1 based on the overall brightness of face image 1; crop face image 2 from image frame 2 captured at time t2, and determine exposure parameter 2 based on the overall brightness of face image 2; crop face image 3 from image frame 3 captured at time t3, and determine exposure parameter 3 based on the overall brightness of face image 3... Then, at time t(1+j), image frames 1+j are captured based on exposure parameter 1; at time t(2+j), image frames 2+j are captured based on exposure parameter 2; and at time t(3+j), image frames 3+j are captured based on exposure parameter 3.
[0054] In such Figure 1 For example, if the image displayed in the preview frame 02 shown in (b) is the i-th frame image, if the mobile phone detects that the overall brightness of the face area (such as the rectangular area surrounded by the face frame) in the i-th frame image is dark, the mobile phone can increase the exposure value and / or gain value of the camera, and capture the i+j-th frame image based on the increased exposure value and / or gain value, and then in the following example, Figure 1 The preview box shown in (c) in the figure displays the i+jth frame image. The i+jth frame image may be the next frame image of the i-th frame image, or there may be at least one frame image between the i+jth frame image and the i-th frame image. It should be noted that the present application uses grayscale to present the brightness of the facial image in the accompanying drawings. When the grayscale is darker, it represents that the brightness of the facial image is lower, and when the grayscale is lighter, it represents that the brightness of the facial image is higher. It is understandable that different types of filling patterns or other drawing methods may also be used to present the brightness of the facial image, and this application does not make any specific limitations.
[0055] After the mobile phone adjusts the exposure parameters of the camera, the brightness of the image displayed in the preview box 02 will also change accordingly. As an example, when the user adjusts the exposure parameters of the camera Figure 1 When the selfie portrait with increased brightness shown in (c) is satisfactory, Figure 1 As shown in (d) in the figure, the user can click on the photo control 05. In response to the user's click operation on the photo control 05, the camera application will notify the image signal processor through the camera framework and the camera driver to convert the data stream collected by the camera into a second format to obtain a second data stream (also called a photo stream). For example, the data stream collected by the camera is in RAW format, and the second data stream is in RGB format. The resolution of the photo stream is relatively large, such as several thousand pixels multiplied by several thousand pixels. Since the resolution of each frame image in the photo stream is higher and the imaging effect is better, the camera application will usually generate the final photo or video based on the photo stream. For example, in Figure 1 The photos in the album interface (f) are generated based on the photo stream.
[0056] It should be noted that there is a one-to-one mapping between the frames of the tiny stream and the photo stream. Each frame in the tiny stream and the photo stream is marked with a timestamp. Frame 001 of the tiny stream and frame 101 of the photo stream are both marked with timestamp t0, indicating that frames 001 and 101 are converted from the data stream captured by the camera at time t0; frame 002 of the tiny stream and frame 102 of the photo stream are both marked with timestamp t1, indicating that frames 002 and 102 are both converted from the data stream captured by the camera at time t1; frame 003 of the tiny stream and frame 103 of the photo stream are both marked with timestamp t2, indicating that frames 003 and 103 are both converted from the data stream captured by the camera at time t2; ...; frame 012 of the tiny stream and frame 112 of the photo stream are both marked with timestamp t11, indicating that frames 012 and 112 are both converted from the data stream captured by the camera at time t11. Furthermore, the formats of the tiny stream and the photo stream can be the same or different. For example, the tiny stream and the photo stream can both be in RGB format. The difference between the tiny stream and the photo stream is that the tiny stream has a lower resolution than the photo stream. Taking frames 001 and 101 as an example, since both frames 001 and 101 are converted from the data stream captured by the camera at time t0, theoretically, the image content of frames 001 and 101 should be identical. Frame 001 is obtained by the image signal processor performing a first format conversion on the data stream, resulting in a lower resolution, such as 360*720 pixels. Frame 101 is obtained by the image signal processor performing a second format conversion on the data stream, resulting in a higher resolution, such as 1024*2048 pixels.
[0057] Furthermore, if Figure 1 As shown in (e) in FIG, after the mobile phone completes the photo shooting, the mobile phone updates the thumbnail in the photo preview control 04 to the thumbnail of the photo shot this time. Since the thumbnail in the photo preview control 04 is small in size, the user may not be able to see the specific content of the photo. In this case, Figure 1 As shown in (f) in FIG, the user can click on the photo preview control 04. In response to the user's click operation on the photo preview control 04, the mobile phone opens the gallery application and displays the photo in the following manner: Figure 1 The interface shown in (f) shows the photo 06 taken this time.
[0058] In the above solution, when a human portrait is included in the preview video stream captured by the camera, the electronic device can calculate the overall brightness of the facial region and then perform AE processing based on this overall brightness. This allows the AE processing to capture a brighter preview image when the facial image is darker, and a darker preview image when the facial image is brighter. Therefore, the brightness of the image captured by the electronic device better meets user needs. The specific implementation of the electronic device performing AWB processing based on the overall brightness of the facial region can be found in the description of the related art and will not be elaborated here.
[0059] Typically, the overall brightness of the above-mentioned face area is the average brightness calculated based on the brightness of each pixel in the rectangular area surrounded by the face frame. In some application scenarios, the rectangular area surrounded by the face frame may include not only the face image, but also the background image. The face image may include not only the facial area (i.e., the exposed skin area not covered by the occluder, such as the forehead, glasses, nose, mouth, cheeks, ears and chin, etc.), but also the area covered by the occluder. If the color of the background image or the occluder is significantly different from the color of the facial area, the overall brightness of the face area calculated by the electronic device may not be accurate, making the exposure parameters calculated based on the overall brightness of the face area also inaccurate, resulting in overexposure or underexposure of the camera, ultimately affecting the image quality.
[0060] The following combination Figures 3 to 6 The application scenarios of the exposure parameter adjustment method provided in this application are illustrated with examples.
[0061] As an example scenario, Figure 3 As shown in (a) in the figure, when the subject is not wearing a mask, the electronic device captures the image according to the normal exposure parameters. Figure 3 As shown in (b) in the figure, when the subject wears a light-colored mask (i.e., the color of the mask is lighter than the color of the subject's facial area), the average brightness calculated by the electronic device based on the brightness of each pixel in the facial area is too high, making the exposure value and gain value obtained based on the average brightness too small, resulting in underexposure (i.e., the overall brightness of the image is darker than the actual brightness), resulting in a dark photo or video. Figure 3 As shown in (c) in the figure, when the subject wears a dark mask (i.e., the color of the mask is darker than the color of the subject's facial area), the average brightness calculated by the electronic device based on the brightness of each pixel in the facial area is relatively low, making the exposure value and gain value obtained based on the average brightness relatively large, resulting in overexposure (i.e., the overall brightness of the image is brighter than the actual brightness), causing the final photo or video to be brighter.
[0062] As another example scenario, Figure 4As shown in (a) in FIG, when the subject is wearing transparent lenses (or not wearing glasses), the electronic device captures the image according to normal exposure parameters. Figure 4 As shown in (b) in the figure, when the subject is wearing light-colored lenses (i.e., the color of the lenses is lighter than the color of the subject's facial area), the average brightness calculated by the electronic device based on the brightness of each pixel in the facial area is too high, making the exposure value and gain value obtained based on the average brightness too small, resulting in underexposure (i.e., the overall brightness of the image is darker than the actual brightness), resulting in a dark photo or video. Figure 4 As shown in (c) in the figure, when the subject wears dark lenses (i.e., the color of the lenses is darker than the color of the subject's facial area), the average brightness calculated by the electronic device based on the brightness of each pixel in the facial area is relatively low, making the exposure value and gain value obtained based on the average brightness relatively large, resulting in overexposure (i.e., the overall brightness of the image is brighter than the actual brightness), causing the final photo or video to be brighter.
[0063] As another example scenario, Figure 5 As shown in (a) in the figure, when the subject does not have a beard, the entire facial area is not blocked, and the electronic device captures the image according to normal exposure parameters. Figure 5 As shown in (b) in the figure, when the subject has a beard, part of the facial area is obscured, leaving only the remaining unobstructed facial area visible. Because the color of the area obscured by the beard is darker than the color of the area not obscured by the beard, the electronic device calculates the average brightness based on the brightness of each pixel in the entire face frame and calculates it to be lower. This results in a larger exposure value and gain value based on the average brightness, resulting in overexposure (i.e., the overall image brightness is brighter than the actual brightness), resulting in a brighter photo or video.
[0064] As another example scenario, Figure 6 As shown in (a) in FIG, when the subject is at position P1, the upper left corner area of the face frame 1 is the background image 1. Figure 6 As shown in (b) in FIG, when the subject moves to position P2, the face frame 2 has no background image. Figure 6 As shown in (c) in FIG, when the subject moves to position P3, the upper right corner area of the face frame 3 is the background image 2. Figure 6As shown in (d), when the subject moves to position P4, the left area of the face frame 4 is the background image 3. In the scene where the subject moves from position P1 to position P2, position P3, and position P4 in sequence, the background image in the face frame keeps changing. For example, the relative position of the background image in the face frame, the image content of the background image, and the brightness of the background image may all change. Referring to the description of the above embodiment, the electronic device calculates the average brightness based on the brightness of each pixel in the entire face frame. When the background image in the face frame keeps changing, the exposure parameters calculated by the electronic device based on the average brightness of the face frame also keep changing. In this way, the electronic device needs to constantly adjust the exposure parameters of the camera, making the overall exposure unstable, resulting in flickering (i.e., flickering) in the captured image.
[0065] It should be noted that the above scene examples are merely illustrative and do not limit the present application. It is understood that in other scenes, problems of overexposure, underexposure, or overall unstable exposure may also occur. For example, when the subject's hair obscures the face, or when the subject is wearing a face scarf, problems of overexposure or underexposure may also occur. For another example, when the subject and the electronic device remain relatively still, but the background image is constantly changing, problems of overall unstable exposure may also occur.
[0066] In view of the above problems, an embodiment of the present application provides an exposure parameter adjustment method. This method can be applied to scenarios where an electronic device is used to capture facial images. When displaying a preview image based on a preview stream, if a facial image is detected from the preview image, the electronic device can first segment the real facial part (i.e., facial area) from the facial image, and then calculate the exposure parameters based on the average brightness of the real facial part, and then adjust the exposure parameters of the camera. In this process, by segmenting the real facial part from the facial image, the influence of facial obstructions such as masks, glasses, beards, or hair, and background images on the overall brightness calculation can be reduced, so that the exposure parameters calculated based on the average brightness of the real facial part are more accurate, so that the brightness of the image finally captured is more in line with user needs, and the problem of constant flickering of the captured image is avoided, thereby improving the quality of the film and enhancing the user's photo-taking experience.
[0067] In some embodiments, the electronic device may be a terminal device. The terminal device is also referred to as a terminal or user equipment (UE). For example, the terminal device may be a mobile phone, a personal computer (PC), a smart screen, a smart TV, a tablet computer (Pad), a wearable device, a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, or a wireless terminal in a smart home, etc., or may be other devices or apparatuses.
[0068] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0069] like Figure 7 As shown, the electronic device 100 may include a processor 110, a memory 120, a button 130, a sensor module 140, a display 150, an audio module 160, a speaker 160A, a receiver 160B, a microphone 160C, an earphone interface 160D, a camera 170, a depth estimation device 180, etc.
[0070] The processor 110 can be used to execute the exposure parameter adjustment method in the embodiment of the present application based on the data stream collected by the camera 170. The processor 110 may include one or more processing units. For example, the processor 110 may include a central processing unit (CPU), a neural processing unit (NPU), a graphics processing unit (GPU), an application processor (AP), a digital signal processor (DSP), an image signal processor (ISP), etc. Different processing units can be independent devices; they can also be integrated into one or more processors. For example, the NPU can be set in the DSP.
[0071] The memory 120 can be used to store computer executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the memory 120. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application (application, APP) required for at least one function, such as a camera application, a gallery application, etc. The data storage area can store configuration files of each APP, as well as data created during the use of the electronic device 100, such as a YUV image after the RAW data stream collected by the camera is processed by basic 3A processing and image signal processing (image signal processor, ISP). Among them, 3A processing includes AE processing, AWB processing, and auto focus (auto focus, AF) processing.
[0072] The buttons 130 include a power button, a volume button, and the like. The buttons 130 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100, such as key signal inputs for triggering the camera 170 to capture an image.
[0073] The sensor module 140 may include an image sensor, a touch sensor, etc. Among them, the image sensor, also known as a photosensitive element, is a device that uses the photoelectric conversion function of a photoelectric device to convert the light image on the photosensitive surface into an electrical signal that is proportional to the light image. For example, the image sensor may be a complementary metal oxide semiconductor image sensor (CMOS image sensor, CIS). The touch sensor is also called a "touch panel". The touch sensor can be set on the display screen 150. The touch sensor and the display screen 150 form a touch screen, also known as a "touch screen". The touch sensor is used to detect touch operations acting on or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event and provide visual output related to the touch operation through the display screen 150. In other embodiments, the touch sensor can also be set on the surface of the electronic device 100, at a different location from the display screen 150. It should be noted that the above-mentioned sensor can be an independent functional module in the electronic device, or it can be set in certain functional devices. For example, the image sensor can be integrated into the camera 170.
[0074] The display screen 150 includes a display panel for displaying the desktop, a photo preview interface, various images in the gallery, and the like.
[0075] The electronic device 100 can implement audio functions through the audio module 160, the speaker 160A, the receiver 160B, the microphone 160C, the headphone jack 160D, and the application processor. For example, music playback, recording, etc. Among them, the audio module 160 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 160 can also be used to encode and decode audio signals. The speaker 160A, also known as the "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 160A. The receiver 160B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a call or voice message, the voice can be heard by placing the receiver 160B close to the human ear. The microphone 160C, also known as the "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting his mouth close to the microphone 160C, and the voice signal is input into the microphone 160C. The earphone interface 160D is used to connect a wired earphone.
[0076] The camera 170 is used to capture still images or videos. The light from the object passes through the lens to generate an optical image and is projected onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 170, where N is a positive integer greater than 1.
[0077] The depth estimation device 180 is used to collect depth maps and estimate the distance information of objects in the depth maps through computer vision technology, that is, the distance from the photographed object to the camera. The depth estimation device 180 may include an emitting unit, an optical lens, an imaging unit, a control unit and a computing unit. Through these components, the depth estimation device 180 can collect a depth map, and each pixel in the depth map corresponds to the depth information of a target object, and these pixels together constitute a depth image. In some embodiments, the depth estimation device 180 is an independent functional device in an electronic device. In other embodiments, the depth estimation device 180 may also be a sub-device of a functional device in an electronic device. For example, the depth estimation device 180 may be a TOF sensor in a camera module, and the camera module may also include a camera 170.
[0078] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0079] The following is an example of an electronic device's software system. The electronic device's software system can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, or a microservice architecture. This embodiment of the application takes the Android system with a layered architecture as an example to illustrate the software system architecture of the electronic device.
[0080] For example, Figure 8 A schematic diagram of the architecture of an electronic device provided in an embodiment of the present application is shown.
[0081] like Figure 8 As shown, electronic devices can adopt a layered architecture to divide the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software layers of the software architecture are divided from top to bottom into: application layer, application framework (framework, FW) layer, system library (FWK LIB), hardware abstraction layer (HAL) layer and kernel layer. The above software architecture runs on the hardware layer, which may include a display screen (also called a screen), NPU, depth estimation device and camera, etc.
[0082] The application layer, also known as the application layer, can include a series of application packages. For example, the application layer may include a camera application, a gallery application, etc. The camera application is used to call the camera to capture the tiny stream (micro stream) and the photo stream, and display the captured data on the preview interface; the gallery application is used to store and display the final photos obtained by the camera application. When these application packages are running, they can access the various service modules provided by the application framework layer through the application programming interface (API) and perform corresponding intelligent services.
[0083] The application framework layer provides an API and programming framework for applications in the application layer. The application framework layer includes some predefined functions. The application framework layer can include an activity manager service (AMS), a window manager service (WMS), and a camera service. Among them, the AMS manages the life cycle of each application. The WMS manages all windows in the system. The camera service includes a face detection module, a mask module, and a depth estimation module. For the specific implementation of each module included in the camera service, please refer to the description of the following embodiment.
[0084] The system library can include multiple functional modules, such as a surface manager, media libraries, a two-dimensional (2D) graphics engine (e.g., SGL), and a three-dimensional (3D) graphics processing library (e.g., OpenGLES). The surface manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of various common audio and video formats, as well as static image files. The media library can support multiple audio and video encoding formats. The 2D graphics engine is a drawing engine for 2D graphics. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0085] Within the system libraries, the Android Runtime includes the core libraries and a virtual machine. The Android Runtime is responsible for scheduling and management of the Android system. The core libraries consist of two parts: one containing the functional functions required by the Java language and the other the Android core libraries. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine is responsible for performing functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0086] The hardware abstraction layer has standard interfaces implemented by hardware vendors. For example, the hardware abstraction layer may include camera HAL and audio HAL.
[0087] The kernel layer is the layer between hardware and software, belonging to the bottom layer of the Android system. The kernel layer can contain various driver interfaces, such as camera drivers, display drivers, audio drivers, etc. In addition, the kernel layer can also include 3A modules. 3A modules specifically include AE module, AWB module, and AF module.
[0088] It should be noted that although the embodiments of the present application are described using the Android system as an example, its basic principles are also applicable to electronic devices based on operating systems such as iOS or Windows.
[0089] Below Figure 7 and Figure 8 Based on the functional modules provided, combined Figures 9 to 17 , an example is given to illustrate the specific implementation of the exposure parameter adjustment method provided in the embodiment of the present application.
[0090] For example, Figure 9 A module interaction diagram for a portrait shooting scenario is shown.
[0091] like Figure 9 As shown, the camera may include at least a lens and a photosensitive element (such as a CIS), and the CPU may include at least a face detection module, a MASK module, an AE module and a camera driver.
[0092] In response to the user opening the camera application, the camera application sends a capture request / image acquisition request to the camera service. This capture request / image acquisition request may include parameters such as the camera identity (ID), frame rate range, and capture mode corresponding to the current capture scene. The camera service sends the camera ID and frame rate range corresponding to the current capture scene to the camera driver through the camera HAL. The camera driver can then open the corresponding camera based on the camera ID corresponding to the current capture scene. The opened camera begins to capture data streams.
[0093] The light emitted or reflected by the photographed object (including at least one person) generates an optical image through the lens and is projected onto the CIS. The CIS converts the light signal into an electrical signal in RAW format (i.e., a RAW image), and then passes the RAW image to the 3A module. The 3A module performs basic 3A processing on the RAW image and passes it to the ISP. The ISP converts it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format (such as RGB and YUV formats), such as an RGB image and a YUV image. It should be noted that the RGB image and the YUV image can be converted to each other. In the RGB color encoding method, each pixel has three colors: red, green, and blue. In the YUV color encoding, Y represents brightness, U and V represent chroma, and brightness and chroma specify the color of the pixel. For display screens, images are usually displayed through the RGB model, but the YUV model is usually used when processing image data because the YUV model saves more storage space. Based on this, the DSP executes the following processing flow: On the one hand, the DSP sends the preview stream in RGB format (i.e., RGB image) for display, so that the user can see what is shown. Figure 1On the other hand, the DSP puts the preview stream in YUV format (i.e., the YUV image) into the memory so that the CPU can execute the exposure parameter adjustment method provided in the embodiment of the present application. It can be understood that since the RGB image and the YUV image are images of different formats generated based on the same digital image signal, the image content of the two is consistent, such as containing the same face image. In this way, in the process of displaying the preview stream in RGB format, the CPU can obtain the preview stream in YUV format from the memory, and based on the brightness of the face area in the preview stream in YUV format, adjust the exposure parameters of the camera in real time to change the brightness of the preview video stream collected subsequently.
[0094] Specifically, after processing the i-1th YUV image frame, the face detection module can read the i-th YUV image frame from the memory, detect the i-th YUV image frame to obtain the face frame information in the current image (such as the number of faces and the coordinates of the face frame), and then pass the i-th YUV image frame and the face frame information of the i-th YUV image frame to the MSAK module. Based on the face frame information of the i-th YUV image frame and the i-th YUV image frame, the MSAK module crops the mask data of the real face corresponding to the maximum face area from the i-th YUV image frame, and then passes the mask data to the AE module. Based on the mask data and the i-th YUV image frame, the AE module calculates the brightness of the real face, then determines the exposure parameters corresponding to the brightness of the real face, and sends the exposure parameters to the camera through the camera driver. The camera can readjust the exposure parameters and collect a new data stream based on the adjusted exposure parameters.
[0095] Figure 10 For Figure 9 Schematic diagram of the specific method flow corresponding to the module interaction diagram.
[0096] This method can be applied to Figure 7 and Figure 8 The electronic device shown in the figure may include at least a CPU, an ISP, a DSP, and a camera. The CPU may include at least functional modules such as a camera application, a face detection module, a MASK module, a 3A module (including an AE module), and a camera driver.
[0097] like Figure 10 As shown, the method may include the following S101 to S124.
[0098] S101: A camera application receives an operation from a user to open the camera application or a request from another application to call the camera application.
[0099] For example, the camera application can receive a user click operation on the camera application icon, or receive a user selection operation on a shooting mode in the preview interface of the camera application, or receive a user trigger operation to open the camera in other applications (such as a video call operation, a photo operation, or a scan operation), etc.
[0100] S102: The camera application sends a shooting request to the camera driver to instruct the camera to capture a data stream.
[0101] S103: In response to the shooting request, the camera driver sends a shooting command to the camera head to instruct the camera head to capture a data stream.
[0102] Exemplarily, in response to the user's operation of opening the camera application or the request of other applications to call the camera application, the camera application of the application layer identifies the current shooting scene, determines the camera corresponding to the current shooting scene (for example, opening the rear camera in the long-range shooting scene, and opening the front camera in the selfie scene), the camera's frame rate range (including the maximum frame rate and the minimum frame rate), the shooting mode, etc. The camera application sends the camera ID, frame rate range, shooting mode and other parameters corresponding to the current shooting scene to the camera driver through the camera service and the camera HAL in sequence. The camera driver can open the corresponding camera according to the camera ID corresponding to the current shooting scene. The opened camera collects data streams based on the frame rate range and shooting mode. With reference to the description of the above embodiment, the data stream collected by the camera is an electrical signal in RAW format, that is, a RAW image.
[0103] S104: The camera returns the collected data stream, ie, the RAW image, to the camera driver.
[0104] S105 , the camera driver transmits the RAW image to the 3A module.
[0105] S106 , the 3A module performs basic 3A processing on the RAW image and transmits the processed RAW image to the ISP.
[0106] The above basic 3A processing may include the simplest pre-processing operations such as white balance correction and image cropping, but does not include AE processing.
[0107] S107 , the ISP converts the processed RAW image into a digital image signal, and transmits the digital image signal to the DSP.
[0108] S108, the DSP converts the digital image signal into an image signal in a standard format (such as RGB and YUV format), such as an RGB image and a YUV image.
[0109] Referring to the description of the above embodiment, RGB image and YUV image are two different color encoding methods, and the two can be converted into each other. For display screens, images are usually displayed through the RGB model. When processing image data, the YUV model that saves more storage space is usually used. Based on this, after the DSP converts the digital image signal into RGB image and YUV image, the DSP can perform the following processing flow: On the one hand, the DSP sends the preview stream (i.e., RGB image) in RGB format to the display, so that the user can see the image. Figure 1 On the other hand, the DSP stores the preview stream (i.e., YUV image) in YUV format into the memory. Then, for each frame of the preview stream stored in the memory, the face detection module can execute the following S109 according to the acquisition order of each frame. In other words, when the electronic device displays the preview interface based on the acquired preview stream, the electronic device can also adjust the camera exposure parameters based on the brightness of the face in the acquired preview stream to adjust the brightness of the preview stream to be acquired.
[0110] The following description is made by taking the YUV image read from the memory by the face detection module as the i-th frame YUV image as an example.
[0111] S109: The face detection module reads the i-th YUV image frame (referred to as the second image) from the memory.
[0112] S110: The face detection module determines whether the i-th YUV image frame includes a face.
[0113] The face detection module can use a face detection algorithm, such as a template matching-based method, a singular value feature-based method, a subspace analysis method, a local preservation projection method or a principal component analysis method, to determine whether the i-th frame YUV image includes a face. If the i-th frame YUV image includes a face, then the exposure parameter adjustment method provided in the embodiment of the present application can be used, that is, the following S111-S124 are executed. If the i-th frame YUV image does not include a face, then there is no need to execute the exposure parameter adjustment method provided in the embodiment of the present application, and the i+1-th frame YUV image can continue to be read from the memory, and it can be determined whether the i+1-th frame YUV image includes a face. As a possible implementation method, when the i-th frame YUV image does not include a face, the electronic device can also use other exposure parameter adjustment methods provided by related technologies to adjust the exposure parameters of the camera.
[0114] S111, the face detection module obtains face frame information from the i-th YUV image frame.
[0115] The above-mentioned face frame information is information about the face image in the YUV image. For example, the face frame information may include the number of faces in the YUV image and the coordinates of each face frame. The number of faces in the YUV image may be one or more. If the YUV image includes one face, a face frame is determined to enclose the face area. If the YUV image includes multiple faces, multiple face frames are determined to enclose the face areas. The coordinates of a face frame can be used to represent the position and size of a face area in the YUV image. For example, if the face frame is rectangular in shape, if a two-dimensional coordinate system is established with a vertex corner of the YUV image as the origin, the coordinates of a face frame include the coordinates of the four vertices of the rectangular frame.
[0116] S112: The face detection module transmits the i-th YUV image and the face frame information of the i-th YUV image to the MSAK module. In this way, the MSAK module can obtain the MASK data of the real face in the i-th YUV image based on the face frame information of the i-th YUV image.
[0117] As a first possible implementation, the MSAK module can obtain the mask data of the real face in the i-th YUV image frame based on the face frame information of the i-th YUV image frame, i.e., S113-S118 described below. As a second possible implementation, the MSAK module can obtain the mask data of the real face in the i+k-th YUV image frame based on the face frame information of the i-th YUV image frame. The specific implementation of the second possible implementation is similar to that of the first possible implementation, and can be referred to the description of S113-S118 described below, which is not repeated here. Wherein, i and k are positive integers.
[0118] S113, the MSAK module obtains the face region with the largest face frame area in the i-th YUV image based on the face frame information of the i-th YUV image.
[0119] When the i-th YUV image frame includes a single face frame, the "face region with the largest face frame area in the i-th YUV image frame" refers to the face region enclosed by the single face frame. When the i-th YUV image frame includes multiple face frames, the "face region with the largest face frame area in the i-th YUV image frame" refers to the face region enclosed by the face frame with the largest face frame area among the multiple face frames.
[0120] like Figure 11As shown, a two-dimensional coordinate system xoy is established with the upper left corner of the i-th YUV image frame as the origin. Based on the face detection algorithm, the face detection module can determine two rectangular face frames in the i-th YUV image frame: face frame 1 and face frame 2. The coordinates of the four vertices of face frame 1 are (x1, y1), (x2, y2), (x3, y3), and (x4, y4), and the coordinates of the four vertices of face frame 2 are (x5, y5), (x6, y6), (x7, y7), and (x8, y8). Based on these vertex coordinates, the MSAK module can calculate the area of the face region enclosed by face frame 1 as S1 = (y7 - y5) * (x6 - x5), and the area of the face region enclosed by face frame 2 as S2 = (y3 - y1) * (x2 - x1). If S1 > S2, then the face region enclosed by face frame 1 is the face region with the largest face frame area in the i-th YUV image frame.
[0121] S114, the MSAK module determines whether the size of the largest facial region in the face frame is greater than or equal to a preset size.
[0122] As an example, the MSAK module determines whether the area of the face region is greater than or equal to a preset area.
[0123] As another example, the MSAK module determines whether the height of the facial region is greater than or equal to a preset height, and whether the width of the facial region is greater than or equal to a preset width. The preset height and preset width can be equal or different. For example, the preset height and preset width can both be 64 pixels.
[0124] It will be appreciated that the size of each YUV image frame is typically fixed. If the face region with the largest face frame area in the i-th YUV image frame is smaller than a preset size, this indicates that the face occupies a relatively small proportion of the i-th YUV image frame, has little impact on the overall image brightness, and is not the focal point. Therefore, the brightness of this face region should not be used as a basis for adjusting exposure parameters, and S115 below need not be executed. If the face region with the largest face frame area in the i-th YUV image frame is greater than or equal to the preset size, this indicates that the face occupies a relatively large proportion of the i-th YUV image frame and is likely the focal point, and S115 below can be continued.
[0125] S115, the MSAK module crops an image of the face region with the largest face frame area from the i-th frame YUV image (referred to as the first face image).
[0126] S116, the MSAK module converts the image of the face region with the largest face frame area into RGB format.
[0127] Since artificial intelligence (AI) models usually obtain MSAK images based on images in RGB format, the image of the face area needs to be converted from the YUV domain to the RGB domain before inputting it into the AI model.
[0128] S117, the MSAK module adjusts the image of the face area converted into the RGB domain to a preset size.
[0129] Taking the preset size of 64 pixels * 64 pixels as an example, referring to the description of S114 in the above embodiment, since the size of the face area with the largest area in the face frame may far exceed 64 pixels * 64 pixels, directly inputting it into the AI model may result in a large amount of AI model computation. To reduce the amount of AI model computation, the MSAK module can first resize the image of the face area to 64 pixels * 64 pixels before inputting it into the AI model.
[0130] It should be noted that the embodiment of the present application does not specifically limit the execution order of S116 and S117. For example, the MSAK module may execute S117 first and then execute S116. After S116 and S117, the MSAK module executes the following S118.
[0131] S118, the MSAK module calls the AI model to obtain the MSAK image.
[0132] The above-mentioned MSAK image includes data corresponding to the exposed skin area not covered by the occluder, and data corresponding to the area covered by the occluder and / or the background image. The MASK image is a matrix with the same dimension as the original image (i.e., the i-th frame YUV image), and the elements in the MASK image determine whether to retain, modify or block the image pixels at the corresponding position in the i-th frame YUV image. In an embodiment of the present application, the MASK image can be used to select the real face area and filter the non-face area (such as the background image and the area covered by the occluder). As an example, in the MSAK image of the real face, the pixel value of a pixel point is 0, which indicates that the pixel point is in the real face area, and the pixel value of a pixel point is 1, which indicates that the pixel point is in the non-face area. As another example, the pixel value of a pixel point is 1, which indicates that the pixel point is in the real face area, and the pixel value of a pixel point is 0, which indicates that the pixel point is in the non-face area.
[0133] In some embodiments, to reduce the overall performance and power consumption of electronic devices, some algorithms can be deployed on the CPU side, while model reasoning and some pre-processing algorithms can be placed on the DSP. For example, the DSP is equipped with an NPU, and the NPU's AI model can perform model reasoning. After the MSAK module obtains an image of the facial area, the image of the facial area can be input into the AI model to obtain an MSAK image containing real facial data from the AI model.
[0134] For example, continue as Figure 11 As shown, the area of the area surrounded by face frame 1 is larger than the area of the area surrounded by face frame 2, so the MSAK module can crop the image of the area surrounded by face frame 1 from the i-th frame YUV image. The image of the area surrounded by face frame 1 includes the real face area (i.e., the exposed skin area not covered by the occluder, such as the forehead, glasses, nose, mouth, cheeks, ears and chin, etc.), the background image, and the area covered by the occluder (such as the face area covered by the mask). Referring to the description of the above embodiment, if the color of the background image or the area covered by the occluder is significantly different from the color of the real face area, the calculated overall brightness of the face area may not be accurate, so that the exposure parameters calculated based on the overall brightness of the face area are also not accurate enough. Based on this, the MSAK module can call the AI model to obtain the MSAK image carrying real face data to ignore or delete the image of the non-face area.
[0135] In some embodiments, the above-mentioned AI model can be a portrait segmentation model. Specifically, after the image of the face area is adjusted to RGB of a preset size through S116 and S117, the MSAK module can pass the image into the portrait segmentation model, and the portrait segmentation model continues to segment the image. After the segmentation is completed, data of size 2*64*64 is obtained, where 2 is the dimension and 64 represents the width and height. Then, the portrait segmentation model reduces it to data of size 1*64*64 through the argMax() operation, turning it into a 0-1MSAK image, and then transforms (mapping) the 1*64*64 MSAK image into data of size 1*16*16. Among them, argMax() is the maximum independent variable point set function, which means finding the parameter with the maximum score.
[0136] S119, the MSAK module transmits the MSAK image and the i-th frame YUV image to the 3A module.
[0137] Specifically, the MSAK module transfers the MSAK image and the i-th frame YUV image to the AE module of the 3A module, so that the AE module performs AE processing, that is, executes the following S120 and S121.
[0138] S120, 3A module calculates the brightness of the real face based on the MASK image and the i-th frame YUV image.
[0139] The "brightness of the real face" mentioned above refers to the brightness of the area corresponding to the mask image in the i-th YUV image frame. It is understandable that since the position of the facial region with the largest face frame area in the i-th YUV image frame is known, and the position of the real face region within the mask is also known, the 3A module can determine the position of the real face region in the i-th YUV image frame and obtain the brightness value of each pixel in the real face region. Based on the brightness value of each pixel, the brightness of the real face can be calculated. For example, the average brightness value of all pixels in the real face region can be used as the brightness value of the real face.
[0140] S121, the 3A module determines an exposure parameter corresponding to the brightness of a real face.
[0141] The above exposure parameters include exposure value and gain value. The exposure value is related to the exposure time. The longer the shutter is open, the more light enters, and the image appears brighter. The gain value is related to ISO. The sensitivity is used to describe the sensitivity of the photosensitive element to light. The higher the sensitivity, the higher the sensitivity. In bright light conditions, set the ISO value to a relatively low value. In low light conditions, the ISO value can be appropriately increased.
[0142] In the embodiment of the present application, the brightness of the face is negatively correlated with the exposure and gain values. When the brightness of the face is high, the calculated exposure and gain values are smaller, reducing the brightness of the subsequent captured images. When the brightness of the face is low, the calculated exposure and gain values are larger, increasing the brightness of the subsequent captured images. When the brightness of the face is just right, there is no need to adjust the camera's exposure parameters.
[0143] S122, the 3A module sends exposure parameters to the camera driver.
[0144] S123: The camera driver instructs the camera to adjust exposure parameters.
[0145] The camera can capture a new data stream (ie, a new RAW image) based on the adjusted exposure parameters.
[0146] At step S124, the camera returns the captured new RAW image to the camera driver. The camera driver can then pass the new RAW image to the 3A module for a new round of processing. For details, refer to the description of steps S106 to S123 in the above embodiment and will not be repeated here.
[0147] In the above solution, when displaying a preview image based on the preview stream, if a face image is detected in the preview image, the actual face portion can be segmented from the face image. Exposure parameters are then calculated based on the average brightness of the actual face portion, and the camera's exposure parameters are then adjusted. By segmenting the actual face portion from the face image, the impact of facial obstructions such as masks, glasses, beards, or hair, as well as the background image, on the overall brightness calculation is reduced. This makes the exposure parameters calculated based on the average brightness of the actual face portion more accurate, resulting in a brightness that better meets user needs and avoids the problem of constant flickering in the captured image, improving the quality of the resulting image and enhancing the user's photography experience.
[0148] The above embodiment uses the example of obtaining an MSAK image of the i-th YUV frame based on the face frame information of the i-th YUV frame, and then obtaining the brightness of the real face of the i-th YUV frame based on the MSAK image of the i-th YUV frame. In actual implementation, depending on the scheduling logic from the face detection module to the mask module, the MSAK image obtained by the mask module may not be the MSAK image of the preview frame currently displayed on the display screen. For example, the preview frame currently displayed on the display screen is the i+k-th image, while the MSAK image obtained by the mask module is the MSAK image of the i-th YUV frame. When the image content of the i-th YUV frame and the i+k-th image is consistent, the above exposure parameter adjustment method can meet the requirements for accurate exposure parameter adjustment. However, if the image content of the i-th YUV frame and the i+k-th image differs (for example, due to the movement of the subject), obtaining the brightness of the real face of the i+k-th YUV frame based on the MSAK image of the i-th YUV frame may result in sudden exposure changes.
[0149] For example, the preview frame currently being displayed on the display screen is the i+kth frame RGB image (called the first image). Figure 12 As shown, the i-th frame YUV image (called the second image) is converted from the RAW image collected at time ti, the i+k-th frame YUV image (called the third image) and the i+k-th frame RGB image ( Figure 12 (not shown) is converted from the RAW image collected at time t(i+k). The face detection module can obtain the face frame information of the i-th frame YUV image, such as the coordinates of face frame 1, through the face detection algorithm. Then, the MASK module crops the face image (called the first face image) from the i-th frame YUV image based on the coordinates of face frame 1, and then obtains the MSAK image of the i-th frame YUV image based on the face image of the i-th frame YUV image. Then, the 3A module obtains the face area of the i+k-th frame image (such as the MSAK image of the i-th frame YUV image) based on the MSAK image of the i-th frame YUV image. Figure 12It is understandable that if the subject is in motion, the specific position of the subject in the YUV image of the i-th frame and the YUV image of the i+k-th frame will be different. This will cause the facial image cropped based on face frame 1 to not cover the entire face, resulting in an inaccurate MASK image (called the first mask image). This in turn makes the exposure parameters obtained based on the MASK image inaccurate, resulting in a sudden exposure change.
[0150] Furthermore, in certain scenarios, when the preview stream includes a distant image, and the distant image is a relatively large facial image (such as a facial image displayed on a light box), if the exposure parameter adjustment method provided in the above embodiment is used, the electronic device may determine that the facial image displayed on the light box is the facial image with the largest face frame area and obtain exposure parameters based on the facial image displayed on the light box. Although the facial image displayed on the light box is relatively large, users are generally more concerned with close-up images than distant images. Therefore, the exposure parameters obtained based on the facial image displayed on the light box may also be inaccurate.
[0151] Based on the above reasons, based on the exposure parameter adjustment method provided in the above embodiment, the embodiment of the present application provides an improved exposure parameter adjustment method. When displaying a preview image based on a preview stream, the electronic device can obtain the depth data of the face image, and then search for the face area with the largest face frame area whose depth is less than a preset depth and whose face frame area is larger than a preset size, so as to avoid misjudging the larger face image in the distant image as the face image used to extract the MASK image, so that the final screening of the face area is more accurate. In addition, the AI model adds a timing processing function on the basis of the traditional portrait segmentation algorithm, and the AI model internally stores multiple frames of historical portrait data, so that the AI model can predict the MASK image of the i+kth frame YUV image based on the multiple frames of historical portrait data and the i-th frame YUV image, thereby compensating for the frame difference in the path transmission, reducing the erroneous exposure operation, and improving the quality of the film.
[0152] The following combination Figures 13 to 17 An example is given to illustrate the improved exposure parameter adjustment method.
[0153] exist Figure 9 Based on the module interaction diagram provided, Figure 13 Another module interaction diagram in a portrait shooting scenario is shown.
[0154] and Figure 9 The difference is that in Figure 13 The module interaction diagram shown here adds a depth estimation component, which is used to estimate the distance information of objects in the depth information image, that is, the distance from the photographed object to the camera, using computer vision technology.
[0155] In response to the user opening the camera application, the camera application sends a capture request to the camera driver. In response to the capture request, the camera driver can issue capture commands to the camera and the depth estimation device respectively.
[0156] On the one hand, the camera starts to collect data streams, and the light emitted or reflected by the photographed object generates an optical image through the lens and is projected onto the CIS. The CIS converts the light signal into an electrical signal in RAW format, namely a RAW image, and then passes the RAW image to the 3A module. The 3A module performs basic 3A processing on the RAW image and passes it to the ISP. The ISP converts it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format (such as RGB and YUV format), such as an RGB image and a YUV image. Please refer to the relevant description of the above embodiment and will not repeat it here.
[0157] On the other hand, the depth estimation device can collect a data stream consisting of a depth information map (also called a depth image, depth map), each frame of which carries depth information representing the distance from the photographed object to the camera. For example, Figure 14 FIG. 1 shows a schematic diagram of a depth map collected by a depth estimation device. Figure 14 As shown, the depth map consists of an array of m*n pixels. The larger the depth value of a pixel, the farther the object is from the depth estimation device. It should be noted that since the depth estimation device is only used to roughly estimate the distance between the object and the camera, and does not require complex calculation steps such as portrait detection, the resolution of the depth map collected by the depth estimation device is lower than the resolution of the RAW image collected by the camera. For example, m = 30, n = 40.
[0158] It should be noted that the frames in the data stream captured by the camera have a one-to-one mapping relationship with the frames in the data stream captured by the depth estimation device. For example, each frame in the data stream captured by the camera and the data stream captured by the depth estimation device is marked with a timestamp. Two frames marked with the same timestamp represent two different forms of data captured at the same time, representing the same captured content.
[0159] For example, the preview frame currently displayed on the display is the i+kth RGB image. The face detection module can read the i-th YUV image from the memory and perform detection on the i-th YUV image to obtain face frame information (such as the number of faces and the coordinates of the face frames) in the i-th YUV image. The MSAK module can obtain the face frame information of the i+kth YUV image and the i-th YUV image from the memory / face detection module, and obtain the i-th depth map from the depth estimation device. The i-th depth map is acquired by the depth estimation device at time ti, the i-th YUV image is converted from the RAW image acquired at time ti, and the i+kth YUV image and the i+kth RGB image are converted from the RAW image acquired at time t(i+k). Based on the face frame information of the i-th YUV image and the i-th depth map, the MSAK module selects the face region with the largest face frame area, where the depth of the face image is less than a preset depth and the face frame area is larger than a preset size. The MSAK module then extracts mask data based on the filtered facial regions and multiple frames of historical portrait data stored within the AI model. This mask data is then passed to the AE module. The AE module calculates the brightness of the true face based on this mask data and the YUV image of the (i+k)th frame. It then determines the exposure parameters corresponding to the true face brightness and sends them to the camera through the camera driver. The camera can then readjust the exposure parameters and capture a new data stream based on the adjusted exposure parameters.
[0160] Figure 15 For Figure 13 The specific method flow chart corresponding to the module interaction diagram. Figure 10 The difference is that in Figure 15 In the method flow shown, the electronic device may further include a depth estimation device, and the AI model stores multiple frames of historical portrait data. Figure 15 As shown, the method may include the following S201 to S226.
[0161] S201: The camera application receives an operation of opening the camera application by a user or a request from another application to call the camera application.
[0162] S202: The camera application sends a shooting request to the camera driver to instruct the camera to capture a data stream.
[0163] For the specific implementation of S201 - S202 , reference may be made to the description of S101 - S102 in the above embodiment.
[0164] S203: In response to the shooting request, the camera driver sends a shooting command to the camera and the depth estimation device to instruct them to collect a data stream.
[0165] In response to the user opening the camera application or a request from another application to call the camera application, the camera application at the application layer identifies the current shooting scene and determines the shooting parameters corresponding to the current shooting scene. The camera application sends the shooting parameters corresponding to the current shooting scene to the camera driver through the camera service and the camera HAL in sequence. The camera driver can issue shooting commands to the camera and the depth estimation device respectively to instruct them to collect data streams. In response to the shooting instruction, the camera begins to collect a data stream consisting of RAW images. At the same time, the depth estimation device can collect a data stream consisting of depth maps. Each frame of the depth map carries depth information used to represent the distance from the object being photographed to the camera / camera.
[0166] S204: The camera returns the collected data stream, ie, the RAW image, to the camera driver.
[0167] S205 , the camera driver transmits the RAW image to the 3A module.
[0168] S206 , the 3A module performs basic 3A processing on the RAW image and transmits the processed RAW image to the ISP.
[0169] S207 , the ISP converts the processed RAW image into a digital image signal, and transmits the digital image signal to the DSP.
[0170] S208 , the DSP converts the digital image signal into an image signal in a standard format (such as RGB and YUV format), such as an RGB image and a YUV image.
[0171] For the specific implementation of S204-S208, reference may be made to the description of S104-S108 in the above embodiment.
[0172] After the DSP converts the digital image signal into an RGB image and a YUV image, the DSP can execute the following processing flow: on the one hand, the DSP sends the preview stream in RGB format (i.e., the RGB image) to the display; on the other hand, the DSP puts the preview stream in YUV format (i.e., the YUV image) into the memory, and then for each frame image in the preview stream stored in the memory, the face detection module can execute the following S209 according to the acquisition order of each frame image.
[0173] Taking the preview image currently displayed on the display screen as the i+kth frame RGB image (referred to as the first image) based on the data stream as an example, the face detection module may execute the following S209-S211.
[0174] S209: The face detection module reads the i-th YUV image frame from the memory.
[0175] S210: The face detection module determines whether the i-th YUV image frame includes a face.
[0176] The face detection module can use a face detection algorithm, such as a template matching-based method, a singular value feature-based method, a subspace analysis method, a locality preserving projection method, or a principal component analysis method, to determine whether the i-th YUV image frame includes a face. If the i-th YUV image frame includes a face, the exposure parameter adjustment method provided in the embodiment of the present application can be used, that is, the following S211-S226 are executed. If the i-th YUV image frame does not include a face, there is no need to execute the exposure parameter adjustment method provided in the embodiment of the present application, and the i+1-th YUV image frame can be read from the memory to determine whether the i+1-th YUV image frame includes a face.
[0177] S211, the face detection module obtains face frame information from the i-th YUV image frame.
[0178] The above-mentioned face frame information is the information of the face image in the i-th frame YUV image. For example, the face frame information may include the number of faces in the YUV image and the coordinates of each face frame.
[0179] For the specific implementation of S209 - S211 , reference may be made to the description of S109 - S111 in the above embodiment.
[0180] At step S212, the MSAK module obtains the YUV image of frame i+k from the memory and obtains the face frame information of the YUV image of frame i from the face detection module. The YUV image of frame i is converted from the RAW image captured at time ti (referred to as the second time point), and the YUV image of frame i+k (referred to as the third image) and the RGB image of frame i+k are converted from the RAW image captured at time t(i+k) (referred to as the first time point).
[0181] S213: The MSAK module obtains the i-th frame depth map from the depth estimation device, wherein the i-th frame depth map is collected by the depth estimation device at time ti (referred to as the second time).
[0182] S214, the MSAK module obtains the face region with the largest face frame area in the i-th YUV frame image, whose depth is less than a preset depth and whose face frame area is greater than a preset size, based on the face frame information of the i-th YUV frame image and the i-th depth map.
[0183] For example, when the i-th YUV image frame includes at least one face frame, the MSAK module may first filter out face regions with a depth less than a preset depth from the at least one face frame, then filter out face regions with a face frame area greater than a preset size from the face regions less than the preset depth, and finally select the face region with the largest face frame area. In other embodiments, the MSAK module may also first filter out face regions with a depth less than a preset depth from the at least one face frame, then select the face region with the largest face frame area from the face regions less than the preset depth, and finally determine whether the size of the face region with the largest face frame area is greater than or equal to a preset size. The specific implementation of filtering out the face region with the largest face frame area greater than the preset size from the face frames can be referred to the description of S113 and S114 in the above embodiment and will not be repeated here.
[0184] Since the i-th depth map and the i-th YUV image are captured at the same time, the image content of the two frames is consistent. For example, the same face appears in the same relative position in both frames, meaning that there is a mapping relationship between the pixels of the i-th YUV image and the pixels of the i-th depth map. Based on this mapping relationship between the pixels of the i-th YUV image and the pixels of the i-th depth map, the depth values of each face frame in the i-th YUV image can be obtained. The depth value of a face frame can be the average of the depth values of all pixels in the face frame. It can be understood that the greater the depth value of a face frame, the farther the object corresponding to the face frame is from the depth estimation device. If the depth value of a face frame is greater than or equal to a preset depth, it indicates that the distance between the object corresponding to the face frame and the depth estimation device exceeds the preset distance and is considered distant content. In this case, the face area of the face frame can be ignored or deleted.
[0185] S215, the MSAK module crops the image of the face area from the i-th frame YUV image.
[0186] S216, the MSAK module converts the image of the face area into RGB format.
[0187] Since AI models usually obtain MSAK images based on images in RGB format, the image of the face area needs to be converted from the YUV domain to the RGB domain before inputting it into the AI model.
[0188] S217, the MSAK module adjusts the image of the face area converted to the RGB domain to a preset size to obtain RGB domain data of the face area.
[0189] Taking the preset size of 64 pixels * 64 pixels as an example, since the size of the face area image converted to the RGB domain may far exceed 64 pixels * 64 pixels, directly inputting it into the AI model may result in a large amount of AI model computation. To reduce the AI model's computational complexity, the MSAK module can first resize the face area image to 64 pixels * 64 pixels, thereby obtaining the RGB domain data of the face area.
[0190] S218, the MSAK module adjusts the face area of the i-th frame depth image to a preset size to obtain depth information data of the face area.
[0191] The face region in the above-mentioned “face region in the i-th frame depth map” and the face region cropped from the i-th frame YUV image are regions corresponding to the same face. The above-mentioned “depth information data of the face region” includes the depth value of each pixel in the face region.
[0192] It should be noted that the preset size of S217 and S218 is the same, for example, 64 pixels * 64 pixels. By adjusting the face area of the i-th frame depth image to the preset size, the face area in the "i-th frame depth image face area" can be made the same size as the face area cropped from the i-th frame YUV image, thereby facilitating the rearrangement of the depth information data of the face area and the RGB domain data of the face area.
[0193] S219, the MSAK module rearranges the depth information data of the face area and the RGB domain data of the face area to obtain a face depth map.
[0194] For example, Figure 16 As shown, the face detection algorithm detects that the i-th YUV image frame contains two face images: one located in face frame 4, closer to the camera, and the other located in face frame 3, farther from the camera and on a light box. Because the depth A of face frame 4 is greater than a preset depth, the depth B of face frame 3 is less than a preset depth, and the size of face frame 4 is larger than a preset size, the MSAK module can crop the face image of face frame 4 from the i-th YUV image frame. This face image is then converted to the RGB domain and resized to 64 pixels by 64 pixels, thereby obtaining depth information data for the face region. Furthermore, the MSAK module resizes the same region in the i-th depth image to 64 pixels by 64 pixels, obtaining depth information data for the face region. The MSAK module then rearranges the depth information data and the RGB domain data for the face region to obtain a face depth map.
[0195] The RGB domain data for the face area is 3*64*64 in size, where 3 represents the dimension and 64 represents the width and height. After rearranging the depth information data and the RGB domain data for the face area, the 3D input data (3*64*64) becomes 4D input data (4*64*64). In other words, the face depth map is 4D input data (4*64*64).
[0196] S220, the MSAK module calls the AI model to obtain a MASK image based on the facial depth map obtained through S219 and the multiple frames of historical portraits saved in the AI model.
[0197] In some embodiments, the above-mentioned AI model may be a portrait segmentation model.
[0198] To reduce the overall performance and power consumption of electronic devices, some algorithms can be deployed on the CPU side, while model reasoning and some pre-processing algorithms can be placed in the DSP. For example, the DSP is equipped with an NPU, and the NPU's AI model can perform model reasoning. After the MSAK module obtains the facial depth map through S219, the facial depth map can be input into the AI model to obtain an MSAK image containing real facial data from the AI model.
[0199] With reference to the description of the above embodiment, the MSAK image obtained based on the MASK module is not the MSAK image of the preview frame currently being displayed on the display screen. For example, the preview frame currently being displayed on the display screen is the i+k frame image, and the MSAK image obtained by the MASK module is the MSAK image of the i-th frame YUV image. When there is a difference in the image content of the i-th frame YUV image and the i+k-th frame image (such as the movement of the photographed person), if the MSAK image of the i-th frame YUV image is obtained based on the face frame information of the i-th frame YUV image, then there may be a problem of exposure mutation. In order to solve this problem, the AI model adds a timing processing function on the basis of the traditional portrait segmentation algorithm, and the AI model internally stores multiple frames of historical portrait data. In this way, the MSAK module can call the AI model to predict the MASK image corresponding to the i+k-th frame YUV image based on the face depth map obtained through S219 and the multiple frames of historical portraits stored in the AI model.
[0200] For example, take the example of saving two frames of historical portraits in the AI model. Figure 17As shown, the AI model stores the facial depth map of the YUV image frame i-2 and the facial depth map of the YUV image frame i-1. The MSAK module can pass the facial depth map of the YUV image frame i-1 into the AI model. Based on the facial depth maps of the YUV image frame i-2, the YUV image frame i-1, and the YUV image frame i, the AI model can estimate the movement speed of the subject, and then estimate the MASK image corresponding to the YUV image frame i+k based on the movement speed of the subject. The portrait segmentation model then uses the argMax() operation to reduce the image to 1*64*64 data, converting it into a 0-1 MSAK image. The 1*64*64 MSAK image is then transformed (mapped) into 1*16*16 data. argMax() is the maximum independent variable point set function, which indicates the search for the parameter with the maximum score.
[0201] S221, the MSAK module transmits the MSAK image obtained in S220 and the (i+k)th frame YUV image to the 3A module.
[0202] Specifically, the MSAK module transfers the MSAK image and the (i+k)th frame YUV image to the AE module of the 3A module, so that the AE module performs AE processing, that is, executes the following S222 and S223.
[0203] S222, 3A module calculates the brightness of the real face in the i+kth frame YUV image based on the MASK image.
[0204] The above “brightness of the real face” refers to the brightness of the area corresponding to the MASK image in the i+k-th frame YUV image.
[0205] S223, module 3A determines exposure parameters corresponding to the brightness of the real face.
[0206] The above exposure parameters include exposure value and gain value.
[0207] S224, the 3A module sends exposure parameters to the camera driver.
[0208] S225: The camera driver instructs the camera to adjust exposure parameters.
[0209] The camera can capture a new data stream, i.e., a new RAW image, based on the adjusted exposure parameters.
[0210] In step S226 , the camera returns the acquired new RAW image to the camera driver, which then passes the new RAW image to the 3A module to execute a new round of processing.
[0211] For the specific implementation of S221-S226, reference may be made to the description of S119-S124 in the above embodiment.
[0212] In the above scheme, when displaying a preview image based on a preview stream, the electronic device can obtain the depth data of the facial image, and then search for the facial region with the largest facial frame area whose depth is less than a preset depth and whose facial frame area is larger than a preset size, thereby avoiding misjudging the larger facial image in the distant image as the facial image used to extract the MASK image, making the final screening of the facial region more accurate. In addition, the AI model adds a timing processing function based on the traditional portrait segmentation algorithm. The AI model internally stores multiple frames of historical portrait data, so that the AI model can predict the MASK image of the i+kth frame YUV image based on the multiple frames of historical portrait data and the i-th frame YUV image, thereby compensating for the frame difference in the channel transmission, reducing the misexposure operation, and improving the quality of the film.
[0213] An embodiment of the present application further provides an electronic device, including a processor, wherein the processor is coupled to a memory, and the processor is configured to execute a computer program or instruction stored in the memory, so that the electronic device implements the methods in the above embodiments.
[0214] The embodiment of the present application also provides a computer-readable storage medium having computer instructions stored therein. When the computer instructions are run on a computer, the computer is caused to execute the method shown above. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more media that can be integrated. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk or tape), an optical medium or a semiconductor medium (e.g., a solid state drive (SSD)), etc.
[0215] An embodiment of the present application further provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the methods in the above embodiments.
[0216] The present application also provides a chip that is coupled to a memory and is used to read and execute computer programs or instructions stored in the memory to perform the methods in the above embodiments. The chip can be a general-purpose processor or a dedicated processor. It should be noted that the chip can be implemented using the following circuits or devices: one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this application.
[0217] The electronic device, computer-readable storage medium, computer program product and chip provided in the above-mentioned embodiments of the present application are all used to execute the methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the methods provided above, and will not be repeated here.
[0218] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0219] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0220] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0221] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0223] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for adjusting exposure parameters, characterized in that: The method comprises: Displaying a first image based on the first data stream on a preview interface; cropping a first facial image from a second image based on the first data stream; Acquire a first mask image corresponding to the first facial image, where the first mask image at least includes data of a facial area not covered by an obstruction; Determining a first facial region in the third image based on the data of the facial region not covered by the occlusion; Obtaining a first exposure parameter based on the brightness of the first facial area; Controlling the camera to collect a second data stream based on the first exposure parameter; Among them, the first image is an RGB image converted from the RAW image collected at the first moment, the second image is a YUV image converted from the RAW image collected at the second moment, and the third image is a YUV image converted from the RAW image collected at the first moment.
2. The method according to claim 1, characterized in that The step of cropping a first facial image from a second image based on the first data stream includes: Obtaining face frame information using a face detection algorithm, the face frame information including the number of faces in the second image and the coordinates of each face frame, where one face frame corresponds to one face; determining a second face region based on the face frame information, where the second face region is a face frame region in the second image having a face frame area larger than a preset size and a largest face frame area; When the size of the second facial area is greater than or equal to a preset size, the first facial image is cropped from the second image, and the first facial image is an image of the second facial area.
3. The method according to claim 2, characterized in that After cropping the first facial image from the second image, the method further includes: Converting the first facial image into RGB format; The first facial image converted into the RGB format is adjusted to the preset size.
4. The method according to claim 3, characterized in that The acquiring of a first mask image corresponding to the first facial image includes: The first facial image is input into an artificial intelligence model to obtain the first mask image, where the first facial image is an image converted to a preset size and RGB format.
5. The method according to claim 4, characterized in that Inputting the first facial image into the artificial intelligence model to obtain the first mask image includes: Arrange the RGB domain data of the first facial image and the depth information data of the first depth map to obtain a first facial depth map, where the first depth map is a depth map collected by the depth estimation device at the second moment; The first face depth map is input into the artificial intelligence model to obtain the first mask image.
6. The method according to claim 5, characterized in that The artificial intelligence model stores multiple frames of historical portraits, where the multiple frames of historical portraits are facial depth maps obtained based on RAW images collected before the second moment; Inputting the first face depth map into the artificial intelligence model to obtain the first mask image includes: Inputting the first face depth map into the artificial intelligence model; The artificial intelligence model is called to obtain the first mask image corresponding to the third image based on the first facial depth map and the multiple frames of historical portraits.
7. The method according to claim 5, characterized in that Before displaying the first image based on the first data stream on the preview interface, the method further includes: The first depth map is collected by the depth estimation device.
8. The method according to any one of claims 2 to 7, characterized in that The determining of the second face region based on the face frame information includes: Based on the face frame information and the first depth map, the second face area is determined. The second face area is specifically a face frame area in the second image having a depth less than a preset depth, a face frame area greater than a preset size, and a largest face frame area.
9. The method according to any one of claims 1 to 8, characterized in that The determining, based on the data of the facial region not covered by the occluder, the first facial region in the third image includes: Based on the coordinates of the pixel points in the facial area not covered by the occlusion, a first facial area is determined in the third image, where the coordinates of the pixel points in the first facial area correspond to the coordinates of the pixel points in the facial area not covered by the occlusion.
10. The method according to any one of claims 1 to 9, characterized in that The second moment is a moment before the first moment, and the second image and the third image are different YUV images; or, the second moment and the first moment are the same moment, and the second image and the third image are the same YUV image.
11. The method according to any one of claims 1 to 10, characterized in that The first mask image further includes at least one of data of a face region covered by an occluder and data of a background image.
12. An electronic device, characterized in that: The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions, which, when executed on an electronic device, enable the electronic device to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Photographing effect adjusting method of mobile terminal and mobile terminal
CN105933607A
Direct method unsupervised monocular image scene depth estimation method
CN112085776A
Mask face recognition algorithm based on low-rank attention mechanism
CN114581984A
Image processing method and device, electronic equipment and storage medium
CN115205172A
Image processing method, model training method, electronic equipment and readable storage medium
CN115661912A