A photographing method, an electronic device, and a storage medium

By detecting the direction of the gaze using the front-facing camera, the rear camera is controlled to autofocus, solving the problem of camera shake during manual focusing and improving shooting results and user experience.

CN117278839BActive Publication Date: 2026-01-27HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210666645.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2026-01-27
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

During manual focusing, the shaking or vibration of electronic devices can affect the shooting results, and existing technology cannot meet the photographer's focusing needs in real time.

Method used

The front-facing camera detects the photographer's gaze direction and controls the rear cameras to focus, displaying sharp and blurred areas, providing gaze guidance and distance prompts, and achieving automatic focusing.

Benefits of technology

It improves shooting results, reduces the impact of device shake, enhances user experience and the immediacy of focusing, and meets the needs of photographers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117278839B_ABST
    Figure CN117278839B_ABST
Patent Text Reader

Abstract

The application provides a photographing method, an electronic device and a storage medium, relates to the technical field of image applications, and can improve the photographing effect of the electronic device. The method can determine the line of sight of a photographer and a region corresponding to a first preview image, i.e., a first target ROI, based on an image collected by a front camera after the electronic device detects a video recording instruction. Then, the electronic device can control a rear first camera and a rear second camera to focus according to the first target ROI and display a preview image after focusing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technology, and in particular to a shooting method, electronic device and storage medium. Background Technology

[0002] Currently, more and more people are using electronic devices (such as mobile phones, tablets, cameras, etc.) to take photos and videos to record the moments of life. To meet the needs of photographers, electronic devices generally have a manual focus function. By manually touching the screen of the electronic device, the photographer selects the object to focus on, so that the object falls exactly on the focal point of the camera, thus obtaining the clearest image.

[0003] However, during manual focusing, when the photographer touches the screen of the electronic device, the device will shake (or vibrate), affecting the shooting results. Summary of the Invention

[0004] This application provides a shooting method, an electronic device, and a storage medium, which can improve the shooting effect of the electronic device.

[0005] The embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, a shooting method is provided, applied in an electronic device, the electronic device including a rear first camera, a rear second camera, a front camera, and a display screen; the method includes: after detecting a recording command, the electronic device displays a first preview interface on the display screen; the first preview interface includes a first preview image; the electronic device determines a first target region of interest (ROI) on the first preview image based on an image captured by the front camera; the first target ROI is the area corresponding to the viewer's line of sight; the electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI, and displays a second preview interface; the second preview interface includes a second preview image; the second preview image includes a first sharp area and a first blurred area, the first sharp area corresponding to the first target ROI; the second preview image is generated by the electronic device processing a first image frame and a second image frame; the first image frame is captured by the rear first camera, and the second image frame is captured by the rear second camera.

[0007] Based on the first aspect, after the electronic device detects the recording command, it can determine the area on the display screen corresponding to the position of the photographer's gaze on the first preview image based on the image captured by the front-facing camera, thus determining the first target ROI. Then, the electronic device controls the rear first and second cameras to focus according to the first target ROI and displays the second preview interface. Since the electronic device determines the first target ROI based on the image captured by the front-facing camera, i.e., the electronic device focuses based on the photographer's gaze, it can solve the problem of shaking (or jarring) during manual focusing, thereby improving the shooting effect.

[0008] In one implementation of the first aspect, the method further includes: when the electronic device detects a change in the photographer's gaze, the electronic device determines a second target ROI on a second preview image based on an image captured by a front-facing camera; the second target ROI is the area corresponding to the photographer's gaze; the electronic device controls a rear-facing first camera and a rear-facing second camera to focus according to the second target ROI, and displays a third preview interface; the third preview interface includes a third preview image; the third preview image includes a second clear area and a second blurred area, the second clear area corresponding to the second target ROI; the third preview image is generated by the electronic device processing a third image frame and a fourth image frame; the third image frame is captured by the rear-facing first camera, and the fourth image frame is captured by the rear-facing second camera; wherein the first clear area is different from the second clear area, and the first blurred area is different from the second blurred area.

[0009] In this implementation, when the electronic device detects a change in the photographer's gaze, it can determine a second target ROI based on the image captured by the front-facing camera, and control the first and second rear cameras to focus according to the second target ROI; that is, the electronic device can refocus based on the changed direction of the photographer's gaze to further improve the shooting effect.

[0010] In one implementation of the first aspect, after displaying the first preview interface on the display screen, the method further includes: the electronic device receiving a first event input by the user and displaying a fourth preview interface on the display screen; the first event is used to trigger the electronic device to enter a large aperture mode; the fourth preview image includes a third sharp area and a third blurred area; the fourth preview image is generated by the electronic device processing a fifth image frame and a sixth image frame; the fifth image frame is captured by a rear-mounted first camera, and the sixth image frame is captured by a rear-mounted second camera; wherein the third sharp area is different from the first sharp area, and the third blurred area is different from the first blurred area.

[0011] In this implementation, after the electronic device detects the recording command, it can receive the first event input by the user, triggering the electronic device to enter the large aperture mode. As a result, the fourth preview image displayed by the electronic device includes both sharp and blurred areas, improving the user experience.

[0012] In one implementation of the first aspect, the electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI and displays the second preview interface, including: the electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI, determines the first focus area in the second preview interface, and displays the second preview interface; wherein the first focus area corresponds to the first clear area.

[0013] In one implementation of the first aspect, the second preview interface includes a first prompt message, which is used to indicate the position of the first target ROI on the display screen; or, the first prompt message is used to guide the photographer to focus their gaze on the target position on the display screen.

[0014] In this implementation, the electronic device can prompt the photographer with the location of the first target ROI on the display screen through the first prompt information; or, the electronic device can guide the photographer to focus their gaze on the target location on the display screen through the first prompt information, thereby improving the shooting effect and enhancing the user experience.

[0015] In one implementation of the first aspect, the second preview interface includes a mask area, which is divided into multiple preset ROIs, each corresponding to a second preview image; a first prompt message is located within a target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI; wherein, the mask area is used to display an image captured by the front-facing camera; or, the mask area is used to display a scaled-down second preview image.

[0016] In this implementation, the electronic device can set a mask area in the second preview interface and divide the mask area into multiple preset ROIs. Since the multiple preset ROIs correspond to the second preview image, the first prompt information is located within the target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI. Therefore, the electronic device can intuitively prompt the photographer about the position of the first target ROI on the display screen through the first prompt information within the mask area, and guide the photographer to focus their gaze on the target position on the display screen through the relationship between the first prompt information and the multiple preset ROIs. This improves the shooting effect and further enhances the user experience.

[0017] In one implementation of the first aspect, the second preview interface includes multiple preset ROIs, which correspond to the second preview image; the first prompt information is located within the target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI.

[0018] In this implementation, the electronic device can further divide the second preview interface into multiple preset ROIs. Since the multiple preset ROIs correspond to the second preview image, and the first prompt information is located within the target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI, the electronic device can intuitively prompt the photographer on the position of the first target ROI on the display screen through the first prompt information in the second preview interface; and guide the photographer to focus their gaze on the target position on the display screen through the relationship between the first prompt information and the multiple preset ROIs, thereby improving the shooting effect and further enhancing the user experience.

[0019] In one implementation of the first aspect, when the mask area is used to display the image captured by the front-facing camera, the method further includes: when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device displays a preset face area in the mask area; the preset face area is used to prompt the photographer to adjust the distance between themselves and the electronic device.

[0020] In this implementation, the electronic device displays a preset face area in the mask area to prompt the photographer to adjust the distance between themselves and the electronic device, thereby improving the shooting effect and further enhancing the user experience.

[0021] In one implementation of the first aspect, the second preview interface further includes a second prompt message; the second prompt message is used to indicate to the photographer the position of the first focus area on the display screen.

[0022] In this implementation, since the electronic device can indicate the position of the first focus area on the display screen to the photographer through the second prompt information, the shooting effect can be improved, thereby further enhancing the user experience.

[0023] In one implementation of the first aspect, the method further includes: when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device issues a text prompt; the text prompt is used to prompt the photographer to adjust the distance between themselves and the electronic device; or, when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device issues a voice prompt; the voice prompt is used to prompt the photographer to adjust the distance between themselves and the electronic device.

[0024] In this implementation, when the electronic device detects that the distance between the photographer and the electronic device is not within the preset range, the electronic device can also prompt the photographer to adjust the distance between the photographer and the electronic device through text prompts or voice prompts, thereby improving the shooting effect and further enhancing the user experience.

[0025] In one implementation of the first aspect, the method further includes: the electronic device identifying a focused object within the first focus area and tracking the focused object; when the electronic device detects that the duration for which the tracked focused object does not correspond to the first target ROI is longer than a preset duration, the electronic device redetermines the first focus area.

[0026] In this implementation, the electronic device can identify and track the focus object within the first focus area. When the electronic device detects that the tracked focus object does not correspond to the first target ROI for a duration longer than a preset duration, the electronic device re-determines the first focus area. In other words, the electronic device uses a delay strategy for focusing, which avoids affecting the shooting effect when the electronic device focuses frequently.

[0027] In one implementation of the first aspect, the method further includes: the electronic device identifying a focused object within the first focus area and tracking the focused object; when the electronic device detects that the tracked focused object does not correspond to the first target ROI and the photographer's gaze is not on the display screen, the electronic device maintains the first focus area.

[0028] In this implementation, the electronic device identifies and tracks the focus object within the first focus area. When the electronic device detects that the tracked focus object does not correspond to the first target ROI and the photographer's gaze is not on the display screen (i.e., the photographer's gaze is outside the display screen), the electronic device maintains the first focus area unchanged to avoid the problem of the electronic device being unable to focus due to the photographer's gaze being outside the display screen.

[0029] In one implementation of the first aspect, the second preview interface includes an end-recording control; the method further includes: the electronic device generating a video file in response to the photographer's operation on the end-recording control; wherein the video file includes a first clear area and a first blurred area; the video file is generated by the electronic device processing a first image frame and a second image frame.

[0030] In this implementation, since the electronic device uses the first target ROI for focusing during video recording, it can improve the shooting effect. Therefore, after the electronic device finishes recording video, it can also improve the shooting effect of the generated video file.

[0031] In one implementation of the first aspect, the electronic device determines a first target region of interest (ROI) on a first preview image based on an image captured by a front-facing camera. This includes: the electronic device performing face recognition processing on the image captured by the front-facing camera to determine face information and eye information in the image; the face information includes the coordinates of the subject's facial contour, and the eye information includes one or more of the following: interpupillary distance, pupil size, pupil size variation, pupil brightness contrast, corneal radius, spot information, and iris information; the electronic device inputs the face information and eye information into a preset model and outputs the position where the subject's gaze falls on the display screen; the preset model is trained by the electronic device based on sample face information and sample eye information; and the electronic device determines the first target ROI on the first preview image based on the area corresponding to the subject's gaze in the first preview image.

[0032] In this implementation, the electronic device first performs face recognition based on the image captured by the front-facing camera to determine the face and eye information of the image captured by the front-facing camera. In this way, the electronic device does not need to detect the entire image, but only the image related to the face, thereby narrowing the detection range, improving detection accuracy, and reducing device power consumption.

[0033] In one implementation of the first aspect, the electronic device controls the rear first camera and the rear second camera to focus based on the first target ROI to determine the first focus area, including: the electronic device performs autofocus (AF) processing on the first target ROI, controls the rear first camera and the rear second camera to focus, and determines the first focus area.

[0034] In one implementation of the first aspect, the method further includes: an electronic device preprocessing a second image frame; the preprocessing is used to make the field of view of the second image frame and the first image frame the same; the electronic device calculating the depth of field based on the preprocessed second image frame and the first image frame; and the electronic device determining a first blurred region based on the target ROI and the depth of field.

[0035] In this implementation, the electronic device preprocesses the second image frame to make the field of view of the second image frame the same as that of the first image frame. This ensures that the electronic device can calculate the depth of field more accurately using the second image frame and the first image frame, thereby further improving the shooting effect.

[0036] In one implementation of the first aspect, before the electronic device displays the second preview interface, the method further includes: the electronic device performing image conversion processing on the first image frame and the second image frame; the image conversion processing includes: the electronic device converting the first image frame into a first image frame of a target format, and converting the second image frame into a second image frame of a target format; the bandwidth of the first image frame during transmission is higher than the bandwidth of the first image frame of the target format during transmission, and the bandwidth of the second image frame during transmission is higher than the bandwidth of the second image frame of the target format during transmission.

[0037] In this implementation, since the electronic device performs image conversion processing on the first image frame and the second image frame, it can reduce the bandwidth of the first image frame during transmission and the bandwidth of the second image frame during transmission, thus further reducing the power consumption of the device.

[0038] In one implementation of the first aspect, the method further includes: an electronic device performing image simulation transformation processing on a first image frame of the target format; the image simulation transformation processing is used to enhance the first image frame of the target format.

[0039] In this implementation, the electronic device can enhance the image of the first image frame in the target format by performing image simulation transformation processing on the first image frame in the target format, thereby further improving the image effect.

[0040] In one implementation of the first aspect, the method further includes: the electronic device, in response to a zoom operation input by the photographer, performing zoom processing on a first image frame of the target format to generate a first image frame of the target format corresponding to the target zoom factor.

[0041] In this implementation, during the shooting process, the electronic device can also perform zoom processing on the first image frame of the target format to generate a first image frame of the target format corresponding to the target zoom factor, thereby further improving the shooting effect.

[0042] In a second aspect, an electronic device is provided, which has the functions described in the first aspect. These functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions.

[0043] Thirdly, an electronic device is provided, comprising: a rear-facing first camera, a rear-facing second camera, a front-facing camera, a display screen, a memory, and one or more processors; the display screen is used to display images captured by the rear-facing first camera, the rear-facing second camera, and the front-facing camera; or the display screen is used to display images generated by the processor; the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the following steps: after detecting a recording instruction, the electronic device displays a first preview interface on the display screen; the first preview interface includes a first preview image; the electronic device determines a first target region of interest (ROI) on the first preview image based on the image captured by the front-facing camera; the first target ROI is the area corresponding to the viewer's line of sight; the electronic device controls the rear-facing first camera and the rear-facing second camera to focus according to the first target ROI, and displays a second preview interface; the second preview interface includes a second preview image; the second preview image includes a first sharp area and a first blurred area, the first sharp area corresponding to the first target ROI; the second preview image is generated by the electronic device processing a first image frame and a second image frame; the first image frame is captured by the rear-facing first camera, and the second image frame is captured by the rear-facing second camera.

[0044] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device further performs the following steps: when the electronic device detects a change in the photographer's gaze, the electronic device determines a second target ROI on the second preview image based on the image captured by the front-facing camera; the second target ROI is the area corresponding to the photographer's gaze; the electronic device controls the rear first camera and the rear second camera to focus according to the second target ROI, and displays a third preview interface; the third preview interface includes a third preview image; the third preview image includes a second clear area and a second blurred area, the second clear area corresponding to the second target ROI; the third preview image is generated by the electronic device processing a third image frame and a fourth image frame; the third image frame is captured by the rear first camera, and the fourth image frame is captured by the rear second camera; wherein, the first clear area is different from the second clear area, and the first blurred area is different from the second blurred area.

[0045] In one implementation of the third aspect, after the first preview interface is displayed on the screen, when the computer instruction is executed by the processor, the electronic device further performs the following steps: the electronic device receives a first event input by the user and displays a fourth preview interface on the screen; the first event is used to trigger the electronic device to enter the large aperture mode; the fourth preview image includes a third sharp area and a third blurred area; the fourth preview image is generated by the electronic device processing a fifth image frame and a sixth image frame; the fifth image frame is captured by the rear first camera, and the sixth image frame is captured by the rear second camera; wherein, the third sharp area is different from the first sharp area, and the third blurred area is different from the first blurred area.

[0046] In one implementation of the third aspect, when the computer instruction is executed by the processor, the electronic device specifically performs the following steps: the electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI, determines the first focus area in the second preview interface, and displays the second preview interface; wherein, the first focus area corresponds to the first clear area.

[0047] In one implementation of the third aspect, the second preview interface includes a first prompt message, which is used to indicate the position of the first target ROI on the display screen; or, the first prompt message is used to guide the photographer to focus their gaze on the target position on the display screen.

[0048] In one implementation of the third aspect, the second preview interface includes a mask area, which is divided into multiple preset ROIs, each corresponding to a second preview image; the first prompt information is located within a target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI; wherein, the mask area is used to display the image captured by the front camera; or, the mask area is used to display a scaled-down second preview image.

[0049] In one implementation of the third aspect, the second preview interface includes multiple preset ROIs, which correspond to the second preview image; the first prompt information is located within the target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI.

[0050] In one implementation of the third aspect, when the mask area is used to display the image captured by the front-facing camera, when the computer instruction is executed by the processor, the electronic device also performs the following steps: when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device displays a preset face area in the mask area; the preset face area is used to prompt the photographer to adjust the distance between themselves and the electronic device.

[0051] In one implementation of the third aspect, the second preview interface also includes a second prompt message; the second prompt message is used to indicate to the photographer the position of the first focus area on the display screen.

[0052] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device further performs the following steps: when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device issues a text prompt; the text prompt is used to prompt the photographer to adjust the distance between themselves and the electronic device; or, when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device issues a voice prompt; the voice prompt is used to prompt the photographer to adjust the distance between themselves and the electronic device.

[0053] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device further performs the following steps: the electronic device identifies the focus object in the first focus area and tracks the focus object; when the electronic device detects that the duration for which the tracked focus object does not correspond to the first target ROI is longer than a preset duration, the electronic device redetermines the first focus area.

[0054] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device also performs the following steps: the electronic device identifies the focus object in the first focus area and tracks the focus object; when the electronic device detects that the tracked focus object does not correspond to the first target ROI and the photographer's gaze is not on the display screen, the electronic device maintains the first focus area.

[0055] In one implementation of the third aspect, the second preview interface includes an end-recording control; when a computer instruction is executed by the processor, the electronic device further performs the following steps: the electronic device generates a video file in response to the photographer's operation of the end-recording control; wherein the video file includes a first clear area and a first blurred area; the video file is generated by the electronic device processing a first image frame and a second image frame.

[0056] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device performs the following steps: The electronic device performs face recognition processing on the image captured by the front-facing camera to determine the face information and eye information of the image captured by the front-facing camera; the face information includes the coordinates of the facial contour of the photographer, and the eye information includes one or more of the following: interpupillary distance, pupil size, pupil size variation, pupil brightness contrast, corneal radius, spot information, and iris information; the electronic device inputs the face information and eye information into a preset model and outputs the position where the photographer's gaze falls on the display screen; the preset model is trained by the electronic device based on sample face information and sample eye information; the electronic device determines the first target ROI on the first preview image based on the area where the photographer's gaze falls on the first preview image.

[0057] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device specifically performs the following steps: the electronic device performs autofocus (AF) processing on the first target ROI, controls the rear first camera and the rear second camera to focus, and determines the first focus area.

[0058] In one implementation of the second aspect, when the computer instructions are executed by the processor, the electronic device further performs the following steps: the electronic device preprocesses the second image frame; the preprocessing is used to make the field of view of the second image frame and the first image frame the same; the electronic device calculates the depth of field based on the preprocessed second image frame and the first image frame; the electronic device determines the first blurred region based on the target ROI and the depth of field.

[0059] In one implementation of the third aspect, before the electronic device displays the second preview interface, when the computer instruction is executed by the processor, the electronic device further performs the following steps: the electronic device performs image conversion processing on the first image frame and the second image frame; the image conversion processing includes: the electronic device converts the first image frame into a first image frame of a target format, and converts the second image frame into a second image frame of a target format; the bandwidth of the first image frame during transmission is higher than the bandwidth of the first image frame of the target format during transmission, and the bandwidth of the second image frame during transmission is higher than the bandwidth of the second image frame of the target format during transmission.

[0060] In one implementation of the third aspect, when the computer instructions are executed by the processor, the electronic device further performs the following steps: the electronic device performs image simulation transformation processing on the first image frame of the target format; the image simulation transformation processing is used to enhance the first image frame of the target format.

[0061] In one implementation of the third aspect, when a computer instruction is executed by a processor, the electronic device responds to a zoom operation input by the photographer, performs zoom processing on a first image frame of the target format, and generates a first image frame of the target format corresponding to the target zoom factor.

[0062] Fourthly, a computer-readable storage medium is provided, which stores computer instructions that, when executed on a computer, enable the computer to perform the photographing method described in any one of the first aspects.

[0063] Fifthly, a computer program product containing instructions is provided, which, when executed on a computer, enables the computer to perform the shooting method described in any one of the first aspects.

[0064] The technical effects of any of the design methods in the second to fourth aspects can be found in the technical effects of different design methods in the first aspect, and will not be repeated here. Attached Figure Description

[0065] Figure 1 A schematic diagram of an imaging principle provided in an embodiment of this application;

[0066] Figure 2 This is a schematic diagram illustrating a photographer taking a picture using a handheld electronic device, as provided in an embodiment of this application.

[0067] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0068] Figure 4 This is a schematic diagram of the structure of a mobile phone camera provided in an embodiment of this application;

[0069] Figure 5 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 1 ;

[0070] Figure 6 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 2 ;

[0071] Figure 7 A schematic diagram of the principle of image processing provided in this application embodiment. Figure 1 ;

[0072] Figure 8 A schematic diagram of the principle of image processing provided in this application embodiment. Figure 2 ;

[0073] Figure 9a A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 3 ;

[0074] Figure 9b A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 4 ;

[0075] Figure 10 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 5 ;

[0076] Figure 11 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 6 ;

[0077] Figure 12 A schematic diagram of the principle of image processing provided in this application embodiment. Figure 3 ;

[0078] Figure 13 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 7 ;

[0079] Figure 14 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 8 ;

[0080] Figure 15 A schematic diagram nine of a shooting interface provided for an embodiment of this application;

[0081] Figure 16 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 ;

[0082] Figure 17 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 one;

[0083] Figure 18 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 two;

[0084] Figure 19 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 three;

[0085] Figure 20 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 Four;

[0086] Figure 21 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 five;

[0087] Figure 22 A flowchart illustrating a shooting method provided in this application embodiment. Figure 1 ;

[0088] Figure 23 A schematic diagram of a shooting interface provided in an embodiment of this application. Figure 10 six;

[0089] Figure 24 A flowchart illustrating a shooting method provided in this application embodiment. Figure 2 ;

[0090] Figure 25 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0091] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" are not necessarily different. Meanwhile, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.

[0092] To facilitate understanding of the solutions provided in the embodiments of this application, some terms involved in the embodiments of this application will be explained below.

[0093] Bokeh: In an image, there are sharp parts and blurred parts (or indistinct parts); the imaging of the blurred parts is called bokeh. Specifically, the blurred parts in an image can include foreground blur or background blur.

[0094] Focal point: Cameras in electronic devices are generally composed of at least one lens, including convex and concave lenses. Taking a convex lens as an example, when a light beam reflected (or emitted) from the subject is projected onto the convex lens, the beam gradually converges to a single point, which is the focal point. After the beam converges to a point, it will diverge again as it continues to propagate.

[0095] Depth of field: When an electronic device takes a picture, the process of making a subject that is a certain distance away from the camera appear sharp is called focusing. The point where the subject is located is called the focus point. Within a certain range before and after the focus point, the electronic device can still obtain a sharp image; that is, within a certain range before and after the focus point, the image of the subject remains sharp. Therefore, the range within which the image is sharp before and after the focus point is called the range of sharpness. This range of sharpness is called the depth of field. It should be noted that within the range of sharpness, the area closer to the camera than the focus point is called the foreground depth of field, and the area farther away from the camera than the focus point is called the background depth of field.

[0096] Understandably, during the imaging process of an electronic device, the light beam reflected from the object being photographed propagates to the imaging surface, thus forming an image of the object on the imaging surface. Typically, the light beam reflected from the object is focused at a single point (i.e., the focal point) after passing through the camera; this focal point may be located in front of, behind, or even on the imaging surface. Taking the focal point being located on the imaging surface as an example, for instance... Figure 1 The diagram illustrates an electronic device capturing an image of a subject. O represents the optical axis of camera L, F is the focal point, and f is the focus point. Within the range M1 to M2 before and after focus point f, the electronic device can still acquire a clear image; therefore, the range M1 to M2, i.e., the distance S, represents the depth of field. (The diagram continues...) Figure 1 As shown, within the range from M1 to M2, there exist near point A and far point B. The light beams reflected from both near point A and far point B can pass through camera L to the imaging plane. Figure 1 In the diagram, A' is the imaging point of the near point A, and B' is the imaging point of the far point B.

[0097] Aperture: An optical device used to control the amount of light passing through a camera, typically located inside the camera. Generally, an aperture consists of several leaf-shaped metal blades forming an adjustable opening in the center. By rotating these blades, the size of the opening can be adjusted, thus changing the aperture size. It's important to note that a larger opening allows more light to pass through the camera, resulting in a larger aperture (also called a large aperture); conversely, a smaller opening allows less light to pass through the camera, resulting in a smaller aperture (also called a small aperture).

[0098] In related technologies, electronic devices use autofocus to select the subject to be photographed, ensuring that the subject falls on the focal point of the camera. However, sometimes the photographer does not want the electronic device to automatically select the subject, meaning autofocus cannot always meet the photographer's needs. Therefore, electronic devices can allow the photographer to manually switch the focus point by operating the screen, thus providing a manual focus function.

[0099] In some embodiments, when the photographer holds the electronic device with both hands (e.g., in landscape mode) to take a picture, the photographer's focusing operation will affect the photographer's holding motion of the electronic device when manual focusing is required. For example, as... Figure 2 As shown, when the photographer needs to manually focus, they switch from holding the device with both hands to one hand and touch the screen to change the focus object. During this process, the electronic device will experience some shaking (or vibration), thus affecting the shooting effect. Furthermore, since the photographer's focusing operation requires touching the screen, this touch operation introduces a certain delay. Therefore, when shooting a dynamic scene, the photographer's focusing operation will affect the immediacy of the focus, thereby impacting the quality of the dynamic scene shot.

[0100] In some embodiments, when an electronic device enters a large aperture mode for shooting, the photographer's focus operation on the object being focused on can also affect the bokeh effect in the image, resulting in a significant change in the final image.

[0101] Based on this, related technologies can also switch the focus object by detecting the behavior of the subject (such as turning its head, turning around, or moving). However, detecting the behavior of the subject may not meet the photographer's requirements and may even contradict them. Therefore, this method still cannot meet the photographer's needs in real time.

[0102] This application provides a shooting method that solves the problem of camera shake (or shakiness) during manual focusing, thereby improving the shooting effect. For example, when the electronic device is shooting, it activates the front-facing camera to detect the direction of the photographer's gaze, and controls the electronic device to switch the focus object according to the direction of the target's gaze, thus achieving focus.

[0103] The imaging method provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0104] For example, the electronic device in this application embodiment can be an electronic device with shooting function. For example, the electronic device can be a mobile action camera (GoPro), digital camera, tablet computer, desktop, laptop, handheld computer, notebook computer, in-vehicle device, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc. The specific form of the electronic device is not particularly limited in this application embodiment.

[0105] like Figure 3 The diagram shown is a structural schematic of an electronic device 100. The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0106] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0107] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0108] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0109] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0110] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0111] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device. In other embodiments, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0112] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0113] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0114] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0115] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0116] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.

[0117] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0118] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0119] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1.

[0120] In this embodiment, the electronic device may include multiple cameras; for example, the multiple cameras may include a front-facing camera and a rear-facing camera. The front-facing camera is used to capture images of the photographer; the rear-facing camera is used to capture images of the subject being photographed. For example, the electronic device performs gaze detection on the image of the photographer captured by the front-facing camera to determine the direction of the photographer's gaze, thereby obtaining the region of interest (ROI). Then, the electronic device controls the rear-facing camera to determine the focus object in the image of the subject being photographed based on the ROI, thereby achieving focusing.

[0121] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device is selecting a frequency, a DSP can perform a Fourier transform on the frequency energy.

[0122] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, the electronic device can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0123] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0124] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0125] Audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, audio module 170 may be located in processor 110, or some functional modules of audio module 170 may be located in processor 110. Speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Receiver 170B, also called a "handset," is used to convert audio electrical signals into sound signals. Microphone 170C, also called a "microphone" or "microphone," is used to convert sound signals into electrical signals.

[0126] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0127] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, audio and video files can be stored on the external memory card.

[0128] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 121. For example, in this embodiment, processor 110 can execute instructions stored in internal memory 121, which may include a program storage area and a data storage area.

[0129] The program storage area can store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area can store data created during the use of the electronic device (such as audio data, phonebook, etc.). Furthermore, the internal memory 121 can include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0130] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device. The electronic device can support one or N SIM card interfaces, where N is a positive integer greater than 1. SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.

[0131] The methods described in the following embodiments can all be implemented in the electronic device 100 having the above-described hardware structure. The following embodiments use a mobile phone as an example to specifically illustrate the technical solutions provided by the embodiments of this application.

[0132] The mobile phone provided in this application embodiment is equipped with cameras of different focal lengths. In some embodiments, the back of the mobile phone includes four rear cameras, and the front of the mobile phone includes two front cameras. For example, as shown... Figure 4 As shown, the four rear cameras can be, for example, a first rear camera 201, a second rear camera 202, a third rear camera 303, and a fourth rear camera 304; the two front cameras can be, for example, a first front camera 301 and a second front camera 302.

[0133] In some embodiments, the first rear camera 201 may be a main camera (or rear main camera); the second rear camera 202 may be a macro camera; the third rear camera 203 may be a wide-angle camera; and the fourth rear camera 204 may be a depth camera. The second rear camera 202, the third rear camera 203, and the fourth rear camera 204 may also be referred to as rear secondary cameras. In some embodiments, the first front camera 301 may be referred to as a main camera, and the second front camera 302 may be referred to as a front secondary camera; or, the first front camera 301 may be referred to as a secondary camera, and the second front camera 302 may be referred to as a front main camera. This application does not limit this aspect.

[0134] For example, a camera app (or an app with shooting capabilities) can be installed on the mobile phone. In some embodiments, when the mobile phone runs the camera app, the mobile phone displays a first preview image. The first preview image can be a photo preview image or a video preview image. The video preview image can be a preview image before recording video, or a preview image during the recording process. The first preview image includes a first focus area.

[0135] In some embodiments, the first focus area is determined by the phone's autofocus (AF). It should be noted that examples of autofocus can be found in relevant technical documents, and will not be repeated here.

[0136] In other embodiments, the first focus area is determined by the phone after focusing based on the ROI (Region of Interest) of the photographer's gaze. For example, when the phone runs a camera application, it simultaneously activates the rear main camera (e.g., the first rear camera 201) and the front camera (e.g., the first front camera 301 or the second front camera 302). The front camera is used to capture the user image of the photographer, and the rear main camera is used to capture the image of the subject. The phone performs gaze detection on the user image captured by the front camera to determine the direction of the photographer's gaze, thus obtaining the ROI of the gaze. Then, the phone controls the rear main camera to focus based on the ROI of the gaze.

[0137] After the phone displays the first preview image, it acquires a user image of the photographer captured by the front-facing camera and performs gaze detection based on this image to determine the direction of the photographer's gaze, thus obtaining the ROI of the gaze. Then, the phone controls the rear main camera to focus based on the ROI of the gaze and displays a second preview image. This second preview image includes a second focus area, which is different from the first focus area.

[0138] In this way, the phone performs gaze detection on the user image captured by the front-facing camera to determine the direction of the photographer's gaze, i.e., the position of the photographer's line of sight on the display screen (i.e., the phone screen). Then, the phone determines the Region of Interest (ROI) of the gaze based on the area corresponding to the direction of the photographer's gaze in the first preview image. This allows the phone to switch the focus area in real time based on the ROI, thereby improving the shooting effect. In other words, the solution described in this application embodiment can solve the problem in related technologies where the photographer needs to manually operate the phone screen when switching the focus object, which affects the shooting effect.

[0139] In some embodiments, the mobile phone can use eye-tracking technology to detect the user image of the photographer captured by the front-facing camera to determine the ROI of eye gaze. In other embodiments, the mobile phone can use image capture or scanning extraction functions to detect the user image of the photographer captured by the front-facing camera to determine the photographer's facial information (or facial contour information) and eye information (or pupil information); then, the mobile phone inputs the facial information and eye information into a preset model to obtain the ROI of eye gaze. Examples of facial information and eye information can be found in the following embodiments.

[0140] It should be noted that the mobile phone can run the camera application by receiving input from the user. For example, this input can be one of the following: touch operation, button operation, gesture operation, or voice operation. Touch operation, for example, can be a tap or a swipe.

[0141] In some embodiments, the mobile phone includes multiple shooting modes, and the preview images generated by the mobile phone in different shooting modes have different effects. For example, the multiple shooting modes include one or more of the following: photo mode, video mode, portrait mode, large aperture mode, slow motion mode, and panorama mode. When the mobile phone enters large aperture mode, it can activate another rear camera as a secondary rear camera to capture the original image. This secondary rear camera (or secondary rear camera) can be a second rear camera 202, a third rear camera 203, or a fourth rear camera 204; this embodiment does not limit the specific camera used.

[0142] It should be noted that the preview image generated by the phone in large aperture mode includes both the sharp and blurred portions. The focused object is the sharp portion of the image, while other objects are blurred.

[0143] The following describes in detail the technical solution provided in this application embodiment, taking the scenario of recording video with a mobile phone as an example. In order to enrich the style or effect of mobile phone video recording, the mobile phone can use a movie mode to record video.

[0144] In some embodiments, such as Figure 5 As shown in (1), in response to the photographer's operation of the camera application icon 401 on the main screen of the phone, the phone displays as shown in the image. Figure 5 Interface 402 is shown in (2). Interface 402 is the preview interface for mobile phone photography. Interface 402 also includes portrait mode, video recording mode, movie mode and professional mode. Among them, movie mode is a mode for mobile phone video recording. Movie mode includes multiple LUTs, and different LUTs correspond to different shooting scenes. The photographer can select the corresponding LUT according to different shooting scenes so that the images and styles (or effects) of different shooting scenes are different, thereby enriching the style (or effect) of mobile phone shooting and making the shooting style more diverse and personalized.

[0145] Still Figure 5 As shown in (2), in response to the photographer's selection of movie mode 403, the phone displays the following: Figure 6 Interface 404 is shown in Figure (1). Interface 404 is the preview interface before recording video on the phone. In interface 404, the phone displays a prompt message 405. This prompt message 405 is used to remind the photographer to put the phone in landscape mode. For example, the prompt message 405 could be "Movie mode landscape shooting effect is better." Then, when the photographer places the phone in landscape mode, the phone displays as shown... Figure 6The interface 406 shown in (2) is a preview interface before recording video on a mobile phone in landscape mode. In some embodiments, the interface 406 further includes a LUT control 407 and a large aperture control 408.

[0146] In some embodiments, when the phone enters movie mode, the phone displays a first preview image, which includes a first focus area. It should be noted that the examples illustrating the phone displaying the first preview image and how to determine the first focus object can be found in the above embodiments, and will not be repeated here.

[0147] After the phone displays the first preview image, or when the phone starts recording video in movie mode, the phone can acquire a user image of the photographer captured by the front-facing camera, and perform gaze detection based on this user image to determine the direction of the photographer's gaze, thus obtaining the ROI of the gaze. Then, the phone controls the rear main camera to focus based on the gaze ROI and displays a second preview image. The second preview image includes a second focus area, which is different from the first focus area.

[0148] It should be understood that in the embodiments of this application, the ROI of gaze refers to the area on the mobile phone screen where the photographer's gaze is located, corresponding to the area of ​​the first preview image.

[0149] For example, such as Figure 7 As shown, the front-facing camera is connected to both a face recognition module and a gaze detection module, with the face recognition module also connected to the gaze detection module. In some embodiments, the front-facing camera inputs the captured user image into the gaze detection module, which then detects the user image to determine the ROI of the gaze. For example, the gaze detection module uses gaze tracking technology to track the gaze of the photographer in the user image, thereby determining the ROI of the gaze. In other embodiments, the front-facing camera inputs the captured user image into the face recognition module, which performs face recognition processing on the user image to determine the ROI of the face. For example, the ROI of the face includes face information and eye information. For example, face information may include, for instance, the coordinates of the facial contour; eye information may include, for instance, one or more of the following: interpupillary distance, pupil size, pupil size variation, pupil brightness contrast, corneal radius, spot information, and iris information. Of course, eye information may also include other features used to characterize subtle changes in the eyes, which will not be listed here. In some embodiments, the face recognition module can determine the facial and eye information of the photographer through functions such as image capture or scanning extraction.

[0150] Then, the face recognition module inputs facial and eye information into the gaze detection module. The gaze detection module performs gaze detection based on the facial and eye information, thereby outputting the ROI of the gaze. For example, the gaze detection module includes a preset model; wherein, this preset model is pre-trained by the mobile phone based on sample facial and eye information. Based on this, the target gaze detection module inputs the facial and eye information into the preset model and outputs the ROI of the gaze.

[0151] It should be noted that the front-facing camera inputs the user's image to the face recognition module, which performs face recognition processing on the user's image to determine the ROI of the face. In this way, the gaze detection module does not need to detect the entire user image, but only needs to detect the ROI image related to the face (i.e., the face ROI), thereby narrowing the detection range, improving detection accuracy, and reducing device power consumption.

[0152] For example, such as Figure 8 As shown, the rear main camera is connected to both the first ISP front-end module and the AF algorithm module, and the AF algorithm module is connected to the gaze detection module. For example, the gaze detection module inputs the ROI of the target gaze into the AF algorithm module. The AF algorithm module controls the rear main camera to focus based on the ROI of the gaze. For instance, the AF algorithm module determines the second focus area based on the area corresponding to the photographer's gaze position on the display screen and the first preview image; then, the AF algorithm module controls the rear main camera to focus the object included in the second focus area onto the focal point of the rear main camera, thereby achieving focus. Based on this, after the rear main camera completes focusing, it inputs the acquired raw image frame into the first ISP front-end module, which performs target processing on the raw image frame, converting it into a target format raw image frame (or target raw image frame).

[0153] For example, target processing can be "YUV domain" processing, and the original image frame in the target format can be, for example, a raw image frame in YUV format. Here, YUV format is an image color encoding method, where Y represents luminance, and U and V represent chrominance.

[0154] Still Figure 8As shown, the first ISP front-end module is connected to both the first and second image stabilization modules. The first image stabilization module is connected to the first ISP back-end module, and the second image stabilization module is connected to the second ISP back-end module. For example, the first ISP front-end module can be connected to the first and second image stabilization modules via a video graphics array (VGA) interface, the first image stabilization module can be connected to the first ISP back-end module via a VGA interface, and the second image stabilization module can be connected to the second ISP back-end module via a VGA interface. In this way, the ISP front-end module, image stabilization module, and ISP back-end module connected via the VGA interface can output full high definition (FHD) images, thereby improving the clarity of the preview image.

[0155] In conjunction with the above embodiments, such as Figure 8 As shown, after the first ISP front-end module converts the original image frame into the target original image frame, it splits the target original image frame into two data outputs. One data stream is a preview stream, which is output to the display screen so that the display screen shows the second preview image. The other data stream is a video stream, which is used to generate a video file and save the video file on the mobile phone. For example, the mobile phone encodes the video stream data to generate a video file.

[0156] For example, such as Figure 8 As shown, the first ISP front-end module inputs the target raw image frame into the first stabilization module, which performs electronic image stabilization (EIS) processing on the target raw image frame. The first stabilization module then inputs the processed target raw image frame into the first ISP back-end module, which performs image enhancement on the target raw image frame and outputs a preview stream. Subsequently, the first ISP back-end module outputs the preview stream to the display screen, which displays a second preview image based on the preview stream.

[0157] Accordingly, the first ISP front-end module inputs the target raw image frame to the second stabilization module, which performs EIS processing on the target raw image frame. The second stabilization module then inputs the processed target raw image frame to the second ISP back-end module, which enhances the target raw image frame and outputs a video stream. Based on this, after the phone finishes recording video, the phone encodes the video stream output by the second ISP back-end module to generate a video file.

[0158] In some embodiments, the second image stabilization module performs EIS delay processing on the target raw image frames. That is, the second image stabilization module can buffer multiple target raw image frames, and perform EIS processing on each of these multiple target raw image frames to achieve better image stabilization. For example, the first and second image stabilization modules can be equipped with inertial measurement units (IMUs) to perform EIS processing on the target raw image frames.

[0159] In some embodiments, a preset image algorithm can be configured in the first ISP front-end module to process the target raw image frame. For example, such as... Figure 8 As shown, the first ISP front-end module has a pre-set Graph Transformation Matching (GTM) algorithm module; the GTM module is used to perform target processing on the original target image frame. The first ISP back-end module and the second ISP back-end module have pre-set Render Image Rendering (WRAP) algorithm modules; the WRAP module is used to enhance the original target image frame.

[0160] In some embodiments, to further improve the shooting effect, the photographer can select a large aperture mode to record video according to specific needs, so that the recorded video file has the effect of clearly displaying the focused object and blurring other objects. Based on this, when the phone is shooting in large aperture mode, the phone detects the direction of the photographer's gaze in the user image captured by the front camera, and determines the ROI of the gaze; then, the phone switches the first focus object in the first preview image according to the ROI of the gaze, thus achieving focusing based on the ROI of the photographer's gaze.

[0161] For example, with the main rear camera and front camera already activated, when the phone enters large aperture mode, in order to display a preview image with a bokeh effect, the phone needs to activate the secondary rear camera based on the user's selection of large aperture mode, and control the main rear camera and secondary rear camera to perform focusing processing through the ROI of the user's gaze. Furthermore, the phone needs to perform depth-of-field calculations using the original image frames captured by the main rear camera and secondary rear camera, and generate a second preview image with a bokeh effect based on the original image frame captured by the main rear camera. It should be noted that, for ease of understanding, in the following embodiments of this application, the original image frame captured by the main rear camera can be referred to as the first original image frame, and the original image frame captured by the secondary rear camera can be referred to as the second original image frame.

[0162] In some embodiments, the two rear cameras, a main rear camera and a secondary rear camera, can form a logical camera unit (Logical Camera Id) for depth calculation. For example, the main rear camera can be a mid-range camera (or a standard camera), and the secondary rear camera can be a wide-angle camera or a telephoto camera. Furthermore, the logical camera can be a combination of a mid-range camera and a wide-angle camera, or a combination of a mid-range camera and a telephoto camera.

[0163] In some embodiments, such as Figure 9a As shown in (1), the mobile phone responds to the photographer's operation of the large aperture control 408, displaying as shown in the image. Figure 9a Interface 409 is shown in (2). This interface 409 is the interface when the phone enters the large aperture mode. For example, as shown in (2). Figure 9a As shown in Figure (2), interface 409 includes a control 410 for adjusting the aperture size. For example, when the mobile phone receives an operation from the photographer on control 410 (e.g., the photographer can slide a number or dot icon in control 410), the mobile phone adjusts the aperture size of the preview image in interface 409. It should be noted that the aperture size determines how much of the image is clearly displayed. In some embodiments, the aperture size is adjusted between f0.95 and f16; wherein, the smaller the aperture value, the larger the aperture, and the more of the image is clearly displayed. Figure 9a The aperture size shown in (2) is f4.

[0164] In some embodiments, after the mobile phone adjusts the aperture size of the preview image in the interface 409 through the photographer's operation of the control 410, the mobile phone can receive the photographer's operation of the control 410 again and retract the control 410 for adjusting the aperture size displayed in the interface 409.

[0165] In other embodiments, after the mobile phone adjusts the aperture size of the preview image in the interface 409 by the photographer's operation of the control 410, if the mobile phone does not detect the photographer's operation of the control 410 within a preset time period (such as 5s, 1s, etc.), the mobile phone automatically retracts the control 410 for adjusting the aperture size displayed in the interface 409.

[0166] Subsequently, if the electronic device needs to adjust the aperture size of the preview image again, the electronic device can receive the photographer's operation on the control 408 again, display the control 410 for adjusting the aperture size, and adjust the aperture size by the user's operation on the control 410.

[0167] It should be noted that in the following embodiments, the control 410 for adjusting the aperture size displayed on the phone's retractable interface 409 is used as an example for illustration. In this way, the influence of the control 410 on the preview image displayed on the phone's interface 409 can be avoided.

[0168] In some embodiments, when the phone is not in large aperture mode, the second preview image displayed by the phone is always the area with clear visibility. For example, such as... Figure 9b As shown, the second preview image displayed on the phone includes a "tree" and a "bird"; by Figure 9b As can be seen, the "tree" and "bird" in the second preview image displayed on the phone are both clearly visible objects.

[0169] In other embodiments, when the phone enters large aperture mode, the second preview image displayed by the phone includes a sharp display area and a blurred display area. The sharp display area corresponds to the second focus area. This means that the focused object included in the second focus area is the portion that is clearly visible. In some embodiments, the sharpest part of the second focus area is called the in-focus area; that is, the in-focus position is the sharpest, and the farther away from the focal plane of the in-focus area, the greater the degree of blurring.

[0170] For example, in the second preview image, when the focused object included in the second focus area is in the foreground portion, the foreground is clearly displayed and the background is blurred; or, when the focused object included in the second focus area is in the background portion, the background is clearly displayed and the foreground is blurred. For example, as... Figure 10 As shown, when the second focus area includes a "bird" as the focus object, and the focus object is in the foreground, by... Figure 10 As can be seen, the "bird" in the foreground is clearly visible, while the "tree" in the background is blurred. For example, as... Figure 11 As shown, when the second focus area includes a "tree" as the focus object, and the focus object is in the background, by Figure 11 As can be seen, the "tree" in the background is clearly visible, while the "bird" in the foreground is blurred.

[0171] It should be noted that, Figure 10 and Figure 11 The use of "dashed lines" to represent blurred display and "solid lines" to represent sharp display is for illustrative purposes only and does not constitute a limitation on the blurred and sharp display methods in this application. The effect of blurred display is subject to the specific implementation.

[0172] For example, after the phone enters large aperture mode, the front-facing camera captures the user's image; the phone performs facial recognition processing on the user's image to generate the ROI (Region of Interest) of the face. Then, the phone performs gaze detection based on the ROI of the face to generate the ROI of gaze. Based on this, the phone controls the rear main camera and the rear secondary camera to perform AF focusing processing according to the ROI of gaze, determining the second focus area.

[0173] It should be noted that for an example illustrating how the phone determines the ROI of eye gaze based on the user's image captured by the front-facing camera in large aperture mode, please refer to [link to example]. Figure 7 As described in the above embodiments, they will not be listed one by one here.

[0174] In some embodiments, such as Figure 12 As shown, the rear main camera and the rear secondary camera are connected to the AF algorithm module, which in turn is connected to the eye gaze detection module. For example, the eye gaze detection module inputs the ROI (Region of Interest) of eye gaze into the AF algorithm module. The AF algorithm module controls the rear main camera and the rear secondary camera to focus based on the ROI. For instance, the AF algorithm module determines the second focus area based on the area corresponding to the photographer's gaze position on the display screen and the first preview image; then, the AF algorithm module controls the rear main camera and the rear secondary camera to focus the object included in the second focus area onto the focal point of the rear main camera and the rear secondary camera, thereby achieving focus.

[0175] Still Figure 12 As shown, the rear main camera is connected to the first ISP front-end module, and the rear secondary camera is connected to the second ISP front-end module; the second ISP front-end module is also connected to the preprocessing module. Simultaneously, the first ISP front-end module and the preprocessing module are also connected to the depth-of-field processing module. For example, the rear main camera inputs a first raw image frame to the first ISP front-end module, which performs target processing on the first raw image frame, converting it into a first raw image frame of a target format (or target first raw image frame). Correspondingly, the rear secondary camera inputs a second raw image frame to the second ISP front-end module, which performs target processing on the second raw image frame, converting it into a second raw image frame of a target format (or target second raw image frame).

[0176] For example, the first ISP front-end module and the second ISP front-end module are pre-configured with a GTM module, which is used to perform target processing on the first raw image frame and the second raw image frame.

[0177] In some embodiments, the smaller the difference in field of view (FOV) between the rear main camera and the rear secondary camera, the more accurate the depth of field calculated by the phone will be. Based on this, in this embodiment, the target second original image frame can be preprocessed to make its FOV the same as that of the target first original image frame, thereby ensuring that the image frames output by the rear main camera and the rear secondary camera have consistent effects, which is beneficial to improving the accuracy of depth of field calculation. For example, the second ISP front-end module inputs the target second original image frame into the preprocessing module, which performs FOV switching on the target second original image frame to make its FOV the same as that of the target first original image frame. For example, assuming the FOV of the target first original image frame output by the rear main camera is 85° and the FOV of the target second original image frame output by the rear secondary camera is 115°. After the second original image frame of the target output by the rear main camera is processed by the preprocessing module, the field of view of the second original image frame of the target can be, for example, 85°, which is consistent with the field of view of the first original image frame of the target.

[0178] In some embodiments, still as Figure 12 As shown, the preprocessing module inputs the processed second original image frame of the target to the depth calculation module; at the same time, the first ISP front-end module inputs the first original image frame of the target to the depth calculation module; the depth calculation module calculates the depth based on the second original image frame and the first original image frame of the target.

[0179] In conjunction with the above embodiments, such as Figure 12 As shown, after the first ISP front-end module converts the first raw image frame into a target first raw image frame, it splits the target first raw image frame into two data outputs. One data stream is a preview stream, which is output to the display screen so that the display screen shows a second preview image on the preview interface. The other data stream is a video stream, which is used to generate a video file and save the video file in the mobile phone. For example, the mobile phone encodes the video stream data to generate a video file.

[0180] For example, such as Figure 12 As shown, the first ISP front-end module inputs the first original image frame of the target to the first image stabilization module, which performs EIS processing on the first original image frame of the target. The first image stabilization module then inputs the processed first original image frame of the target to the first ISP back-end module, which performs image enhancement on the first original image frame of the target and outputs a preview stream. Subsequently, the first ISP back-end module outputs the preview stream to the display screen, which displays a second preview image based on the preview stream.

[0181] Accordingly, the first ISP front-end module inputs the target first raw image frame to the second stabilization module, which performs EIS processing on the target first raw image frame. The second stabilization module then inputs the processed target first raw image frame to the second ISP back-end module, which performs image enhancement on the target first raw image frame and outputs a video stream. Subsequently, when the phone finishes recording video, the phone encodes the video stream output by the second ISP back-end module to generate a video file.

[0182] It should be noted that the second image stabilization module can perform EIS delay processing on the first original image frame of the target. For an example of EIS delay processing, please refer to the above embodiments. In addition, for an example of the first image stabilization module and the second image stabilization module, please refer to the above embodiments. They will not be described in detail here.

[0183] For example, the first ISP backend module and the second ISP backend module are pre-configured with a WRAP module; the WRAP module is used to perform image enhancement on the target first original image frame and the target second original image frame.

[0184] In some embodiments, such as Figure 12 As shown, the first ISP backend module and the second ISP backend module also have a preset bokeh algorithm module. The bokeh algorithm module is used to blur the target first original image frame according to the depth of field and the ROI of eye gaze. For example, the bokeh algorithm module is connected to the eye gaze detection module and the depth of field algorithm module; the eye gaze detection module inputs the ROI of eye gaze into the bokeh algorithm module, and the depth of field algorithm module inputs the calculated depth of field into the bokeh algorithm module. The bokeh algorithm module blurs the target first original image frame according to the ROI of eye gaze and the depth of field. For example, the bokeh algorithm module uses the second focus area corresponding to the ROI of eye gaze as the subject plane, the subject plane is clearly displayed, and the area outside the subject plane is blurred. In some embodiments, the bokeh algorithm module can also determine the range of blurred display within the subject plane according to the depth of field. In some embodiments, the strategy adjusted by the bokeh algorithm module is: the subject plane is clearly displayed, and the area outside the subject plane is blurred to a higher degree the farther away from the focal plane.

[0185] It should be understood that after the AF algorithm module focuses on the ROI (Region of Interest) based on the gaze, the determined second focus area is located within the subject plane, and this second focus area happens to fall at the focal point. Based on this, the area outside the subject plane is blurred, which can also be called bokeh, i.e., blurred image outside the focal point; or it can be called a virtual image, out of focus, etc.

[0186] In some embodiments, the first ISP backend module and the second ISP backend module further include a zoom module, which is used to zoom the first original image frame of the target after it has been blurred. For example, the mobile phone can receive a zoom operation input by the photographer to adjust the size of the second preview image displayed on the screen. The zoom operation instructs the mobile phone's screen to display the second preview image corresponding to the target zoom ratio. In some embodiments, before the mobile phone receives a zoom operation input by the photographer, the zoom ratio displayed by the mobile phone can be the reference zoom ratio of the rear main camera (e.g., 1×).

[0187] It should be noted that the aforementioned zoom ratio can be either optical zoom or digital zoom. For example, the zoom ratio can be 1×, 3×, 4×, 4.5×, 4.9×, or 5×, etc. Here, "1×" represents a zoom ratio of 1x; "3×" represents a zoom ratio of 3x; and "4×" represents a zoom ratio of 4x. Furthermore, the zoom ratio in this embodiment can also be referred to as a magnification factor. That is, the aforementioned zoom ratio can also be called a zoom magnification factor.

[0188] For example, when a mobile phone records video in large aperture mode, the phone displays as follows: Figure 13 The interface 411 shown includes a zoom control 412 for adjusting the zoom ratio. For example, the interface 411 currently displays a zoom ratio of 4.5×. When the phone responds to the photographer's "+" operation on the zoom control 412, the phone increases the current zoom ratio, for example, to 5.0×. When the phone responds to the photographer's "-" operation on the zoom control 412, the phone decreases the current zoom ratio, for example, to 4.0×.

[0189] For example, when the mobile phone receives the zoom operation input by the photographer, the preset zoom modules in the first ISP backend module and the second ISP backend module enlarge (or reduce) the target first original image frame according to the target zoom ratio corresponding to the zoom operation, so that the second preview image finally displayed by the mobile phone is enlarged (or reduced).

[0190] In some embodiments, when the phone focuses based on the ROI (Region of Interest) of the viewer's gaze, the determined focus area may not match the area the photographer wants to focus on, thus affecting the photographer's experience. Therefore, to improve the accuracy and stability of the phone's focusing based on the direction of the photographer's gaze, the phone can mark the ROI of the viewer's gaze in the preview interface (which can be the interface before video recording, during video recording, or the interface when taking a photo). That is, the phone provides feedback on the ROI of the viewer's gaze to the photographer through the user interface (UI), both to indicate the location of the photographer's current gaze (i.e., line of sight) on the display screen and to guide the photographer to focus their gaze on the target location on the display screen, i.e., the area the photographer wants to focus on.

[0191] In some embodiments, such as Figure 14 As shown, the preview interface displays a second preview image and a user image. The second preview image is generated by the phone based on a first raw image frame captured by the rear camera; the user image is generated by the phone based on an image of the photographer captured by the front camera. For example, the preview interface includes a mask area used to display the user image.

[0192] It should be noted that the mask area refers to the area or process by which the mobile phone uses a selected image, graphic, or object to occlude (fully or partially) the second preview image, thereby controlling the processing of the second preview image. The specific image, graphic, or object used for occlusion is called a mask. In image processing, the mask can be a film, filter, etc. In the embodiments of this application, the specific image, graphic, or object can be the user image described in the above embodiments.

[0193] In this embodiment, the specific shape and location of the mask area are not limited. The shape of the mask area can be, for example, a square, a rectangle, a circle, or other regular or irregular shapes. Figure 14 The following is an example of a mask area with a quadrilateral shape (which can be a rectangle or a square).

[0194] Still Figure 14 As shown, in some embodiments, the mobile phone can divide the mask area into multiple preset ROIs; each preset ROI corresponds to a portion of the second preview image; in this way, the portions of the second preview images corresponding to the multiple preset ROIs are merged to form the entire second preview image.

[0195] It should be noted that the specific number of multiple preset ROIs is not limited in the embodiments of this application, and shall be determined according to actual needs. Figure 14 The following example illustrates how a mobile phone divides the mask area into nine preset ROIs.

[0196] For example, such as Figure 14 As shown, when the phone detects that the ROI the photographer is looking at corresponds to the position of the "bird" in the second preview image, the phone determines the target preset ROI among multiple preset ROIs based on the correspondence between multiple preset ROIs and the second preview image. It should be understood that the area corresponding to the target preset ROI in the second preview image is the same area on the display screen where the ROI the photographer is looking at corresponds to the area in the second preview image. Based on this, the phone marks the target preset ROI, that is, the phone marks the ROI the photographer is looking at in the mask area. For example, Figure 14 The ROI being observed is represented by a filled box. Of course, in practice, the ROI being observed can also be represented by a thicker border or by highlighting.

[0197] In some embodiments, based on the ROI being viewed by the photographer, the area of ​​the target preset ROI marked by the mobile phone corresponds to at least two preset ROI areas. For example, Figure 14 The photographer's gaze is directed towards the "bird." At this point, the phone, based on the correspondence between multiple preset ROIs and the second preview image, determines that the target preset ROI corresponds to two of the preset ROIs. For example, the target preset ROI marked by the phone covers a portion of both preset ROIs. Specifically, the ratio of the target preset ROI marked by the phone to the area of ​​the two preset ROIs can be 8:2. That is, the target preset ROI occupies approximately 80% of the area of ​​one preset ROI and approximately 20% of the area of ​​the other.

[0198] For example, such as Figure 15 As shown, when the phone detects that the ROI the photographer is looking at corresponds to the position of the "tree" in the second preview image, the phone determines the target preset ROI among multiple preset ROIs based on the correspondence between multiple preset ROIs and the second preview image. It should be understood that the area corresponding to the target preset ROI in the second preview image corresponds to the ROI the photographer is looking at (i.e., the area on the screen where the photographer's gaze is directed). Based on this, the phone marks the target preset ROI; that is, the phone marks the ROI the photographer is looking at within the mask area. For example, Figure 15 The ROI where the gaze is focused is represented by a filled box.

[0199] For example, when the phone detects that the direction of the gaze is at the location of a "tree," the phone determines that the focus object is the tree. However, the object the photographer actually wants to focus on is not the "tree." Based on this, the photographer can adjust the direction of their gaze according to the ROI marked in the preview interface. For instance, the photographer can use the ROI marked in the preview interface and the object they actually want to focus on to move (e.g., left, right, up, or down) the direction of their gaze accordingly, so that the direction of their gaze ultimately falls within the focus area they want to focus on.

[0200] In some embodiments, such as Figure 16 As shown, the preview interface displays a second preview image and a mask image. The preview image is generated by the phone based on a first original image frame captured by the rear camera; the mask image is generated by the phone after scaling down the second preview image. For example, the preview interface includes a mask area; this mask area is used to display the mask image.

[0201] It should be noted that the explanation of the mask area and the examples of its specific shape can be found in the above embodiments, and will not be repeated here.

[0202] Still Figure 16 As shown, in some embodiments, the mobile phone can divide the mask area into multiple preset ROIs; each preset ROI corresponds to a portion of the second preview image; in this way, the portions of the second preview images corresponding to the multiple preset ROIs are merged to form the entire second preview image. Figure 16 The following is an example of how a mobile phone divides the mask area into 9 preset ROIs.

[0203] For example, such as Figure 16 As shown, when the phone detects that the ROI the photographer is looking at corresponds to the position of the "bird" in the second preview image, the phone determines the target preset ROI among multiple preset ROIs based on the correspondence between multiple preset ROIs and the second preview image. It should be understood that the area corresponding to the target preset ROI in the second preview image corresponds to the ROI the photographer is looking at (i.e., the area on the screen where the photographer's gaze is directed). Based on this, the phone marks the target preset ROI; that is, the phone marks the ROI the photographer is looking at within the mask area. For example, Figure 16 The ROI where the gaze is focused is represented by a filled box.

[0204] In some embodiments, such as Figure 17 As shown, the phone divides the preview interface into multiple preset ROIs (…). Figure 17(Taking a division into 9 preset ROIs as an example). For instance, when the phone detects that the ROI the photographer is looking at corresponds to the position of the "bird" in the second preview image, the phone determines the target preset ROI among the multiple preset ROIs based on the correspondence between the preset ROIs and the second preview image. It should be understood that the area corresponding to the target preset ROI in the second preview image corresponds to the ROI the photographer is looking at (i.e., the area on the screen where the photographer's gaze is directed). Based on this, the phone marks the target preset ROI; that is, the phone marks the ROI the photographer is looking at in the preview interface. For example, Figure 17 Use a filled box to represent the ROI where the gaze is focused.

[0205] It should be noted that, Figures 14-17 The UI interfaces shown are merely illustrative examples of embodiments in this application and do not constitute a limitation of this application. Other UI interfaces that use the schemes described in the embodiments of this application to mark the ROI that the eye is focused on should also fall within the protection scope of the embodiments of this application.

[0206] In some embodiments, when the distance between the photographer and the mobile phone is greater than a first preset value (e.g., the distance between the photographer and the mobile phone is too far), or when the distance between the photographer and the mobile phone is less than a second preset value (e.g., the distance between the photographer and the mobile phone is too close), the mobile phone may not be able to detect the direction of the photographer's gaze, or the direction of the photographer's gaze detected by the mobile phone may not be accurate enough, thereby affecting the accuracy of the mobile phone's focusing. Based on this, in some embodiments, during the shooting process, if the mobile phone detects that the distance between the photographer and the mobile phone is too far (or too close), the mobile phone can prompt the photographer with a prompt message, so that the photographer can adjust the distance between themselves and the mobile phone according to the prompt message, thereby ensuring the accuracy of the mobile phone's focusing based on the direction of the photographer's gaze.

[0207] Combination Figure 14 and Figure 15 ,like Figure 18 As shown, for example, the mobile phone has a preset face area divided within the mask area, and the photographer can adjust the distance between themselves and the phone based on this preset face area. For instance, the photographer can adjust the distance based on the preset face area so that their facial image falls precisely within the preset face area. Figure 18 The preset face area is represented by a dashed box.

[0208] In some embodiments, combined with Figures 14-17The phone can also display text prompts to remind the photographer to adjust the distance between the phones; alternatively, it can play voice prompts to remind the photographer to adjust the distance between the phones. In other embodiments, the phone can also use a combination of text and voice prompts to remind the photographer to adjust the distance between the phones.

[0209] For example, such as Figure 19 As shown, the phone displays a text prompt in the preview interface, while simultaneously playing a voice prompt through its speaker. For example, the text prompt (or voice prompt) could be: "Please bring your face close to (or away from) the phone screen." Figures 14-15 In the illustrated embodiment, the text prompt (or voice prompt) could be, for example, "Please place your face within the dashed box."

[0210] by Figure 14 The second preview image shown is used as an example. For instance, before the phone starts recording video, the phone displays something like this. Figure 20 The preview interface 413 shown in Figure (1) includes a recording control 414. Subsequently, in response to the photographer's operation of the recording control 414, the phone displays the following... Figure 20 The preview interface 415 shown in (2) is the interface when the mobile phone starts recording video. The preview interface 415 also includes a pause button 416 and an end button 417. For example, in response to the photographer's operation of the pause button 416, the mobile phone pauses the video recording; correspondingly, in response to the photographer's operation of the end recording button 417, the mobile phone ends the video recording and saves the recorded video file in the mobile phone (such as a photo album application).

[0211] Still Figure 20 As shown in (2), after the phone starts recording video, the phone's preview interface 415 displays a third preview image. The third preview image includes a third focus area, where the focused object is the part that is clearly visible. Specifically, the phone determines the ROI (Region of Interest) of the photographer's gaze based on the user image captured by the front-facing camera, and controls the rear-facing camera to focus based on the ROI to determine the third focus area. For example, as shown in (2), the third focus area is determined by the third focus area. Figure 20 As shown in (2), when the object in the third focus area determined by the mobile phone is "bird", the "bird" is clearly displayed in the third preview image displayed by the mobile phone, and the other areas are blurred.

[0212] In some embodiments, during video recording, the phone's preview interface may also include a mask area. This mask area includes the user image captured by the front-facing camera. As can be seen from the above embodiments, the mask area is divided into multiple preset Regions of Interest (ROIs), such as... Figure 20 As shown in (2), when the mobile phone detects that the ROI being gazed at by the photographer corresponds to the position of the "bird" in the third preview image, the mobile phone determines the target preset ROI among the multiple preset ROIs based on the correspondence between multiple preset ROIs and the third preview image, and marks the target preset ROI. It should be understood that the area corresponding to the target preset ROI in the third preview image is the same as the area corresponding to the ROI being gazed at in the third preview image.

[0213] In this way, the phone can indicate to the photographer the location of the ROI on the screen based on the target preset ROI marked in the mask area; or, it can guide the photographer to focus their gaze on the target location on the screen through the target preset ROI marked in the mask area.

[0214] In some embodiments, the mobile phone can also mark the third focus area using focus markers. These focus markers are used to indicate to the photographer the position of the focused object on the display screen. For example, the focus markers can be focus frames. Figure 20 As shown in (1), the mobile phone can mark the focused objects (such as "birds") included in the third focus area through the focus frame.

[0215] In some embodiments, during video recording, when the phone detects a change in the direction of the photographer's gaze, the phone can redetermine the ROI of the gaze based on the changed direction of the photographer's gaze, and control the rear camera to focus based on the redetermined ROI of the gaze.

[0216] For example, in combination Figure 20 (2) Figure 21 As shown, during video recording, when the phone detects a change in the direction of the photographer's gaze, the preview interface switches from preview interface 415 to preview interface 418. Preview interface 418 displays a fourth preview image, which includes a fourth focus area, where the focused object is the clearly visible portion. Specifically, when the phone detects a change in the direction of the photographer's gaze, it determines the ROI (Region of Interest) of the photographer's gaze based on the user image captured by the front-facing camera, and controls the rear camera to focus based on the ROI to determine the fourth focus area. For example, as... Figure 21As shown, when the fourth focus area determined by the phone includes a "tree," it indicates that the phone has detected a change in the photographer's gaze from "bird" to "tree." At this time, in the fourth preview image displayed by the phone, the "tree" is clearly visible, while the other areas are blurred.

[0217] In some embodiments, such as Figure 21 As shown, when the phone detects that the ROI the photographer is looking at corresponds to the position of the "tree" in the fourth preview image, the phone determines the target preset ROI among multiple preset ROIs based on the correspondence between multiple preset ROIs and the fourth preview image, and marks the target preset ROI. It should be understood that the area corresponding to the target preset ROI in the fourth preview image is the same as the area corresponding to the ROI the photographer is looking at in the fourth preview image.

[0218] In some embodiments, such as Figure 21 As shown, the mobile phone can mark the objects in focus (such as "trees") included in the fourth focus area using focus markers (such as focus frames).

[0219] It should be noted that, in this embodiment of the application, before the mobile phone starts recording video, the mobile phone can also mark the focus objects included in the focus area using focus markers (such as focus frames). See details for further information. Figures 20-21 As described in the above embodiments, they will not be repeated here.

[0220] Considering that the photographer's gaze may frequently change direction during filming (e.g., talking, looking at other objects), if the phone focuses in real-time based on the photographer's gaze and switches the focus area, it would cause frequent switching of the focus area in the generated video file, resulting in frequent switching between sharp and blurred areas, affecting the filming quality. Therefore, in some embodiments, after the phone focuses based on the photographer's gaze and determines the focus area, it can track the objects within that focus area. For example, the phone can identify objects (such as faces, bodies, or prominent subjects) within the focus area and track those successfully identified. When the ROI (Region of Interest) of the photographer's gaze is inconsistent with the tracked object, the phone can use methods such as... Figure 22 The steps shown (such as steps 1-1 to 1-4) are used to refocus and identify a new focus object.

[0221] Step 1-1: When the phone tracks the focused object, the phone starts timing.

[0222] For example, a mobile phone can use a timer to keep track of time.

[0223] Steps 1-2: When the tracked focus object and the ROI being looked at remain consistent within the first preset time period, the phone starts to reset the timer.

[0224] For example, a preset timer reset in a mobile phone means that the timer restarts from zero.

[0225] In some embodiments, the first preset duration may be, for example, X seconds (s). Wherein, X may be, for example, 3 seconds, 5 seconds, or other suitable durations. This application does not limit this, but the specific setting shall prevail.

[0226] Steps 1-3: When the tracked focus object is not the same as the ROI that the eye is looking at, the phone starts timing.

[0227] For example, in combination Figure 14 As shown, the phone focuses by tracking the direction the photographer is looking, thus identifying the object in focus. Figure 14 The image shows a "tree". The phone then tracks the focused object (i.e., tracks the "bird"), and starts a timer. When the tracked focused object (the "bird") and the ROI being looked at remain aligned for a preset time (e.g., 3 seconds), the phone restarts the timer. However, when the phone detects that the ROI being looked at has changed... Figure 14 The ROI shown is switched to the gaze focus. Figure 15 When the ROI is observed, the phone detects that the ROI being observed after the eye movement is different from the tracked focus object. For example, the focus object corresponding to the ROI being observed after the eye movement is changed should be a "tree," while the tracked focus object should be a "bird." Based on this, the phone starts timing.

[0228] Steps 1-4: If the tracked focus object and the ROI of the gaze are still inconsistent within the second preset time period, the phone will re-determine the focus object based on the ROI of the gaze after the switch, and repeat steps 1-1 to 1-3.

[0229] For example, the second preset duration can be Y (s). Y can be, for example, 3 seconds, 5 seconds, or other suitable durations. This application does not limit this, but the specific setting shall prevail.

[0230] For example, such as Figure 23As shown in (1), assume that the focus area determined by the mobile phone based on the ROI (Region of Interest) of the photographer's gaze includes a "tree". The mobile phone then tracks this focus object (e.g., tracks the "tree") and starts timing. When the mobile phone detects that the ROI of the gaze is inconsistent with the tracked focus object, it starts timing again. If the mobile phone detects that the ROI of the gaze is still inconsistent with the tracked focus object within a second preset time period, it re-determines the focus object based on the ROI of the gaze. For example, as... Figure 23 As shown in (2), when the mobile phone detects that the ROI of the gaze is still inconsistent with the tracked focus object within 3 seconds, the mobile phone re-determines the focus object based on the ROI of the gaze, such as the re-determined focus object being a "bird".

[0231] For example, combining Figure 23 Neutral (1) and Figure 23 As shown in Figure (2), during video recording, the phone detected at 3 seconds that the ROI being gazed at was inconsistent with the tracked focus object. However, the phone did not refocus on the object and switch, but instead started timing. Afterward, when the phone still detected that the ROI being gazed at was inconsistent with the tracked focus object within 3 seconds, the phone refocused on the object based on the ROI being gazed at. And at 6 seconds, the focus object was switched from "tree" to "bird".

[0232] In some embodiments, when the phone detects that the photographer's gaze is directed outside the phone screen, the phone maintains the tracked focus object consistent with the ROI determined in the previous moment. That is, in the preview image displayed on the phone, the tracked focus object is clearly displayed, while untracked objects are blurred.

[0233] It should be noted that the above embodiments are merely illustrative examples of the phone controlling the rear camera's focus based on the direction the photographer is looking, and do not constitute a limitation of this application. It should be understood that technical solutions where the phone performs other functions based on the direction the photographer is looking should also fall within the protection scope of this application. For example, the phone can also implement zoom functions, or pause / resume shooting functions, etc., based on the direction the photographer is looking.

[0234] It should be understood that the mobile phone shooting scenarios described in the embodiments of this application include photo shooting scenarios and video recording scenarios (or video recording scenarios). Based on this, the mobile phone's control of the rear camera's focus according to the direction of the photographer's gaze in the above embodiments can be applied to both photo shooting scenarios and video recording scenarios. The video recording scenario can be the scene before video recording (i.e., before the photographer clicks the recording control 414); the video recording scenario can also be the scene during video recording (i.e., when the photographer clicks the recording control 414 and the mobile phone starts recording video).

[0235] This application provides a shooting method that can be applied to an electronic device, which includes a rear first camera, a rear second camera, a front camera, and a display screen. Figure 24 This is a flowchart illustrating a shooting method provided in an embodiment of this application, as shown below. Figure 24 As shown, the method includes: S501-S505.

[0236] S501. After the electronic device detects the recording command, it displays the first preview interface on the screen.

[0237] The first preview interface includes a first preview image.

[0238] It should be noted that the electronic device detecting a recording command can be an instruction that the electronic device detects it has entered recording mode. For example, the electronic device responds to a user's instruction... Figure 5 After the operation of Movie Mode 403 shown in Figure (2), the electronic device enters Recording Mode. At this time, the electronic device has not started recording video. Alternatively, the electronic device may detect the recording command after entering Recording Mode (such as Movie Mode), and the electronic device responds to the user's input to the recording control (such as... Figure 20 The recording control 414 shown in (1) is instructed. At this time, the electronic device starts recording video.

[0239] In some embodiments, when the recording command is to enter recording mode but video recording has not yet started, the first preview interface can be, for example, a... Figure 6 , Figure 9a and Figure 9b The interface shown. In other embodiments, when the recording command is to enter recording mode and start recording video, the first preview interface may be, for example, a... Figure 20 The interface shown in (2) is as follows: Figure 21 The interface shown and Figure 23 The interface shown.

[0240] S502, The electronic device determines the first target region of interest (ROI) on the first preview image based on the image captured by the front-facing camera.

[0241] The first target ROI is the area corresponding to the photographer's line of sight.

[0242] It should be noted that the image captured by the electronic device based on the front-facing camera can be, for example, the user image described in the above embodiments.

[0243] S503: The electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI and displays the second preview interface.

[0244] The second preview interface includes a second preview image; the second preview image includes a first clear area and a first blurred area, the first clear area corresponding to the first target ROI; the second preview image is generated by the electronic device processing the first image frame and the second image frame; the first image frame is captured by the rear first camera, and the second image frame is captured by the rear second camera.

[0245] It should be noted that the first clear area can be, for example, the first clear display area described in the above embodiments, and the first blurred area can be, for example, the first blurred display area described in the above embodiments.

[0246] In some embodiments, when the recording command is to enter recording mode but video recording has not yet started, the second preview interface can be, for example, a... Figure 10 The interface shown Figure 14 The interface shown Figures 16-19 The interface shown and Figure 20 The interface shown in (1). In other embodiments, when the recording command is to enter recording mode and start recording video, the second preview interface may be, for example, the interface shown in (1). Figure 20 Neutral (2) and Figure 23 The interface shown in (1) is shown in the middle.

[0247] It should be noted that the first image frame and the second image frame can be, for example, the original image frames described in the above embodiments. Specifically, the first image frame can be, for example, the first original image frame described in the above embodiments, and the second image frame can be, for example, the second original image frame described in the above embodiments.

[0248] S504. When the electronic device detects a change in the photographer's gaze, the electronic device determines a second target ROI on the second preview image based on the image captured by the front-facing camera.

[0249] The second target ROI is the area corresponding to the photographer's line of sight.

[0250] It should be noted that the image captured by the electronic device based on the front-facing camera can be, for example, the user image described in the above embodiments.

[0251] S505: The electronic device controls the rear first camera and the rear second camera to focus according to the second target ROI and displays the third preview interface.

[0252] The third preview interface includes a third preview image; the third preview image includes a second clear area and a second blurred area, the second clear area corresponding to the second target ROI; the third preview image is generated by the electronic device processing the third image frame and the fourth image frame; the third image frame is captured by the rear first camera, and the fourth image frame is captured by the rear second camera; the second target ROI is used to indicate the area on the display screen where the photographer's gaze falls and corresponds to the second preview image; the first clear area is different from the second clear area, and the first blurred area is different from the second blurred area.

[0253] In some embodiments, when the recording command is to enter recording mode but video recording has not yet started, the third preview interface can be, for example, a... Figure 11 , Figure 15 The interface shown. In other embodiments, when the recording command is to enter recording mode and start recording video, the third preview interface may be, for example, […]. Figure 21 , Figure 23 The interface shown in (2) is shown in the middle.

[0254] In some embodiments, the electronic device controls the rear first camera and the rear second camera to focus based on the second target ROI, determines the second focus area, and displays a third preview interface. The second focus area corresponds to the second clear area; that is, the focused object included in the second focus area is clearly displayed, while other objects are blurred. For example, combined with... Figure 11 , Figure 15 , Figure 21 , Figure 23 In the interface shown in (2), the second focus area is the area where the "tree" is located.

[0255] It should be noted that the third and fourth image frames can be, for example, the original image frames described in the above embodiments.

[0256] In some embodiments, after displaying a first preview interface on the display screen, the method further includes: the electronic device receiving a first event input by a user, and displaying a fourth preview interface on the display screen. The first event is used to trigger the electronic device to enter a large aperture mode.

[0257] The fourth preview interface includes a fourth preview image; the fourth preview image includes a third clear area and a third blurred area; the fourth preview image is generated by the electronic device processing the fifth image frame and the sixth image frame; the fifth image frame is captured by the rear first camera, and the sixth image frame is captured by the rear second camera; wherein, the third clear area is different from the first clear area, and the third blurred area is different from the first blurred area.

[0258] It should be noted that the fifth and sixth image frames can be, for example, the original image frames described in the above embodiments.

[0259] For example, in combination Figure 9a As shown in (1), the first event can be, for example, the user's operation on the large aperture control 408.

[0260] In some embodiments, when the recording command is to enter recording mode but video recording has not yet started, the fourth preview interface can be, for example, a... Figure 11 , Figure 15 The interface shown. In other embodiments, when the recording command is to enter recording mode and start recording video, the fourth preview interface may be, for example, [the interface shown]. Figure 21 , Figure 23 The interface shown in (2) is shown in the middle.

[0261] In some embodiments, when the electronic device enters large aperture mode, it can automatically focus on the fourth preview image displayed on the fourth preview interface using autofocus technology, so that the fourth preview image includes a third sharp area and a third blurred area. In other embodiments, when the electronic device enters large aperture mode, it can control the rear first camera and the rear second camera to focus based on the ROI (Region of Interest) being viewed by the photographer, so that the fourth preview image includes a third sharp area and a third blurred area.

[0262] In some embodiments, the electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI and displays a second preview interface, including: the electronic device controls the rear first camera and the rear second camera to focus according to the first target ROI, determines a first focus area in the second preview interface, and displays the second preview interface; wherein the first focus area corresponds to a first clear area.

[0263] It should be understood that the focused object within the first focus area is displayed clearly, while other objects are displayed in a blurred state. For example, combined with... Figure 10 , Figure 14 , Figures 16-19 , Figure 20 and Figure 23 As shown in (1), the first focus area is the area where the "bird" is located.

[0264] The process of the shooting method provided in the embodiments of this application will be described below with reference to the accompanying drawings.

[0265] For example, firstly, after the electronic device enters recording mode (such as the movie mode mentioned above), it displays on the screen as follows: Figure 9b The interface shown; where, in Figure 9bIn the interface shown, the "tree" and "bird" in the preview image displayed by the electronic device are clearly visible objects. Then, the electronic device receives the first user input, enters large aperture mode, and displays the following on the screen: Figure 11 The interface shown; among which, Figure 11 In the interface shown, the focus area in the preview image displayed by the electronic device is the area where the "tree" is located; that is, the "tree" included in the preview image is a clearly displayed object. Then, based on the image captured by the front-facing camera, the electronic device determines the first target ROI on the preview image and controls the first rear camera and the second rear camera to focus according to the first target ROI, displaying the result as shown. Figure 20 The interface shown in (1) is as follows; where, Figure 20 In the interface shown in (1), the focus area in the preview image displayed by the electronic device is the area where the "bird" is located, that is, the "bird" included in the preview image is a clearly displayed object.

[0266] Furthermore, electronic devices respond to user input. Figure 20 The operation of the recording control 414 shown in (1) is displayed as follows: Figure 20 The interface shown in (2) is as follows; among which, Figure 20 In the interface shown in (2), the focus area in the preview image displayed by the electronic device is the area where the "bird" is located, that is, the "bird" included in the preview image is a clearly displayed object. Then, when the electronic device detects a change in the photographer's gaze, it determines the second target ROI based on the image captured by the front-facing camera, and controls the first and second rear cameras to focus according to the second target ROI, displaying as shown in [image description missing]. Figure 21 The interface shown; among which, Figure 21 In the interface shown, the focus area in the preview image displayed by the electronic device is the area where the "tree" is located, that is, the "tree" included in the preview image is a clearly displayed object.

[0267] Finally, electronic devices respond to user input. Figure 21 The operation of the end button 417 shown will stop the video recording and generate a video file.

[0268] In some embodiments, the second preview interface includes a first prompt message; wherein the first prompt message is used to indicate the position of the first target ROI on the display screen; or, the first prompt message is used to guide the photographer to focus their gaze on the target position on the display screen.

[0269] For example, in combination Figures 14-17 As shown, the first prompt can be, for example, the "fill box" of the ROI that the eye is looking at.

[0270] In some embodiments, the second preview interface includes a mask area divided into multiple preset ROIs, each corresponding to a second preview image; a first prompt message is located within a target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI. The mask area is used to display a user image; or the mask area is used to display a scaled-down version of the second preview image.

[0271] For example, in combination Figures 14-16 As shown, the second preview interface includes a mask area, which is divided into multiple preset ROIs (e.g., divided into multiple small squares, each representing a preset ROI). The first prompt refers to the small squares within these smaller squares that are filled (e.g., "filled boxes"). It should be understood that "target ROI corresponds to target preset ROI" means that the area corresponding to the target ROI in the second preview image is the same as the area corresponding to the target preset ROI in the second preview image.

[0272] For example, combining Figure 14 and Figure 15 It can be seen that the mask area is used to display the image captured by the front-facing camera (i.e., to display the user's image); combined with Figure 16 It can be seen that the mask area is used to display the scaled-down second preview image.

[0273] In some embodiments, the second preview interface includes multiple preset ROIs, which correspond to the second preview image; the first prompt information is located within the target preset ROI among the multiple preset ROIs, and the target ROI corresponds to the target preset ROI.

[0274] For example, in combination Figure 17 As shown, the electronic device divides the area displaying the second preview image in the second preview interface into multiple preset ROIs (e.g., into multiple small boxes, each small box representing a preset ROI). The first prompt information refers to the small boxes filled within the multiple small boxes (e.g., "filled boxes").

[0275] In some embodiments, when the mask area is used to display images captured by the front-facing camera, the method further includes: when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device displays a preset face area in the mask area; wherein the preset face area is used to prompt the photographer to adjust the distance between themselves and the electronic device.

[0276] For example, in combination Figure 18As shown, the dashed box within the mask area represents a preset face region. For example, the photographer can adjust the distance between themselves and the electronic device using the preset face region. In some embodiments, the optimal distance between the photographer and the electronic device is indicated when the photographer's facial image falls exactly within the preset face region.

[0277] In some embodiments, the second preview interface further includes a second prompt message; the second prompt message is used to indicate to the photographer the position of the first focus area on the display screen.

[0278] For example, in combination Figure 20 and Figure 21 As shown, the second prompt information can be, for example, a focus indicator (or focus frame).

[0279] In some embodiments, when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device issues a text prompt message; the text prompt message is used to prompt the photographer to adjust the distance between themselves and the electronic device; or when the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device issues a voice prompt message; the voice prompt message is used to prompt the photographer to adjust the distance between themselves and the electronic device.

[0280] For example, in combination Figure 19 As shown, the electronic device can prompt the photographer to adjust the distance between themselves and the device through text prompts (or voice prompts).

[0281] In some embodiments, the electronic device identifies a focused object within a first focus area and tracks the focused object; when the electronic device detects that the duration for which the tracked focused object does not correspond to the first target ROI exceeds a preset duration, the electronic device redetermines the first focus area.

[0282] For example, in combination Figure 23 As shown in (1), the electronic device identifies a "tree" as the focused object within the first focus area, and tracks this "tree". When the electronic device detects that the tracked focused object (i.e., the "tree") does not correspond to the first target ROI for a duration longer than a preset duration (e.g., longer than a second preset duration), the electronic device re-determines the first focus area. For example, if the re-determined first focus area includes the focused object... Figure 23 The bird shown in (2) is shown in the middle.

[0283] In some embodiments, the electronic device identifies a focus object within a first focus area and tracks the focus object; when the electronic device detects that the tracked focus object does not correspond to the first target ROI and the photographer's gaze is not on the display screen, the electronic device maintains the first focus area.

[0284] In some embodiments, the second preview interface includes an end-recording control, and the method further includes: the electronic device generating a video file in response to the photographer's operation of the end-recording control; wherein the video file includes a first clear display area and a first blurred display area; the video file is generated by the electronic device processing a first image frame and a second image frame.

[0285] For example, such as Figure 20 As shown in (2), the second preview interface is the preview interface of the electronic device during the video recording process. The electronic device responds to the photographer's operation on the end recording control 417, generates a video file, and saves the video file in the electronic device, such as in the photo album application of the electronic device.

[0286] In some embodiments, the electronic device determines a first target region of interest (ROI) on a first preview image based on an image captured by a front-facing camera. This includes: the electronic device performing face recognition processing on the image captured by the front-facing camera to determine face information and eye information in the image; the face information includes the coordinates of the subject's facial contour, and the eye information includes one or more of the following: interpupillary distance, pupil size, pupil size variation, pupil brightness contrast, corneal radius, spot information, and iris information; the electronic device inputs the face information and eye information into a preset model and outputs the position where the subject's gaze falls on the display screen; the preset model is trained by the electronic device based on sample face information and sample eye information; and the electronic device determines the first target ROI on the first preview image based on the area corresponding to the subject's gaze in the first preview image.

[0287] For example, in conjunction with the above embodiments and Figure 7 It is understood that the electronic device inputs the image captured by the front-facing camera (such as the first image) into the face recognition module, which performs face recognition processing on the image captured by the front-facing camera to determine the face information and eye information of the image captured by the front-facing camera; then, the electronic device inputs the face information and eye information into the gaze detection module, which determines the first target ROI based on the face information and eye information.

[0288] It should be understood that the gaze detection module has a preset model. After the electronic device inputs the face information and eye information into the gaze detection module, the gaze detection module processes the face information and eye information according to the preset model to obtain the first target ROI.

[0289] In some embodiments, the electronic device controls the rear first camera and the rear second camera to focus based on the first target ROI to determine the first focus area, including: the electronic device performs autofocus (AF) processing on the first target ROI, controls the rear first camera and the rear second camera to focus, and determines the first focus area.

[0290] For example, in conjunction with the above embodiments and Figure 8 and Figure 12 It can be seen that the electronic device inputs the first target ROI into the AF algorithm module, and the electronic device controls the first rear camera and the second rear camera to focus through the AF algorithm module to determine the first focus area.

[0291] In some embodiments, the method further includes: the electronic device preprocessing the second image frame; the preprocessing being used to make the field of view of the second image frame and the first image frame the same; the electronic device calculating the depth of field based on the preprocessed second image frame and the first image frame; and the electronic device determining a first blurred display area based on the target ROI and the depth of field.

[0292] For example, in conjunction with the above embodiments and Figure 12 It can be seen that the electronic device can process the second image frame through the preprocessing module, and the second image frame after preprocessing has the same field of view as the first image frame. Then, the electronic device calculates the depth of field for the processed second and first image frames through the depth calculation module. Based on this, the electronic device inputs the calculated depth of field and the target ROI into the blurring algorithm module, which performs blurring processing to obtain the first blurred display area.

[0293] In some embodiments, before the electronic device displays the second preview interface, the method further includes: the electronic device performing image conversion processing on the first image frame and the second image frame; the image conversion processing includes: the electronic device converting the first image frame into a first image frame of a target format, and converting the second image frame into a second image frame of a target format; the bandwidth of the first image frame during transmission is higher than the bandwidth of the first image frame of the target format during transmission, and the bandwidth of the second image frame during transmission is higher than the bandwidth of the second image frame of the target format during transmission.

[0294] For example, in conjunction with the above embodiments and Figure 12It is understood that the electronic device can perform image conversion processing on the first image frame through the first ISP front-end module and on the second image frame through the second ISP front-end module. For example, both the first and second ISP front-end modules are pre-configured with a GTM algorithm module, which is used to perform image conversion processing on the first and second image frames. This image conversion processing can be, for example, "YUV domain" processing. Based on this, the first image frame obtained after image conversion processing can be, for example, a first image frame in "YUV domain" format, and the second image frame obtained after image conversion processing can be, for example, a second image frame in "YUV domain" format.

[0295] In some embodiments, the method further includes: an electronic device performing image simulation transformation processing on a first image frame of a target format; the image simulation transformation processing is used to enhance the first image frame of the target format.

[0296] For example, in conjunction with the above embodiments and Figure 12 It is understood that the electronic device can perform image simulation transformation processing on the first image frame of the target format through the first ISP backend module; correspondingly, the electronic device can perform image simulation transformation processing on the second image frame of the target format through the second ISP backend module. For example, the first ISP backend module and the second ISP backend module are pre-configured with a WRAP module; the WRAP module is used to enhance the target first original image frame and the target second original image frame.

[0297] In some embodiments, the electronic device, in response to a zoom operation input by the photographer, performs zoom processing on a first image frame of a target format to generate a first image frame of a target format corresponding to a target zoom magnification.

[0298] For example, in conjunction with the above embodiments and Figure 12 It can be seen that, in response to the zoom operation input by the photographer, the electronic device performs zoom processing on the first image frame of the target format through the first ISP backend module; correspondingly, the electronic device performs zoom processing on the first image frame of the target format through the second ISP backend module. For example, the first and second ISP backend modules are pre-configured with zoom modules, which are used to perform zoom processing on the first image frame and the second image frame of the target format.

[0299] This application provides an electronic device that may include a display screen, multiple cameras, a memory, and one or more processors. The display screen is used to display images captured by the multiple cameras or images generated by the processor. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device can perform the various functions or steps performed by the mobile phone in the above embodiments. The structure of this electronic device can be referred to... Figure 3 The structure of the electronic device 100 shown.

[0300] This application also provides a chip system, such as... Figure 25 As shown, the chip system 1800 includes at least one processor 1801 and at least one interface circuit 1802. The processor 1801 can be one of the types described in the above embodiments. Figure 3 The processor 110 is shown. The interface circuit 1802 can be, for example, an interface circuit between the processor 110 and external memory; or an interface circuit between the processor 110 and internal memory 121.

[0301] The processor 1801 and interface circuit 1802 described above can be interconnected via lines. For example, interface circuit 1802 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, interface circuit 1802 can be used to send signals to other devices (e.g., processor 1801). Exemplarily, interface circuit 1802 can read instructions stored in memory and send those instructions to processor 1801. When the instructions are executed by processor 1801, the electronic device can perform the various steps performed by the mobile phone in the above embodiments. Of course, this chip system may also include other discrete components, and this application embodiment does not specifically limit this.

[0302] This application also provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.

[0303] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.

[0304] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0305] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0306] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0307] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0308] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0309] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A shooting method, characterized in that, The method is applied in an electronic device, which includes a rear-facing first camera, a rear-facing second camera, a front-facing camera, and a display screen; the method includes: After detecting a recording command, the electronic device displays a first preview interface on the display screen; the first preview interface includes a first preview image. The electronic device determines a first target region of interest (ROI) on the first preview image based on the image captured by the front-facing camera; the first target ROI is the area corresponding to the photographer's line of sight. The electronic device performs autofocus (AF) processing on the first target ROI, controls the rear first camera and the rear second camera to focus, and determines the first focus area; the first focus area corresponds to the first target ROI; The electronic device acquires a first image frame captured by the rear first camera and a second image frame captured by the rear second camera; The electronic device preprocesses the second image frame; the preprocessing is used to make the field of view of the second image frame and the first image frame the same. The electronic device calculates the depth of field based on the preprocessed second image frame and the first image frame; The electronic device performs a blurring process on the first image frame based on the first target ROI and the depth of field, and displays a second preview interface; The second preview interface includes a second preview image, which includes a first sharp area and a first blurred area; the first sharp area corresponds to the first focus area, and the first blurred area is determined based on the result of the first target ROI and the depth of field. The second preview interface includes a mask area, which is divided into multiple preset ROIs, each corresponding to the second preview image. The second preview interface also includes a first prompt, located within a target preset ROI among the multiple preset ROIs, and the first target ROI corresponds to the target preset ROI. The mask area is used to display the image captured by the front-facing camera; alternatively, the mask area is used to display a scaled-down second preview image. The first prompt guides the photographer to adjust their gaze direction based on the target preset ROI and the multiple preset ROIs. When the mask area is used to display the image captured by the front-facing camera, the method further includes: When the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device displays a preset face area in the mask area; the preset face area is used to prompt the photographer to adjust the distance between himself and the electronic device, and to place the photographer's facial image in the preset face area.

2. The method according to claim 1, characterized in that, The method further includes: When the electronic device detects a change in the photographer's gaze, it determines a second target ROI on the second preview image based on the image captured by the front-facing camera; the second target ROI is the area corresponding to the photographer's gaze. The electronic device controls the rear first camera and the rear second camera to focus according to the second target ROI and displays a third preview interface; the third preview interface includes a third preview image; the third preview image includes a second clear area and a second blurred area, the second clear area corresponding to the second target ROI; the third preview image is generated by the electronic device processing a third image frame and a fourth image frame; the third image frame is captured by the rear first camera, and the fourth image frame is captured by the rear second camera; The first clear region is different from the second clear region, and the first blurred region is different from the second blurred region.

3. The method according to claim 1 or 2, characterized in that, After displaying the first preview interface on the display screen, the method further includes: The electronic device receives a first event input by the user and displays a fourth preview interface on the display screen; the first event is used to trigger the electronic device to enter the large aperture mode. The fourth preview interface includes a fourth preview image; the fourth preview image includes a third clear area and a third blurred area; the fourth preview image is generated by the electronic device processing the fifth image frame and the sixth image frame; the fifth image frame is captured by the rear first camera, and the sixth image frame is captured by the rear second camera; The third clear region is different from the first clear region, and the third blurred region is different from the first blurred region.

4. The method according to any one of claims 1-3, characterized in that, The second preview interface also includes a second prompt message; the second prompt message is used to indicate to the photographer the position of the first focus area on the display screen.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: When the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device also issues a text prompt message; the text prompt message is used to prompt the photographer to adjust the distance between themselves and the electronic device; or... When the electronic device detects that the distance between the photographer and the electronic device is not within a preset range, the electronic device also issues a voice prompt message; the voice prompt message is used to prompt the photographer to adjust the distance between himself and the electronic device.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: The electronic device identifies the focused object within the first focus area and tracks the focused object; When the electronic device detects that the tracked focus object does not correspond to the first target ROI for a duration longer than a preset duration, the electronic device re-determines the first focus area.

7. The method according to any one of claims 1-5, characterized in that, The method further includes: The electronic device identifies the focused object within the first focus area and tracks the focused object; When the electronic device detects that the tracked focus object does not correspond to the first target ROI and the photographer's gaze is not on the display screen, the electronic device maintains the first focus area.

8. The method according to any one of claims 1-7, characterized in that, The second preview interface includes an end-recording control; the method further includes: The electronic device generates a video file in response to the photographer's operation of the end recording control; wherein the video file includes the first clear area and the first blurred area; the video file is generated by the electronic device processing the first image frame and the second image frame.

9. The method according to any one of claims 1-8, characterized in that, The electronic device determines a first region of interest (ROI) on the first preview image based on the image captured by the front-facing camera, including: The electronic device performs face recognition processing on the image captured by the front-facing camera to determine the face information and eye information of the image captured by the front-facing camera; the face information includes the coordinates of the facial contour of the person being photographed, and the eye information includes one or more of the following: interpupillary distance, pupil size, pupil size variation, pupil brightness contrast, corneal radius, spot information, and iris information. The electronic device inputs the facial information and eye information into a preset model and outputs the position of the photographer's gaze on the display screen; the preset model is trained by the electronic device based on sample facial information and sample eye information. The electronic device determines the first target ROI on the first preview image based on the area corresponding to the photographer's line of sight in the first preview image.

10. The method according to any one of claims 1-9, characterized in that, Before the electronic device preprocesses the second image frame, the method further includes: The electronic device performs image conversion processing on the first image frame and the second image frame; The image conversion process includes: The electronic device converts the first image frame into a first image frame of the target format and converts the second image frame into a second image frame of the target format; the bandwidth of the first image frame during transmission is higher than the bandwidth of the first image frame of the target format during transmission, and the bandwidth of the second image frame during transmission is higher than the bandwidth of the second image frame of the target format during transmission.

11. The method according to claim 10, characterized in that, The method further includes: The electronic device performs image simulation transformation processing on the first image frame of the target format; the image simulation transformation processing is used to enhance the first image frame of the target format.

12. The method according to claim 10 or 11, characterized in that, The method further includes: The electronic device responds to the zoom operation input by the photographer, performs zoom processing on the first image frame of the target format, and generates a first image frame of the target format corresponding to the target zoom magnification.

13. An electronic device, characterized in that, include: A rear-facing first camera, a rear-facing second camera, a front-facing camera, a display screen, memory, and one or more processors; The display screen is used to display images captured by the rear first camera, the rear second camera, and the front camera; or, the display screen is used to display images generated by the processor. The memory stores computer program code, which includes computer instructions that, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, Includes computer instructions; when the computer instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Shooting method and shooting device of intelligent terminal

    CN106231185A

  • Video image processing method and related device

    CN111580671A

  • Image-capturing control apparatus and method for controlling same, and storage medium

    CN113452900A