A folding screen shooting method and electronic device
Patent Information
- Application Number
- CN202510334693.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2026-09-22
AI Technical Summary
[0002]日前,折叠屏电子设备因其更大的屏幕、更舒适的视觉体验受到用户越来越广泛的欢迎,通常折叠屏电子设备可以展开为两个屏幕,具备多个摄像头,包括前置摄像头和后置摄像头,使用折叠屏电子设备自拍时,可以使用前置摄像头或后置摄像头,但前置摄像头获取的自拍图像质量不高,例如清晰度较低,而使用后置摄像头自拍时获取的图像虽然质图像量较高,但用户的姿态不够自然、出现偏斜
Smart Images

Figure CN122802612A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a foldable screen shooting method and electronic device. Background Technology
[0002] Recently, foldable screen electronic devices have become increasingly popular among users due to their larger screens and more comfortable visual experience. Foldable screen electronic devices can typically unfold into two screens and have multiple cameras, including a front-facing camera and a rear-facing camera. When taking selfies with foldable screen electronic devices, users can use either the front-facing or rear-facing camera. However, the quality of selfies taken with the front-facing camera is not high, for example, the image is not very clear. While the image quality is higher when taking selfies with the rear-facing camera, the user's posture is not natural and appears skewed. Summary of the Invention
[0003] This invention provides a foldable screen shooting method and electronic device, enabling the acquisition of higher quality selfie images with more realistic user posture using the rear camera of the foldable screen electronic device.
[0004] In a first aspect, embodiments of the present invention provide a foldable screen shooting method applied to an electronic device, the electronic device including a foldable screen, the foldable screen including a first folding body and a second folding body, the first folding body including a first screen and a first camera disposed on different sides, the second folding body including a second screen and a second camera disposed on the same side, the method including: receiving a first operation from a user, wherein a shooting preview interface acquired by the first camera is displayed on the second screen; receiving a second operation from the user, wherein if the inward folding angle between the first folding body and the second folding body is within a preset range, a first image is acquired by the first camera and a second image is acquired by the second camera; the second operation is used to instruct the image displayed in the current shooting preview interface to be saved as a target shooting image; performing a first image processing on the first image and the second image respectively to obtain a third image and a fourth image, wherein the third image and the fourth image have the same scale and field of view; determining whether the third image and the fourth image contain a complete target face, and if so, performing a second image processing on the third image based on the fourth image to obtain a fifth image, wherein the pose corresponding to the target face in the fifth image is the same as the pose corresponding to the target face in the third image; and saving the target shooting image based on the fourth image and the fifth image.
[0005] By implementing the embodiments of this application, higher quality selfie images with more realistic user postures can be obtained when taking photos using foldable screen electronic devices. Specifically, when a user takes a photo using an electronic device, if the first camera (such as a rear camera) and the second camera (such as a front camera) are on the same side, the user can simultaneously activate the front and rear cameras to obtain a first image (such as a rear-camera image) and a second image (such as a front-camera image) with slightly different angles for the same scene. During the selfie process, the user tends to tilt towards the screen with the shooting preview interface and shooting controls (such as the second screen where the front camera is located). This results in the pose and gaze of the target face in the front-camera image being more realistic (usually facing forward or with a smaller degree of tilt relative to the front camera). While the rear-camera image is of higher quality (e.g., higher clarity), the pose and gaze of the target face in the rear-camera image are more significantly tilted compared to the front-camera image. In this embodiment, by simultaneously acquiring both images and referencing and combining them in subsequent image processing, a high-quality rear-camera selfie image with a more realistic posture can be obtained. In the process of mutual reference and combination mentioned above, the rear and front images can be adjusted to the same scale through the first image processing (such as scale alignment), which can ensure the accuracy of subsequent pose correction and gaze correction. Furthermore, by judging whether the scale-aligned rear image (third image) and the scale-aligned front image (fourth image) contain the complete target face, it can be ensured that the folding screen shooting method is executed when both the front and rear images contain the complete target face, which can avoid inaccurate processing results and waste of computing resources.
[0006] In one possible implementation of the first aspect, determining whether the third image and the fourth image contain a complete target face includes: performing facial feature point detection on the third image and the fourth image; and determining whether the third image and the fourth image contain a complete target face based on the result of the facial feature point detection. Implementing this embodiment of the application, determining the completeness of the target face through the result of facial feature point detection makes the target face integrity assessment simpler and faster. Furthermore, it avoids executing the method when the scale-aligned rear-shot image (third image) does not contain a complete target face, thus avoiding wasted computational resources. Simultaneously, judging the scale-aligned rear-shot image (third image) and the front-shot image (fourth image), rather than the scale-aligned rear-shot image (first image) and the front-shot image (second image), avoids executing the method when the target face is incomplete in the rear-shot image and the front-shot image due to scale alignment, thus preventing inaccurate processing results and wasted computational resources.
[0007] In one possible implementation of the first aspect, the step of performing first image processing on the first image and the second image respectively to obtain a third image and a fourth image includes: aligning the first image and the second image by field of view to obtain a first image and a second image after field of view alignment; and performing scale alignment on the first image and the second image after field of view alignment based on the size of the first image after field of view alignment to obtain the third image and the fourth image. By implementing the embodiments of this application, scale alignment can align the field of view of the rear-view image and the front-view image, and can also align the resolution of the front-view image with the rear-view image. Having the rear-view image and the front-view image at the same scale ensures the accuracy of subsequent pose correction and gaze correction.
[0008] In one possible implementation of the first aspect, aligning the field of view of the first image and the second image includes: obtaining first parameters of the first camera and second parameters of the second camera; generating a first cropping box corresponding to the first image and a second cropping box corresponding to the second image based on the first parameters and the second parameters; cropping the first image based on the first cropping box to obtain the first image after field of view alignment; and cropping the second image based on the second cropping box to obtain the second image after field of view alignment. By implementing embodiments of this application and generating cropping boxes of different sizes for front-facing and rear-facing images, it is possible to easily and quickly achieve uniformity of the field of view of the cropped front-facing and rear-facing images.
[0009] In one possible implementation of the first aspect, generating the first cropping box corresponding to the first image and the second cropping box corresponding to the second image includes: calculating the first field of view of the first camera based on the first parameter and calculating the second field of view of the second camera based on the second parameter; obtaining the current zoom ratio of the first camera and calculating the target field of view of the first camera based on the current zoom ratio; generating the first cropping box of the first image based on the target field of view and the first field of view, and generating the second cropping box of the second image based on the target field of view and the second field of view. By implementing the embodiments of this application, the magnification or reduction factor that the user wants to capture can be used as the current zoom ratio, and the field of view of the image that the user wants to capture, i.e., the target field of view, can be generated based on the current zoom ratio. By calculating and analyzing the target field of view with the first field of view (such as the field of view of the rear-facing image) and the second field of view (such as the field of view of the front-facing image), the first cropping frame for cropping the rear-facing image into the image corresponding to the field of view that the user wants to capture and the second cropping frame for cropping the front-facing image into the image corresponding to the field of view that the user wants to capture can be obtained quickly and conveniently.
[0010] In one possible implementation of the first aspect, the step of performing second image processing on the third image based on the fourth image to obtain a fifth image includes: calculating a first pose and a second pose corresponding to the target face in the third image and the fourth image, respectively; and performing pose correction on the first pose in the third image based on the second pose to obtain the fifth image. Implementing embodiments of this application, due to the positional relationship between the first camera (e.g., a rear camera) and the second camera (e.g., a front camera), the target face in the second image acquired by the second camera and the fourth image is usually facing forward or has a small degree of skew relative to the second camera. Correcting the pose of the target face in the third image based on the pose of the target face in the fourth image can obtain a fifth image where the pose and gaze of the target face are more accurately reproduced from the fourth image.
[0011] In one possible implementation of the first aspect, after performing a second image processing on the third image based on the fourth image to obtain a fifth image, the method further includes: determining whether the target face in the fifth image has missing facial content by comparing the third image and the fifth image; if so, performing completion processing on the fifth image to obtain a target captured image. Implementing this embodiment, due to the rotation of the target face, the edges of the target face in the captured image will shift. The greater the rotation, the more obvious the edge shift of the target face will be. Furthermore, a large-scale deflection of the target face will cause missing facial content. Therefore, by comparing the target face edges in the image before pose correction (the third image) with the target face edges in the image after pose correction (the fifth image), it can be determined whether the target face has undergone a large-scale rotation, thus easily and quickly inferring whether missing facial content exists in the target face.
[0012] In one possible implementation of the first aspect, the completion processing of the fifth image includes: generating a completed face based on the target face in the fourth image and the fifth image; replacing the target face in the fifth image with the completed face to obtain a target captured image with completed facial content. By implementing the embodiments of this application, when the pose-corrected fifth image (the pose-corrected rear-facing image) has missing facial content, further completion processing can be performed. This completion processing can be based on the fourth image (the pose-corrected front-facing image), where the target face is usually facing forward or has a relatively small degree of skew relative to the front-facing camera. This can largely ensure a more natural and harmonious facial completion effect, thereby improving the user experience.
[0013] In one possible implementation of the first aspect, the step of saving the target image based on the fourth image and the fifth image includes: performing a third image processing on the fifth image based on the fourth image to obtain the target image, wherein the gaze direction of the target image is the same as that of the target face in the fourth image; saving the target image to a gallery, and displaying a thumbnail of the target image on the shooting preview interface on the second screen. Implementing this embodiment not only corrects the pose of the target face in the rear-facing image, but also further corrects the gaze direction of the target face, making the pose and gaze of the target face in the final acquired target image more realistically reproduce the target face in the real scene, achieving a more natural imaging effect and improving the user experience.
[0014] In one possible implementation of the first aspect, the step of performing third image processing on the fifth image based on the fourth image to obtain the target captured image includes: calculating a gaze transformation matrix corresponding to the target face in the fifth image based on the fourth image and the fifth image; and performing gaze correction on the target face in the fifth image based on the gaze transformation matrix to obtain the target captured image. By implementing the embodiments of this application, performing matrix operations on the pixels in the fifth image (the pose-corrected rear-view image) based on the gaze transformation matrix can achieve gaze correction simply and quickly.
[0015] In one possible implementation of the first aspect, the first camera is a rear-facing camera, and the second camera is a front-facing camera; the first image is acquired through the first camera, and the second image is acquired through the second camera simultaneously. By implementing the embodiments of this application, simultaneously activating the first camera (e.g., a rear-facing camera) and the second camera (e.g., a front-facing camera) allows for the simultaneous acquisition of images from different angles of the same scene (or the image corresponding to the same target face). This largely ensures that the background areas in the first image (rear-facing image) and the second image (front-facing image) are the same or similar, avoiding additional processing of the background area, reducing computational load, and minimizing the impact of the background area on the target face in the front-facing image, thus providing more accurate reference information to assist in processing the rear-facing image.
[0016] In a second aspect, embodiments of the present invention provide an electronic device, the electronic device including a memory and one or more processors; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method described in any of the above aspects.
[0017] Thirdly, embodiments of the present invention provide a chip system applied to an electronic device, characterized in that the chip system includes one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to execute the method described in any of the above aspects.
[0018] Fourthly, embodiments of the present invention provide a computer storage medium, characterized in that the computer storage medium stores a computer program, which, when executed by a processor, implements the method described in any of the above aspects.
[0019] Fifthly, embodiments of the present invention provide a computer program product, characterized in that the computer program product includes instructions that, when executed by a computer, enable the computer to perform the method described in any of the preceding aspects. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.
[0021] Figure 1A This is a schematic diagram of an electronic device at different angles provided in the embodiments of this application.
[0022] Figure 1B This is a schematic diagram of another electronic device at a different angle provided in the embodiments of this application.
[0023] Figure 1C This is a schematic diagram of another electronic device at a different angle provided in the embodiments of this application.
[0024] Figure 1D This is a schematic diagram of another electronic device at a different angle provided in the embodiments of this application.
[0025] Figure 2A This is a schematic diagram of a first camera selfie provided in an embodiment of this application.
[0026] Figure 2B This is a schematic diagram of the second screen when the electronic device provided in this application takes a picture.
[0027] Figure 3 This is a schematic diagram of a foldable screen shooting method provided in an embodiment of this application.
[0028] Figure 4A This is a schematic diagram of the main interface of an electronic device provided in an embodiment of this application.
[0029] Figure 4B This is an electronic device shooting preview interface provided in the embodiments of this application.
[0030] Figure 5 This is a schematic diagram of the target face rotation state provided in an embodiment of this application.
[0031] Figure 6 This is a schematic diagram of another foldable screen shooting method provided in the embodiments of this application.
[0032] Figure 7 This is a schematic diagram of a field-of-view alignment process provided in an embodiment of this application.
[0033] Figure 8 This is another field-of-view alignment flowchart provided in the embodiments of this application.
[0034] Figure 9 This is a schematic diagram of a preview pipeline and a shooting pipeline provided in an embodiment of this application.
[0035] Figure 10A This is a schematic diagram of a first cutting frame provided in an embodiment of this application.
[0036] Figure 10B This is a schematic diagram of a second cutting frame provided in an embodiment of this application.
[0037] Figure 11 This is a schematic diagram of a face integrity assessment process provided in an embodiment of this application.
[0038] Figure 12A This is a schematic diagram of a standard facial feature point provided in an embodiment of this application.
[0039] Figure 12B This is a schematic diagram of the target facial feature point detection result provided in an embodiment of this application.
[0040] Figure 12C This is a schematic diagram of another target facial feature point detection result provided in an embodiment of this application.
[0041] Figure 12D This is a schematic diagram of another target facial feature point detection result provided in an embodiment of this application.
[0042] Figure 13 This is a schematic diagram of a posture correction process provided in an embodiment of this application.
[0043] Figure 14A This is a schematic diagram of a camera coordinate system provided in an embodiment of this application.
[0044] Figure 14B This is a top view of a camera coordinate system provided in an embodiment of this application.
[0045] Figure 15A This is a third-image schematic diagram provided in an embodiment of this application.
[0046] Figure 15B This is a fourth schematic diagram provided in an embodiment of this application.
[0047] Figure 15C This is a schematic diagram of a target face mask provided in an embodiment of this application.
[0048] Figure 15D This is a fifth schematic diagram provided in an embodiment of this application.
[0049] Figure 15E This is a schematic diagram of an eye mask for a target face in a fifth image provided in an embodiment of this application.
[0050] Figure 15F This is a schematic diagram of a target image captured according to an embodiment of this application.
[0051] Figure 16A This is another third-image schematic diagram provided in the embodiments of this application.
[0052] Figure 16B This is another fourth schematic diagram provided in the embodiments of this application.
[0053] Figure 16C This is another fifth schematic diagram provided in the embodiments of this application.
[0054] Figure 16D This is a schematic diagram of an edge difference provided in an embodiment of this application.
[0055] Figure 16E This is a schematic diagram of another target image provided in an embodiment of this application.
[0056] Figure 17A This is a schematic diagram of a vision correction process provided in an embodiment of this application.
[0057] Figure 17B This is a schematic diagram of another vision correction process provided in the embodiments of this application.
[0058] Figure 17C This is an eye image of a target human face provided in an embodiment of this application.
[0059] Figure 17D This is a schematic diagram of the eye feature points of a target human face provided in an embodiment of this application.
[0060] Figure 17E This is a schematic diagram of the gaze corresponding to a target face provided in an embodiment of this application.
[0061] Figure 18A This is a schematic diagram of a fifth image completion process provided in an embodiment of this application.
[0062] Figure 18B This is another schematic diagram of the fifth image completion process provided in the embodiments of this application.
[0063] Figure 19 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0064] Figure 20 This is a schematic diagram of the system architecture of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0065] This application provides a foldable screen shooting method and an electronic device.
[0066] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “some,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations that include one or more of the listed items.
[0067] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0068] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0069] The electronic device to which this application's embodiments apply may include multiple folding bodies and multiple cameras, wherein the multiple cameras include at least one front-facing camera and one rear-facing camera, and at least one front-facing camera and at least one rear-facing camera are capable of taking pictures on the same plane (or the same side of the foldable screen electronic device), and the front-facing camera and the rear-facing camera may be on different folding bodies. Below, taking a foldable screen phone as an example, a possible electronic device form is described. Combined with... Figures 1A-1C This describes an electronic device that uses the foldable screen shooting method provided in the embodiments of this application.
[0070] Specifically, the electronic device may include a first folding body 1101 and a second folding body 1102, wherein the first folding body 1101 may include a first camera 1081, a third camera 1083 and a first screen 1091, and the second folding body 1102 may include a power switch 105, a second screen 1092 and a third screen 1093.
[0071] Figure 1A This is a schematic diagram of an electronic device at different angles provided in the embodiments of this application, such as... Figure 1A As shown, in the electronic device, the first folding body 1101 and the second folding body 1102 are connected by a linking component ( Figure 1A (not shown in the image) are interconnected and can be connected along virtual fold lines ( Figure 1A (As shown by the dashed line) Folds the first folding body 1101 and the second folding body 1102 inwards by an angle 'a', where 'a' is any angle from 0 to 180 degrees. Figure 1A The right side of the virtual fold line shown (here, the right side refers to...) Figure 1A The right side of the image shows the first folding body 1101, and the back cover of the first folding body 1101 ( Figure 1A The black area shown is the first camera 1081 in the upper left corner. The first camera 1081 is the rear camera. Figure 1A The left side of the virtual folding line shown is the second folding body 1102. The left side of the second folding body 1102 includes a power switch 105. Additionally, as shown... Figure 1A As shown, one side of the second folding body 1102 is a third screen 1093, and the third screen 1093 does not include a camera;
[0072] Figure 1B This is a schematic diagram of another electronic device at a different angle provided in an embodiment of this application. Figure 1B It is Figure 1A The electronic device shown is oriented towards Figure 1A The fold is obtained by flipping in the direction of the middle arrow; at this point, 'a' is 180 degrees, and the electronic device is fully unfolded, as shown below. Figure 1BAs shown, the third screen 1093 in the second folding body 1102 is located to the right of the power switch 105. The first screen 1091 is located to the right of the second screen 1092. The first screen 1091 and the third screen 1093 are separated by a virtual folding line. This virtual folding line is a visual representation to better illustrate the structure of the electronic device; in actual electronic devices, there is no corresponding, physically defined virtual folding line. Above the first screen 1091 is a third camera 1083, which is a front-facing camera. In the unfolded electronic device, the first screen 1091 and the third screen 1093 can together form a larger display screen. A user interface 50 can be displayed on this display screen, and a time display control can be displayed in the center of the user interface 50. Figure 1B The time display control shown, "09:08 Saturday, February 8, 2025", can display the current time.
[0073] Figure 1C This is a schematic diagram of another electronic device at a different angle provided in the embodiments of this application. Figure 1C yes Figure 1B The reverse side of the electronic device shown, wherein Figure 1C The left side of the virtual fold line shown is the back cover of the first folding body 1101. Figure 1C (The black part in the middle) The first camera 1081 is located on the upper left of the back cover. The first camera 1081 is a rear camera. The first screen 1091 is located to the right of the virtual folding line. The second camera 1082 is located above the first screen 1091. The user interface 51 can be displayed on the first screen 1091. The user interface 51 can display time display controls and a status bar. The status bar can include, but is not limited to, battery display controls, time display controls and signal display controls.
[0074] Figure 1D This is a schematic diagram of another electronic device at a different angle provided in the embodiments of this application. Figure 1D yes Figure 1A The electronic device shown is in its folded state. Figure 1A Angle 'a' is shown as 0 degrees. The second screen 1092 and the second camera 1082 above the second screen 1092 are... Figure 1D On the front, the second screen 1092 can display the user interface 52. Figure 1D The electronic device is shown at an angle where the power switch 105 is on the right side of the first screen 1091, and the back of the electronic device is... Figure 1C The back cover shown ( Figure 1A or Figure 1C (The black portion shown) and the first camera 1081, Figure 1D The electronic device at the angle shown cannot be fully displayed.
[0075] It should be noted that the above Figures 1A-1D The electronic devices shown are merely illustrative examples of embodiments of this application and do not constitute a specific limitation on the electronic devices.
[0076] The following is combined with Figure 1B , Figure 1C as well as Figure 2A This application describes the application scenarios and the technical problems that can be solved by the embodiments of this application.
[0077] In electronic devices, front-facing cameras typically have lower specifications than rear-facing cameras. For example, rear-facing cameras usually feature higher pixel counts and more advanced sensors. In some implementations, rear-facing cameras can reach 48MP (megapixel) or higher, and their sensors are usually higher-performance models. In contrast, front-facing cameras typically have around 12MP, and their sensors are usually more common. Rear-facing cameras also generally have better optical performance, such as a larger aperture, higher sensitivity, and autofocus. Front-facing cameras, limited by space and cost, usually have poorer optical performance, such as a smaller aperture, lower sensitivity, and lack of autofocus. Rear cameras, including those with aperture, control the amount of light passing through the lens and reaching the sensor inside the camera body. They are typically located inside the lens and adjust the amount of light entering by changing the size of the aperture opening. The aperture size affects the exposure of the captured image; a larger aperture allows more light in, resulting in a brighter image and better performance in low-light environments, producing clear and bright photos. Furthermore, rear cameras usually feature longer focal length lenses, reducing perspective distortion when shooting portraits and making facial contours and proportions appear more natural. Front cameras, on the other hand, typically have shorter lenses, leading to enhanced perspective distortion and making objects appear more distorted in close-up shots. In conclusion, rear cameras in electronic devices are suitable for a wide range of shooting scenarios and support various shooting modes and functions, such as night mode and ultra-wide-angle shooting. They generally perform better than front cameras when shooting portraits or landscapes.
[0078] When taking a selfie on a foldable electronic device that includes multiple screens and multiple cameras, with at least one front-facing camera and one rear-facing camera, and at least one front-facing camera and at least one rear-facing camera capable of shooting on the same plane (or the same side of the foldable electronic device), the rear-facing camera can be used to capture the image, resulting in a higher quality image (e.g., a clearer image), but the pose and gaze of the person in the image are usually slightly skewed (e.g., the face is turned or the gaze is slanted). Alternatively, the front-facing camera can be used, resulting in a lower quality image, but the pose and gaze of the person in the image are usually straight or slightly skewed. For example, one can use... Figure 1B (or Figure 1C The first camera 1081 on the first folding body 1101 shown in the figure is in accordance with Figure 2A Take a selfie as shown. Figure 2A This is a schematic diagram of a first camera selfie provided in an embodiment of this application. Figure 2A As shown, when taking a selfie using the first camera 1081, a shooting preview interface will be displayed on the second screen 1092. This shooting preview interface is the user interface 53. The shooting preview interface may include the image captured by the first camera 1081, an operation menu bar, a magnification adjustment control, and more operation controls. The operation menu bar may include a "Shoot" control, a "Flip" control, and a "View Image" control. The user can click the "Shoot" control to capture the image captured by the first camera 1081 on the shooting preview interface, and click the "Flip" control to shoot the screen displayed in the preview interface. Furthermore, the user can click different magnification adjustment controls (e.g., ...). Figure 2A The circular controls shown (indicated by "1x", "2x", and "5x") are used to adjust the zoom level during shooting. Users can also click "More Operation Controls" in the upper right corner of the second screen 1092 for other shooting adjustments, resulting in clearer and more natural selfie images. However, during shooting, because the shooting preview interface and the first camera are not on the same folding body, the shooting preview interface is displayed on the second screen of the second folding body 1102, while the first camera 1081 capturing the image is located in the upper left corner of the first folding body 1101 (as shown by...). Figure 2A (as shown in the direction of the electronic device), therefore, when the user previews the image captured by the first camera 1081 and clicks the "shoot" control, as... Figure 2A As shown, the angle, pose, and line of sight of the face relative to the first camera 1081 will be tilted, resulting in low image quality for selfies.
[0079] Therefore, how to obtain higher quality selfie images in foldable screen electronic devices that more realistically reproduce the posture and gaze of the person in the scene is an urgent problem to be solved.
[0080] The following is combined with Figure 3 This application describes the main process of the foldable screen shooting method provided in the embodiments. Figure 3 This is a schematic diagram of a foldable screen shooting method provided in an embodiment of this application.
[0081] like Figure 3 As shown, the foldable screen shooting method mainly includes the following steps:
[0082] S201: Receive the user's first operation and display the shooting preview interface acquired by the first camera on the second screen.
[0083] Specifically, a user can take a selfie using an electronic device. The electronic device receives a first operation from the user and responds to the first operation by displaying a shooting preview interface captured by the first camera on a second screen. The first camera can be a rear camera. The user can instruct the electronic device to display the shooting preview image captured by the first camera (i.e., the rear camera) on the second screen through the first operation. For example, the first operation can be clicking "camera" or other operations that can turn on the first camera and display the preview image captured by the first camera on the preview interface on the second screen.
[0084] In some possible implementations, the user can unfold the electronic device to Figure 4A The shape shown Figure 4A This is a schematic diagram of the main interface of an electronic device provided in an embodiment of this application, such as... Figure 4A As shown, the first screen 1091 and the third screen 1093 of the electronic device together constitute a larger display screen. This larger screen can display the user interface 55 of the electronic device. The user interface 55 may include a time display control, a status bar, and an icon for at least one application. The time display control may include the current time and date; the status bar may include the network status, signal status, battery level, and current time of the electronic device; the main interface may also include an icon for at least one application, such as... Figure 4A The applications shown include: "Music", "Videos", "App Store", "Browser", "Dialer", "Messages", "Contacts", "Camera", "Weather", "Stocks", "Calculator", "Settings", "Gallery", etc. The positions of the application icons and corresponding application names can be adjusted according to the user's preferences, and this application embodiment does not limit this.
[0085] like Figure 4A As shown, users can click the "Camera" application on this interface, and the electronic device will display on a larger screen formed by the first screen 1091 and the third screen 1093. Figure 4B The interface shown. Figure 4B This application provides an electronic device shooting preview interface, such as... Figure 4B As shown, the shooting preview interface (i.e., user interface 56) displayed on the screen jointly formed by the first screen 1091 and the third screen 1093 of the electronic device can include the image captured by the first camera 1081, an operation menu bar, and a parameter menu bar. The operation menu bar can include a "flip" control, a "flip camera" control, a "shoot" control, and a "view image" control. At this time, the user is usually facing the third screen 1093 and the first screen 1091. The user can use the third camera 1083 (i.e., the front-facing camera) to take a selfie. When the user needs to use the first camera 1081 (i.e., the rear-facing camera) to obtain a higher-resolution, clearer, and more natural selfie, such as... Figure 4B As shown, the "Flip" control in the operation menu bar can be clicked. Clicking this control instructs the electronic device to turn off the third screen 1093 and the first screen 1091. Figure 4B The second screen 1092 on the reverse side of the electronic device displays the shooting preview interface and receives further shooting operations from the user. At this time, the user's "clicking the flip control" operation is the first operation.
[0086] Furthermore, in some possible implementations, after the electronic device receives the user's first operation, Figure 4B The screen formed by the first screen 1091 and the third screen 1093 will turn off and go black (without displaying content), while the shooting preview interface will be displayed on the second screen 1092. At this time, the user can manually adjust the way they hold the electronic device. Figure 4B The electronic device shown flips over, causing Figure 2A The side shown, which includes the first camera 1081 and the first screen 1091, faces the user.
[0087] S202: Receive the user's second operation, acquire a first image through the first camera, and acquire a second image through the second camera.
[0088] Specifically, when the inward folding angle between the first folding body and the second folding body of the electronic device is within a preset range, such as when the inward folding angle is between 170 and 180 degrees, the user can preview the shooting preview interface displayed on the second screen and captured by the first camera (i.e., the rear camera), determine the image to be captured, and input a second operation into the electronic device. The second operation is used to instruct the electronic device to save the image displayed on the current shooting preview interface as the target shooting image. For example, the second operation can be clicking the shooting control or other operations that enable the electronic device to save the image displayed on the shooting preview interface into the electronic device. The second camera is a front camera.
[0089] The electronic device receives and responds to the user's second operation, acquiring a rear-facing image (i.e., a first image) through the first camera (i.e., the rear camera) and a second image through the second camera (i.e., the front camera). Simultaneously, the first camera 1081 acquires the first image while the second camera 1082 acquires the front-facing image (i.e., the second image), and the rear camera acquires the first image while the front camera acquires the second image. The rear-facing and front-facing images depict the same scene, but due to the different positions of the front and rear cameras within the electronic device, their content may differ slightly. For example, the user's facial angle, posture, and line of sight may differ in the first and second images.
[0090] For example, the user's second action can be Figure 2A As shown, clicking the "Shoot" control in the operation status bar on the second screen 1092 allows for other operations, including clicking, before the user inputs a second operation and while previewing the shooting preview interface displayed on the second screen 1092. Figure 2A The magnification adjustment controls in the shooting preview interface are used to adjust the magnification or reduction factor when the first camera 1081 is shooting. The user adjusts the magnification of the first camera 1081 using these controls; the magnification of the second camera 1082 will not be adjusted (it will maintain its base magnification). For example, if the user clicks... Figure 2A The zoom adjustment control shown has a circular control labeled "2x". When the circular control labeled "2x" is selected, it can turn yellow or become bold. At this time, the magnification of the first camera 1081 is 2x. The image captured by the first camera 1081 displayed on the second screen 1092 will also be changed accordingly. The zoom ratio is the current zoom level when the user inputs the second operation and the zoom ratio of the first camera is selected by the user through the zoom adjustment control.
[0091] S203: Perform first image processing on the first image and the second image respectively to obtain the third image and the fourth image.
[0092] Specifically, after receiving the user's second operation, the electronic device first performs a first image processing on the rear-facing image (i.e., the first image) and the front-facing image (i.e., the second image) to obtain a third image and a fourth image, respectively. The third image is obtained by processing the first image, and the fourth image is obtained by processing the second image. The scale and field of view of the third image (the rear-facing image after first image processing) and the fourth image (the front-facing image after first image processing) are the same. In one possible implementation, the first image processing can be image scale alignment, used to adjust the field of view and scale of the first and second images to a unified standard to facilitate subsequent image processing of both.
[0093] S204: Determine whether the third and fourth images contain the complete target face.
[0094] Specifically, before proceeding to the next image processing step, the electronic device can first determine whether the third and fourth images contain a complete target face. If they do, step S205 is executed; if they do not, step S207 is executed. The target face refers to a face that appears in the rear-facing image (i.e., the first image) and the front-facing image (i.e., the second image) and can be processed using the folding screen shooting method provided in this application embodiment.
[0095] Since the main purpose of the foldable screen shooting method provided in this application embodiment is to enable the use of the rear camera to obtain images of the target face with higher clarity and more realistic facial angle, pose, and gaze when taking selfies using a foldable screen electronic device, it is necessary to ensure that the image processed includes at least the target face. In addition, in some possible implementations, the first image processing can be scale alignment, which can include image cropping. Therefore, the rear-facing image (i.e., the first image) and the front-facing image (i.e., the second image) before scale alignment may contain the complete target face. However, the target face in the rear-facing image (i.e., the third image) and the front-facing image (i.e., the fourth image) after scale alignment may no longer be complete due to image cropping. It is impossible to perform subsequent image processing based on such third and fourth images. Therefore, it is necessary to determine whether the target face is contained in the third and fourth images after scale alignment, rather than determining whether the target face is contained in the first and second images.
[0096] S205: Perform second image processing on the third image based on the fourth image to obtain the fifth image.
[0097] Specifically, when both the third and fourth images contain complete human faces, the third image can be processed based on the fourth image to obtain the fifth image, where the pose of the target face in the fifth image is the same as the pose of the target face in the third image.
[0098] Pose, in this context, refers to a comprehensive description of the position and orientation of a target face in an image. The position of the target face can be its coordinates in the two-dimensional coordinate system of the image; the orientation can be the rotational state of the face within the image. For example, the orientation of the target face can refer to its rotational state. Figure 5 This is a schematic diagram of the target face rotation state provided in an embodiment of this application, such as... Figure 5As shown, when a user takes a selfie using an electronic device, the target face is directly facing either the first camera 1081 or the second camera 1082. (Due to the angle,) Figure 5 The screen cannot directly display the first camera 1081 and the second camera 1082. The screen jointly formed by the first screen 1091 and the third screen 1093 is in a screen-off state. The coordinate system corresponding to the target face pose is a right-handed coordinate system. The rotation state of the target face can include:
[0099] 1. Yaw Angle: The yaw angle describes the angle of rotation of the target face around the y-axis. Here, the y-axis can be a straight line perpendicularly upward through the center of the head. When the target face rotates left or right around the y-axis, the yaw angle changes accordingly. In some possible implementations, the yaw angle can range from 0 to 360 degrees, representing a full rotation of the target face. Alternatively, the yaw angle can range from -180 to +180 degrees. When the target face is directly facing the camera (or electronic device), the yaw angle is 0 degrees. In the image (two-dimensional image) acquired by the first camera 1081 or the second camera 1082, the yaw angle of the target face is manifested as a horizontal yaw of the head to the left or right. The edges of the target face and the facial features will differ depending on the yaw angle. For example, when the target face rotates around the y-axis... Figure 5 When the target face rotates in the direction of the arrow on the y-axis, it can be assumed that the target face rotates to the left in the image acquired by the first camera 1081 or the second camera 1082. In this case, the right edge of the target face may be incomplete in the image acquired by the first camera 1081 or the second camera 1082, and the edge of the target face will be shifted to the right. In addition, when the target face rotates to the left, its left face will be displayed more completely and occupy a larger area, while the right face may not be displayed completely due to the angle rotation, resulting in missing facial content due to facial rotation.
[0100] 2. Pitch Angle: The pitch angle describes the angle of rotation of a face around the x-axis. The x-axis can be a straight line passing through the center of the head. When the face rotates up or down around the x-axis, the pitch angle can vary from 0 to 360 degrees. Similarly, in some possible implementations, the pitch angle can be in the range of 0-360 degrees, representing a full rotation of the target face. In addition, the pitch angle can also be in the range of -180 to +180 degrees. When the face is facing the camera (or electronic device), the pitch angle yaw is 0 degrees. In the image acquired by the first camera 1081 or the second camera 1082, different pitch angles of the target face are represented by looking up or looking down.
[0101] 3. Roll Angle: The roll angle describes the angle of rotation of a face around the z-axis (usually pointing directly in front of the face). The z-axis can be a straight line passing through the center of the head and extending towards the back of the head. When the face tilts to the left or right around the z-axis, the roll angle can vary from 0 to 360 degrees. Similarly, in some possible implementations, the roll angle can be in the range of 0-360 degrees, indicating that the target face tilts and rotates one full turn. In addition, the roll angle can also be in the range of -180 to +180 degrees. When the face is facing the camera (or electronic device) and the head is not tilted, the roll angle is 0 degrees. In the images acquired by the first camera 1081 or the second camera 1082, different roll angles of the target face are manifested as the head tilting to the left or right.
[0102] For example, the second image processing can be pose correction, that is, based on the pose of the target face in the fourth image, the pose of the target face in the third image is corrected to obtain the fifth image. In the embodiments of this application, an electronic device is used according to... Figure 5 When taking a selfie, since the shooting preview interface is displayed on the second screen 1092 of the second folding body 1102, and the second camera 1082 is above the second screen 1092, the user is usually facing the second camera 1082 when previewing the shooting preview interface and inputting the second operation. Therefore, the pose of the target face in the second image obtained by the second camera 1082 and the second image after processing the first image (i.e., the fourth image) can be used as the correct pose. Therefore, the pose of the target face in the third image can be corrected based on the pose of the target face in the fourth image.
[0103] S206: Based on the fourth and fifth images, save the target image.
[0104] Specifically, in response to the user's second operation, the electronic device needs to save the target image to the electronic device. For example, the pose of the target face in the fifth image (the rear-facing image after pose correction) is aligned with the fourth image (the front-facing image after scale alignment). When the target face is complete in the rear-facing image after pose correction, the image (the fifth image) can be directly identified as the target image and saved to an application such as a gallery where the image can be viewed. A thumbnail of the target image is also displayed on the shooting preview interface on the second screen for the user to click and view.
[0105] S207: Save the third image as the target image.
[0106] Specifically, when neither the third image (i.e., the scale-aligned rear-facing image) nor the fourth image (i.e., the scale-aligned front-facing image) simultaneously contains the complete target face, it is impossible to perform second image processing and third image processing on the third and fourth images. In this embodiment, the third image can be directly saved as the target image.
[0107] In some other possible implementations, when neither the third image nor the fourth image simultaneously contains the complete target face, the process corresponding to the foldable screen shooting method provided in this application embodiment can be exited, and the shooting process can be executed normally. For example, in response to the user's second operation, the normal shooting process can be executed (such as shooting the pipeline; a detailed description of the pipeline can be found later). Figure 9 (Related description), save the image on the second screen's shooting preview interface as the target shooting image to the electronic device.
[0108] The following is combined with Figure 6 This application describes in detail the specific process of obtaining a target image based on a first image and a second image and saving it to an electronic device in the foldable screen shooting method provided in the embodiments of this application.
[0109] Figure 6 This is a schematic flowchart of another foldable screen shooting method provided in an embodiment of this application. Figure 6 As shown, after receiving the user's first operation and then the user's second operation, the electronic device responds to the user's second operation by performing a series of processes on the first and second images to obtain the target image and save the target image to the electronic device. The specific image processing flow may include... Figure 6 The steps S301-S313 shown are as follows, wherein, Figure 6 The steps S301-S302 shown can correspond to Figure 3 Step S203: Perform first image processing on the first image and the second image respectively to obtain the third image and the fourth image; Figure 6 Steps S304 and S306 shown can correspond to Figure 3 Step S205: Perform second image processing on the third image based on the fourth image to obtain the fifth image; In some possible implementations, the folding screen shooting method provided in this application embodiment may also include steps S307 and S309; Figure 6 Steps S308, S310, and S311 shown can correspond to Figure 3 Step S206: Based on the fourth and fifth images, save the target image. The following is a detailed explanation with reference to the accompanying drawings. Figure 6 The process shown is as follows:
[0110] S301: Align the first image and the second image with the field of view to obtain the first image and the second image after the field of view alignment.
[0111] Specifically, the first image is acquired by the first camera, and the second image is acquired by the second camera. Due to the significant differences in specifications between the first and second cameras, the first image acquired by the first camera and the second image acquired by the second camera have significant differences in field of view and scale. If the first and second images are not scaled together, and subsequent pose correction is performed directly based on the first and second images, the pose correction results are likely to be inaccurate, or even cause image distortion. Therefore, in order to ensure the accuracy of subsequent pose correction, the first and second images need to be scaled together, where scale includes the size of the two images and the field of view (FOV).
[0112] First, the two images need to be aligned in terms of field of view to obtain the first and second images after alignment. Field of view (FOV) refers to the angle of the field of view that a camera can capture. Specifically, the field of view is the angle formed by the two edges of the maximum range that the user can see through the lens, with the camera lens as the vertex. The field of view (FOV) determines the width of the image that the camera can capture.
[0113] In one possible implementation, Figure 7 This is a schematic diagram of a field-of-view alignment process provided in an embodiment of this application, such as... Figure 7 As shown, the field-of-view alignment of the first image and the second image may include steps S401-S405. To more clearly illustrate the steps involved in the field-of-view alignment process, the following describes the steps in conjunction with... Figure 7 and Figure 8 Please provide a detailed explanation. Figure 8 This is another field-of-view alignment flowchart provided in the embodiments of this application.
[0114] S401: Obtain the first parameter of the first camera and the second parameter of the second camera.
[0115] Specifically, the first parameter of the first camera refers to a series of parameters, characteristics, and calibration data about the first camera that have been determined or are known during the design and manufacturing process of the camera module of the electronic device. For example, the first parameter of the first camera may include: 1. Sensor parameters: the type and size of the sensor in the first camera, the number and size of pixels; 2. Optical parameters: the focal length, aperture size, etc. of the first camera. In addition, the first parameter may also include some other parameters, which are not limited in this embodiment. Similarly, the specific content of the second parameter of the second camera can refer to the relevant content of the first parameter mentioned above, and will not be repeated here.
[0116] S402: Calculate the first field of view based on the first parameter and the second field of view based on the second parameter.
[0117] Specifically, such as Figure 8 As shown, the first field of view of the first camera can be calculated based on the first parameter, and the second field of view of the second camera can be calculated based on the second parameter. Both the first and second field of view are field of view angles at the base magnification of the corresponding cameras. The relevant description of the field of view angle can be found in the relevant description in step 301 above. The size of the field of view angle is affected by the camera's shooting magnification (i.e., the current zoom ratio). When the shooting magnification is higher (magnified shooting), the field of view angle is smaller, and when the shooting magnification is lower (reduced shooting), the field of view angle is larger.
[0118] For example, the field of view can be calculated based on the size of the camera sensor and the focal length of the camera. The field of view of the camera can include the horizontal field of view (HFOV) and the vertical field of view (VFOV). The horizontal FOV describes the range of the camera's field of view in the horizontal direction, reflecting the width of the scene that the camera can capture in the horizontal direction; the vertical FOV describes the range of the field of view in the vertical direction, reflecting the height of the scene that the camera can capture in the vertical direction. Both can be measured in degrees or radians. For example, for a first camera, its first field of view can include a first horizontal field of view and a first vertical field of view. The first horizontal field of view can be calculated as follows:
[0119]
[0120] Where HFOV is the first horizontal field of view of the first camera, w is the width of the first camera sensor, f is the focal length of the first camera, and arctan() is the arctangent operation.
[0121] Similarly, the first vertical field of view can be calculated as follows:
[0122]
[0123] Where VFOV is the first vertical field of view of the first camera, h is the height of the first camera sensor, f is the focal length of the first camera, and arctan() is the arctangent operation.
[0124] Similarly, the second field of view for the second camera can also include a second horizontal field of view and a second vertical field of view. The calculation method can refer to the calculation method of the first field of view, and will not be elaborated here.
[0125] It should be noted that the above calculation of the field of view assumes that the first camera is ideal, without distortion, and that the light is incident parallel to the sensor plane of the first camera. In practical applications, lens distortion correction and other factors may also need to be considered. Therefore, the above calculation method is only an example and does not constitute a specific limitation on the calculation of the field of view.
[0126] S403: Obtain the current zoom ratio of the first camera and calculate the target field of view of the first camera based on the current zoom ratio.
[0127] Specifically, such as Figure 8 As shown, the current zoom ratio is obtained, and the target field of view angle corresponding to the first camera at the current zoom ratio is calculated based on the current zoom ratio and the first parameter.
[0128] The explanation of the current zoom ratio can be found in the relevant description in step S202 above, and will not be repeated here. It should be noted that the current zoom ratio refers to the magnification or reduction ratio of the first camera when shooting, which is determined by the user through operations such as clicking the magnification adjustment control (or pinching with two fingers). The target field of view refers to the field of view of the first camera at the current zoom ratio, which is usually smaller (or larger) than the first field of view. The relevant description of the field of view can be found in the relevant description in step S402 above, and will not be repeated here.
[0129] Specifically, when the current zoom ratio of the camera is larger, that is, when the magnification of the camera is greater, for example, when the zoom ratio is adjusted from the basic 1x to 2x, the focal length will become longer. According to the relevant explanation of calculating the field of view in step S402 above, it can be determined that the field of view of the first camera will be smaller.
[0130] S404: Generate a first cropping box based on the target field of view and a first field of view, and generate a second cropping box based on the target field of view and a second field of view.
[0131] Specifically, such as Figure 8 As shown, based on the first field of view and the first parameter, a first cropping box corresponding to the first image under the target field of view is generated, wherein the first image is an image with the first field of view at the base magnification. Similarly, based on the second field of view, a second cropping box corresponding to the second image under the target field of view is generated, wherein the second image is an image with the second field of view at the base magnification.
[0132] In electronic devices, a series of processing steps or stages organized in a specific order to process input data and generate output data can be called a pipeline. A pipeline can typically include multiple stages, each of which can handle a specific task.
[0133] In some possible implementations, when using an electronic device to take a picture, the process from displaying a preview shooting interface to capturing and saving the target image may involve two pipelines: a preview pipeline and a shooting pipeline. The preview pipeline refers to the image processing pipeline used to display image data captured by the first camera in real time, while the shooting pipeline refers to the image processing pipeline used to process the image taking function. Figure 9 This is a schematic diagram of a preview pipeline and a shooting pipeline provided in an embodiment of this application, such as... Figure 9 As shown, the preview pipeline can include three stages: image acquisition, image processing, and display output, while the shooting pipeline can include four stages: triggering shooting, image acquisition, image processing, and file saving.
[0134] For example, after a user inputs a first action into the electronic device, the electronic device can begin executing a preview pipeline, such as... Figure 9 As shown, firstly, an image is acquired using the first camera, and the acquired first image is output to the next stage. Further, before image processing, the preview pipeline can obtain the current zoom ratio and the first parameter. The current zoom ratio can be referred to the relevant description in step S202 above, which is adjusted or determined by the user through the second screen. During the image processing stage, the preview pipeline can process the first image based on the current zoom ratio and the first parameter. The first image is the image at the base magnification. The algorithm can generate a cropping box corresponding to the first image at the current zoom ratio to realize the magnification or reduction of the first image, so as to obtain the image that needs to be displayed on the current shooting preview interface, which changes in real time in response to the user's adjustment of the current zoom ratio before the second operation, and is displayed and output on the second screen.
[0135] Specifically, when the user inputs a second operation to instruct the electronic device to save the image currently displayed on the shooting preview interface as the target shooting image, the main process involves the shooting pipeline acquiring and saving the corresponding target shooting image. Figure 6 The process shown mainly involves the execution of the shooting pipeline, where... Figure 6 In the illustrated process, all steps except S305 and S311 can be executed during the image processing stage of the shooting pipeline. For example, in this embodiment, when the user inputs a second operation into the electronic device, the shooting pipeline begins execution, such as... Figure 9As shown, the electronic device's shooting pipeline acquires a first image through a first camera and a second image through a second camera, and then inputs these images into the image processing stage of the shooting pipeline. At this time, the first and second images input into the image processing stage are still images acquired by the corresponding cameras at the base magnification. Therefore, when performing step S301: aligning the first and second images with their field of view to obtain the aligned first and second images, step S404 needs to be executed: generating a first cropping frame based on the target field of view and the first field of view, and generating a first cropping frame based on the target field of view and the second field of view, to obtain the cropping frame corresponding to the first image under the target field of view to crop the first image to obtain the third image displayed in the shooting preview interface, under the current zoom ratio, which is the image that the user wants to shoot. That is, the first image acquired by the shooting pipeline at the base magnification is not the image in the shooting preview interface (with a smaller or larger field of view). It needs to be processed similarly to the preview pipeline to obtain the image in the actual shooting preview interface in the background.
[0136] Among them, such as Figure 9 As shown, the image acquisition stages of the preview pipeline and the shooting pipeline are corresponding by dashed lines, indicating that the two can partially overlap. The "overlap" mentioned here means that both the preview pipeline and the shooting pipeline can acquire the first image from the first camera, and the first image is an electrical signal converted from light captured by the first camera sensor. Therefore, the image acquisition stages partially overlap.
[0137] For example, the field of view of a rear camera is typically smaller than that of a front camera. Figure 10A This is a schematic diagram of a first cutting frame provided in an embodiment of this application. Figure 10B This is a schematic diagram of a second clipping frame provided in an embodiment of this application, as shown below. Figure 10A As shown, the first field of view is the field of view at the base magnification of the first camera, with an angle of 'a'; the second field of view is the field of view at the base magnification of the second camera, with an angle of 'c'; and the target field of view is the field of view corresponding to the current zoom ratio, with an angle of 'a'. The target field of view can be smaller than the first field of view and smaller than the second field of view, as shown below. Figure 10A As shown, when the "current zoom ratio" is greater than 1, for example, 2x, it means that the image is magnified by 2 times during shooting. At this time, the target field of view is smaller than the first field of view. Since the first field of view is usually smaller than the second field of view, the target field of view is also smaller than the second field of view. Based on the first field of view corresponding to the first image and the target field of view, a first cropping box can be generated. Further, as... Figure 10BAs shown, a second cropping frame can be generated based on the second field of view corresponding to the second image and the target field of view. Furthermore, when the "current zoom ratio" is less than 1, for example, 0.5x, it means the zoom was reduced by a factor of 2 during shooting. In this case, the target field of view is greater than the first field of view, and may be greater than the second field of view (or may be less than the second field of view). This embodiment of the application does not limit this.
[0138] S405: Crop the first image and the second image based on the first cropping frame and the second cropping frame respectively to obtain the first image and the second image after field of view alignment.
[0139] Specifically, such as Figure 8 As shown, cropping the first image based on the first cropping box yields the first image after field-of-view alignment, and cropping the second image based on the second cropping box yields the second image after field-of-view alignment.
[0140] S302: Based on the size of the first image after field of view alignment, scale the first image and the second image after field of view alignment to obtain the third image and the fourth image.
[0141] Specifically, such as Figure 8 As shown, the second image, which is also aligned to the field of view, is scaled based on the size of the first image after field of view alignment. The first image, after scale alignment, becomes the third image, and the second image, after scale alignment, becomes the fourth image. Here, "scale" in scale alignment includes not only the image size but also the degree of blur or detail. In one possible implementation, scale alignment can be achieved by upsampling (or downsampling) to align the resolution of the fourth image with that of the third image.
[0142] S303: Determine whether the third and fourth images contain the complete target face.
[0143] Specifically, to ensure the accuracy of subsequent pose correction and gaze correction, both the third and fourth images obtained after scale alignment should contain at least a complete target face (so that they can be used as a reference for pose correction and gaze correction). Therefore, a target face integrity assessment is required for the third and fourth images. Furthermore, since scale alignment involves cropping the first and second images, the target face may be complete in the first (and second) images, but incomplete after scale alignment (including cropping). However, since the target image ultimately saved in this embodiment is obtained based on the scale-aligned third image through subsequent image processing (including pose correction and gaze correction), the determination of target face integrity needs to be based on both the third and fourth images.
[0144] Figure 11 This is a schematic diagram of a face integrity assessment process provided in an embodiment of this application. In one possible implementation, determining whether the third and fourth images contain a complete target face may include... Figure 11 Steps S501-S506 shown in the figure:
[0145] S501: Perform image preprocessing on the third and fourth images.
[0146] Specifically, in order to improve the accuracy of facial feature point recognition, it is necessary to first perform image preprocessing on the third and fourth images, which may include image grayscale conversion and image blurring.
[0147] For example, image grayscale conversion refers to converting a color image into a grayscale image. By performing image grayscale conversion on the third and fourth images respectively, the color information in the third and fourth images can be removed, so that they only contain brightness information. This can reduce the amount of data that needs to be processed, help reduce the complexity of the algorithm, and improve the speed of facial feature point recognition. Image blurring refers to averaging or weighted averaging the pixels of lines or shadow boundaries in an image. This can smooth the image and reduce the influence of noise. Applying image blurring to the third and fourth images can help the algorithm more accurately identify edges and contours in the image, thereby improving the accuracy of recognition.
[0148] S502: Perform facial feature point detection on the third and fourth images.
[0149] Specifically, facial feature point recognition algorithms are used to perform facial feature point recognition on the third and fourth images respectively.
[0150] S503: Determine whether the third and fourth images contain facial feature points.
[0151] Specifically, based on the results of facial feature point recognition, it is determined whether the third and fourth images contain facial key points.
[0152] For example, the third image may contain complete facial feature points or it may contain partial facial feature points. Both of these situations fall under the category of "containing facial feature points". Similarly, the fourth image also needs to be judged for facial feature points. When both the third and fourth images contain facial key points, step S504 is executed: judge whether the third and fourth images contain complete faces based on facial feature points. If the third image does not contain facial feature points or the fourth image does not contain facial key points, or neither of them contains facial feature points, step S506 is executed: save the third image as the target image.
[0153] S504: Determine whether the third and fourth images contain a complete human face based on facial feature points.
[0154] Specifically, when both the third and fourth images contain facial feature points, it is necessary to further determine whether the target face in the third and fourth images is complete. When both the third and fourth images contain a complete target face, step S505 is executed: pose correction is performed on the third and fourth images. When the third image does not contain a complete target face, or the fourth image does not contain a complete target face, or neither the third nor the fourth image contains a complete target face, step S506 is executed.
[0155] Figure 12A This is a schematic diagram of standard facial feature points provided in an embodiment of this application. The schematic diagram of standard facial feature points represents the complete and standard facial feature points of a target face when it is looking straight ahead. In one possible implementation, the completeness of the target face can be determined by comparing the degree of missing facial feature points obtained from facial feature point detection in the third and fourth images with the standard facial feature points. For example, if the facial feature points extracted from the third (or fourth) image lack a certain threshold of facial feature points compared to the standard facial feature points, it can be determined that the third (or fourth) image does not contain a complete target face. The threshold can be a pre-set threshold, for example, it can be set to 20%. When the facial feature points obtained from the third (or fourth) image lack 20% or more of the facial feature points compared to the standard facial feature points, it can be determined that the third (or fourth) image does not contain a complete target face.
[0156] For example, when the target face in the third and fourth images is directly facing (or nearly directly facing) the first camera 1081, the facial feature point detection algorithm can obtain facial feature point detection results that are relatively close to standard facial feature points. Based on these detection results, it can be determined that the third image (or fourth image) contains a complete target face. When the target face in the third or fourth image undergoes a certain degree of deflection, tilting its head up or down, the facial feature point detection algorithm can still detect facial feature points relatively accurately. However, some facial feature points may overlap or become crowded due to the deflection of the face. Based on this type of facial feature point detection result, it can still be determined that the corresponding target face is complete. For example, Figure 12B This is a schematic diagram of the target facial feature point detection result provided in an embodiment of this application, such as... Figure 12B As shown, although the yaw and pitch angles of the target face are not zero at this time, the facial feature point detection algorithm can still detect relatively complete and accurate facial feature points. However, some facial feature points overlap or are crowded. Figure 12B The facial feature point detection results shown indicate that the corresponding third (or fourth) image contains the complete target face; similarly, Figure 12C This is a schematic diagram of another target facial feature point detection result provided in an embodiment of this application, which can also be based on... Figure 12C The facial feature point detection results shown indicate that the corresponding third (or fourth) image contains the complete target face.
[0157] It should be noted that, Figure 12B , Figure 12C This is an exemplary display of the face feature point detection results provided in the embodiments of this application, and does not constitute a specific limitation on the embodiments of this application. The face feature point detection results of the target face in the third or fourth image may also be in other forms, and the embodiments of this application do not limit them.
[0158] Figure 12D This is a schematic diagram of another target facial feature point detection result provided in an embodiment of this application, such as... Figure 12D As shown, when the target face is not fully visible in the first image acquired by the first camera or the second image acquired by the second camera due to angular deviation or displacement during shooting, the facial feature points obtained by facial feature point detection in the third or fourth image will be incomplete; or, when the target face is complete in the first and second images, but becomes incomplete in the third and fourth images after scale alignment processing (including cropping), the facial feature points obtained by facial feature point detection in the third or fourth image will also be incomplete. Based on Figure 12D The facial feature point detection results shown indicate that the third (or fourth) image does not contain a complete target face.
[0159] S505: Perform pose correction on the third and fourth images.
[0160] Specifically, when both the third and fourth images contain the complete target face, pose correction is performed on the third and fourth images, i.e., step S304 is executed.
[0161] S506: Save the third image as the target image.
[0162] Specifically, when the third or fourth image does not contain the target face, since there is no human figure, the folding screen shooting method provided in this application embodiment does not need to be executed. In this case, the third image is directly saved as the target shooting image, which can reduce the amount of computation. When the third image does not contain a complete target face, it is impossible to perform subsequent pose correction on the incomplete target face in the third image. When the fourth image does not contain a complete target face, it is impossible to provide reference information for pose correction through the fourth image. In both of these cases, the folding screen shooting method provided in this application embodiment cannot be executed, and the third image is also directly saved as the target shooting image. In some other possible implementations, when the third and fourth images belong to the above two cases, the process corresponding to the folding screen shooting method provided in this application embodiment can be exited, and the shooting process can be executed normally. For example, in response to the user's second operation, the normal shooting process (such as shooting pipeline; for shooting pipeline, please refer to the above) can be executed. Figure 9 (The relevant description will not be repeated here) The image on the second screen's shooting preview interface is saved as the target shooting image to the electronic device.
[0163] It should be noted that step S506 is actually step S305.
[0164] S304: Calculate the first pose and the second pose in the third and fourth images respectively.
[0165] Specifically, based on the face feature point detection results in step S303, the three-dimensional portrait model, and the first and second parameters, the first pose in the third image and the second pose in the fourth image are calculated respectively. Figure 13 This is a schematic diagram of a pose correction process provided in an embodiment of this application. Step S304: calculating the first pose and second pose in the third and fourth images, and step S306: correcting the first pose in the third image based on the second pose in the fourth image to obtain the fifth image, can be referred to... Figure 13 .
[0166] In one possible implementation, the first pose can refer to the pose of the target face in the third image in the real three-dimensional world. Similarly, the second pose can refer to the pose of the target face in the fourth image in the real three-dimensional world. The pose in the three-dimensional world can be represented by coordinates in a three-dimensional coordinate system. The target face in the third or fourth image is presented by arranging multiple specific pixels. Therefore, calculating the pose of the target face in the two-dimensional image (including the third and fourth images) is to calculate the (three-dimensional) coordinates of the target face in the three-dimensional coordinate system by using the (two-dimensional) coordinates of the pixels corresponding to the target face in the two-dimensional image (including the third and fourth images) in the image's two-dimensional coordinate system.
[0167] For example, the pixels (multiple) of the target face in the third and fourth images are fixed, and their coordinates in the coordinate system corresponding to the two-dimensional image are also fixed. Therefore, the two-dimensional coordinates of the corresponding pixels of the target face in the third and fourth images can be determined first, and then the two-dimensional coordinates of the pixels can be inversely calculated based on the intrinsic parameter matrix of the corresponding camera to obtain the coordinates of the target face in the three-dimensional coordinate system. The intrinsic parameter matrix is a fixed matrix describing the internal geometry and optical characteristics of the camera, capable of mapping any three-dimensional coordinate in the camera coordinate system to a two-dimensional coordinate in the pixel coordinate system. It is a parameter used to calculate the position of an object in the real three-dimensional world in the two-dimensional image acquired by the camera. The intrinsic parameter matrix typically includes the following key parameters: 1. Focal length: Determines the size of the image formed on the sensor. The longer the focal length, the larger the image formed on the sensor. It is usually represented by f. 2. Pixel size: The size of each pixel on the camera sensor. It is usually represented by (dx, dy), where dx represents the length of each pixel and dy represents the width of each pixel. 3. Principal coordinates: The coordinates of the optical center O in the pixel coordinate system. It is usually represented as (u0, v0). Different cameras correspond to different intrinsic parameter matrices. The intrinsic parameter matrix of the first camera can be determined by the first parameter, and the intrinsic parameter matrix of the second camera can be determined by the second parameter.
[0168] In one possible implementation, the coordinate values of the pixels corresponding to the target face region in the third image and the coordinate values of the pixels corresponding to the target face region in the fourth image can be determined first. Further, the intrinsic parameter matrices of the first and second cameras are calibrated based on the first and second parameters. Then, based on the three-dimensional portrait model and the face feature point detection information corresponding to the third and fourth images, the poses of the target face in the third and fourth images are calculated respectively. The first pose can be obtained by calculating the third image, and the second pose can be obtained by calculating the fourth image.
[0169] For example, the position of an object in the real world can be represented by coordinates in a camera coordinate system, which is a three-dimensional world coordinate system with the optical center of the camera as the origin. The three-dimensional coordinate axes of the camera coordinate system can be set differently depending on the camera. Figure 14A This is a schematic diagram of a camera coordinate system provided in an embodiment of this application, such as... Figure 14A As shown, the optical center O of the camera is the origin of the camera coordinate system. The optical center refers to a special point within the camera. When the camera includes a single lens, the optical center can be the geometric center of the lens. When the camera includes multiple lenses, the optical center can be considered the equivalent geometric center or optical center of the entire lens assembly. Furthermore, the y-axis is defined by a straight line passing through the origin O and vertically upwards; the z-axis is defined by a straight line passing through the origin O and extending inwards; and the z-axis is defined by a straight line passing through the origin O and extending horizontally to the right (where "right" refers to the direction of the viewpoint). Figure 14A The straight line (in the right-hand direction) is the x-axis, perpendicular to the z-axis, and the plane at a distance of focal length f from the optical center O is the image plane. Figure 14B This is a top view of a camera coordinate system provided in an embodiment of this application, combined with... Figure 14B This can more clearly illustrate the relationship between the image plane and the camera coordinate system; the image plane refers to the imaging plane, which can be understood as the plane on which the image captured by the camera is located (or directly as the image acquired by the camera). The image plane is a two-dimensional plane, and the position on the image plane can be represented using coordinates in the pixel coordinate system. The pixel coordinate system is a two-dimensional coordinate system, and as... Figure 14A As shown, the origin R of the pixel coordinate system is usually located at the upper left corner of the image plane (or the image acquired by the camera). The straight line extending horizontally through the origin R is the x' axis of the image plane, and the straight line extending vertically downward through the origin R is the y' axis. Any point in the image plane can be represented by two-dimensional coordinates.
[0170] For example, such as Figure 14A As shown, for any point p in the real world, its coordinates in the camera coordinate system can be p(X,Y,Z), which is mapped to point q in the image plane. The coordinates q(u,v) of point q in the pixel coordinate system of the corresponding camera can be obtained through the intrinsic parameter matrix of the corresponding camera. The calculation method is as follows:
[0171]
[0172] Therefore, the pixels and pixel coordinates corresponding to the target face region in the third image are determined. For example, to reduce the computational load, the pixel coordinates corresponding to the facial feature points of the target face in the third image can be determined. Further, based on the intrinsic parameter matrix of the first camera, the pixels corresponding to the target face region in the detected third image are inversely operated on, and combined with the three-dimensional standard portrait, the position (or coordinates) of the target face in the third image in the camera coordinate system can be determined, and the first pose of the target face can be calculated based on the three-dimensional coordinates. Similarly, the pixels and pixel coordinates corresponding to the target face region in the fourth image are determined. To reduce the computational load, the pixel coordinates corresponding to the facial feature points of the target face in the fourth image can be determined. The pixels of the target face region in the detected fourth image are inversely operated on through the intrinsic parameter matrix of the second camera, and combined with the three-dimensional standard portrait, the position (or coordinates) of the pixels of the target face region in the fourth image in the camera coordinate system can be determined, and the second pose of the target face can be calculated based on the three-dimensional coordinates.
[0173] After obtaining the first and second poses, the first pose needs to be corrected based on the second pose. This correction needs to be performed in a two-dimensional image, specifically by correcting individual pixels (multiple pixels) in the two-dimensional images (the third and fourth images). Therefore, to more easily and clearly measure the changes or differences in pose represented by the three-dimensional coordinates in the two-dimensional image, the three-dimensional pose of the target face can be represented or described in different ways in the two-dimensional image. Specific explanations can be found in the relevant descriptions in step S205 above, and will not be repeated here. For example, regarding... Figure 2A The first camera 1081 shown in the figure captures the image, which is then processed to obtain... Figure 15A The third image shown, Figure 15A This is a schematic diagram of a third image provided in an embodiment of this application. It can be determined through calculation. Figure 15A The first pose of the target face in the third image shown is: yaw angle 30 degrees (yaw to the right), pitch angle -20 degrees (head down), roll angle 20 degrees (tilt to the right); Figure 15B This is a fourth image schematic diagram provided in an embodiment of this application, which can be determined through calculation. Figure 15B The second pose corresponding to the target face in the fourth image shown is: yaw angle 5 degrees (yaw to the right), pitch angle -15 degrees (head down), roll angle 0 degrees (tilt to the right).
[0174] It should be noted that the above Figure 15A and Figure 15B This is an exemplary illustration of the third image and the first pose, and does not constitute a specific limitation on the embodiments of this application. In addition, the first pose and the second pose may have other representations. The above description of the pose is only an exemplary illustration and does not constitute a specific limitation on the embodiments of this application.
[0175] S305: Save the third image as the target image.
[0176] For details on the specific procedures, please refer to the relevant description in step S506 above, which will not be repeated here.
[0177] S306: Correct the first pose in the third image based on the second pose in the fourth image to obtain the fifth image.
[0178] Specifically, such as Figure 13 As shown, the pose transformation matrix of the first pose of the target face in the third image relative to the second pose of the target face in the fourth image can be calculated. Based on the pose transformation matrix, matrix calculations can be performed on the corresponding pixels of the target face in the third image to change the position of the corresponding pixels and correct the first pose so that the corrected first pose is aligned with the second pose. In addition, during the correction process, the processing range of pose correction can be limited based on the information determined in the face feature point detection.
[0179] For example, by calculating and generating a pose transformation matrix for the target face in the third image and performing matrix calculations on the third image based on this pose transformation matrix, to avoid incorrectly processing pixels of the background portion (such as other objects, buildings, etc.) in the third image during pose correction, the target face region in the third image can be determined during face feature point detection in step S303 above, and a corresponding target face mask can be generated. This target face mask is used to limit the pose transformation area when performing pose correction on the third image based on the pose transformation matrix. Here, the mask typically refers to a binary image (or grayscale image) with the same size as the original image. The selected region that needs processing or is of interest can be labeled with a specific grayscale value, while the remaining regions can be set to another value. For example, for... Figure 15A The third image shown can determine, as Figure 15C The target face mask shown is... Figure 15C This is a schematic diagram of a target face mask provided in an embodiment of this application, such as... Figure 15C As shown, the target face mask is a binary image with the same size as the third image, where the pixel values of the target face region are set to 0 (white area), and the pixel values of other regions are set to 255 (black area). This is based on the pose transformation matrix... Figure 15A When performing pose correction on the third image shown, it can be done through... Figure 15C The target face mask shown defines the area for pose transformation processing, and only applies to... Figure 15A In the third image shown Figure 15C The white area shown is used to perform pose transformation processing on the target face area, which can avoid image distortion caused by incorrect processing of the background or other parts that do not need to be processed.
[0180] It should be noted that, Figure 15C The target face mask shown is just one example provided by the embodiments of this application. In other possible implementations, the target face region in the target face mask can be set to other pixel values, and can include more or less regions of the image corresponding to the target face. The embodiments of this application do not limit this.
[0181] For example, based on the pose transformation matrix and Figure 15C The target face mask shown will Figure 15A The first pose is aligned in the third image shown. Figure 15B The second pose in the fourth image shown can be obtained Figure 15D The fifth image shown, Figure 15D This is a fifth schematic diagram provided in an embodiment of this application. Figure 15DThe fifth image shown is... Figure 15A The third image shown Figure 15B Compared to the fourth image shown, the pose of the target face has been aligned. Figure 15B The fourth image shown, specifically, Figure 15D The yaw angle of the target face in the fifth image shown is determined by... Figure 15A The 30-degree rightward correction shown is the same as... Figure 15B Around 5 degrees; Figure 15D The pitch angle of the target face in the fifth image shown is determined by... Figure 15A The -20 degree (head down) correction shown is the same as... Figure 15B Around 18 degrees Celsius; Figure 15D The roll angle of the target face in the fifth image shown is determined by... Figure 15A The 20-degree rightward tilt correction shown is the same as... Figure 15B Around 0 degrees Celsius, and Figure 15D Image quality and Figure 15A The third image shown is the same as or similar to the one shown.
[0182] S307: Determine whether the target face in the fifth image has missing facial content.
[0183] Specifically, when a user takes a picture using an electronic device, due to issues such as angle deviation, the target face in the image captured by the first camera may be missing. This missing facial content refers to situations where part of the target face is shown in the camera's image due to head rotation (turning, tilting, or looking down), while another part is obscured. See [reference needed]. Figure 16A , Figure 16A This is another schematic diagram of a third image provided in the embodiments of this application. In this case, the target face still appears completely (without being cropped out by the camera's viewfinder) in the third and fourth images. When the electronic device performs face feature point detection in step S304, it can still obtain complete and relatively accurate face feature points (see the above for details). Figure 12B and Figure 12C (Related descriptions), however, after pose correction, due to missing facial content, the fifth image may have incomplete facial features, making the final saved target image not meet user requirements and causing a decline in user experience. Therefore, it is necessary to determine whether the target face has missing content in the fifth image, and if so, image completion processing is required for the fifth image.
[0184] Specifically, it can be determined whether there is missing facial content in the fifth image by calculating the edge difference of the target face in the fifth image compared with the third image. If there is no missing facial content in the fifth image, step S308 is executed; if there is, step S309 is executed.
[0185] For example, Figure 16A This is a third image obtained by the first camera and scale-aligned, in which the user corresponding to the target face is rotated to the right relative to the first camera during the shooting process (within...). Figure 16A The image shown (right side) displays the user's left profile completely, but the right profile is partially obscured by the left profile due to the image's rotation. Specifically, the right eye, right eyebrow, and half of the right lip are missing from the right profile. Figure 16B This is another fourth schematic diagram provided in the embodiments of this application. Figure 16B The fourth image shown is the fourth image obtained by the second camera after scale alignment. The user corresponding to the target face is tilted to the right relative to the first camera during the capture. Due to the positional relationship between the first and second cameras (see the relevant description in the application scenario above), the target face is basically facing forward in both the second and fourth images obtained by the second camera (the tilt angle of the target face is small). Figure 16B The fourth image shown Figure 16A The pose correction of the third image shown can be obtained Figure 16C The fifth image shown, Figure 16C This is another fifth image schematic diagram provided in the embodiments of this application, such as... Figure 16C As shown, the pose of the target face is aligned through pose correction. Figure 16B The fourth image shown is the second pose of the target face. Here, the target face region includes the user's hair and facial area. Figure 16C In the fifth image shown, the right eye, right eyebrow, and right half of the lip on the right side of the target face are incomplete. Figure 16D This is a schematic diagram of an edge difference provided in an embodiment of this application, such as... Figure 16D As shown, by comparing the third image with the fifth image, the edge difference between the edge of the target face in the fifth image and the edge of the target face in the third image can be determined. When the edge difference is greater than a certain threshold, it indicates that the user corresponding to the target face in the third image has made a large head rotation, and it can be determined that the target face has lost facial content due to the rotation.
[0186] S308: Based on the fourth and fifth images, calculate the gaze transition matrix corresponding to the fifth image.
[0187] Specifically, after pose correction, the pose of the target face in the fifth image is aligned with that of the third image, but the gaze of the eyes on the target face is still skewed. Therefore, gaze correction is needed for the target face. Usually, the target face in the fourth image is looking straight ahead, that is, the target face is not skewed. Therefore, the gaze of the target face in the fourth image is usually also looking straight ahead. The gaze of the target face in the fifth image can be corrected based on the gaze of the target face in the fourth image.
[0188] Figure 17A This is a schematic diagram of a gaze correction process provided in an embodiment of this application. Step S308: Calculate the gaze transformation matrix corresponding to the fifth image based on the fourth and fifth images. Step S310: Perform gaze correction on the target face in the fifth image based on the gaze transformation matrix to obtain the target image. The process can be found in [reference needed]. Figure 17A . Figure 17B This is a schematic diagram of another vision correction process provided in an embodiment of this application. The following is in conjunction with… Figure 17A and Figure 17B This section details the specific process for calculating the line-of-sight transition matrix.
[0189] In one possible implementation, calculating the gaze transition matrix may include steps S601-S604:
[0190] S601: Preprocess the fifth and fourth images.
[0191] Specifically, in order to improve the accuracy of eye feature point detection, the fourth and fifth images are first preprocessed, which may include image grayscale conversion and image blurring. The specific processing procedure can be referred to the relevant description in step S501 above, and will not be repeated here.
[0192] S602: Perform eye feature point detection on the fourth and fifth images.
[0193] Specifically, the algorithm detects eye feature points on the target face in the fourth and fifth images, respectively. For example, the detected eye feature points may include iris boundary points and pupil boundary points, and the eye center point and pupil center point are determined based on the iris boundary points and pupil boundary points.
[0194] Figure 17C This application provides an embodiment of an eye image of a target face, such as... Figure 17C As shown, the iris is a pigmented ring-shaped membrane at the front of the eyeball, with a circular opening in the center called the pupil. The center point of the pupil usually refers to the center of the best-fit circle of the pupil. The rotation of the eyeball can be inferred from the center point of the pupil, thereby further estimating the direction of the line of sight. The center point of the eye is the center point of the eyeball. The line connecting the center point of the pupil and the center point of the eye is the optical axis. Under ideal conditions, the optical axis can be approximated as the direction of the line of sight.
[0195] In one possible implementation, algorithms can be used to detect the eye feature points of the target face. Figure 17D This is a schematic diagram of the eye feature points of a target face provided in an embodiment of this application, such as... Figure 17D As shown, the point on the iris boundary is the iris boundary point, and the point on the pupil boundary is the pupil boundary point. Furthermore, the pupil center point can be determined through the pupil boundary point. Then, based on the iris boundary point, pupil boundary point, and pupil center point, the eye center point can be determined. The eye center point is the center point of the eyeball.
[0196] In some other possible implementations, algorithms can also be used to determine other eye feature points on the target face, including but not limited to corner feature points (including inner and outer corners of the eyes) and eyelid feature points (including upper and lower eyelid feature points). These eye feature points can help to better determine the gaze of the target face, and this application example does not limit this.
[0197] In one possible implementation, when detecting eye feature points in the fifth image, an eye mask for the target face in the fifth image can be generated. The description of the mask can be found in the aforementioned description of the target face mask, and will not be repeated here. The eye mask of the target face in the fifth image can help limit the processing area for gaze correction in step S310, preventing incorrect processing of other parts of the fifth image outside the eye area and causing distortion in the final saved target image. Figure 15E This is a schematic diagram of an eye mask for a target face in a fifth image provided in an embodiment of this application. Figure 15E The target face eye mask shown is Figure 15D The fifth image shown corresponds to the target face eye mask, where the white area represents the target face eye region. It should be noted that... Figure 15E The target face eye mask shown is an exemplary illustration of an eye mask and does not constitute a specific limitation on the embodiments of this application.
[0198] S603: Calculate the first and second lines of sight based on eye feature points.
[0199] Specifically, the line of sight corresponding to the target face in the fourth and fifth images is determined by the center point of the eye and the center point of the pupil, respectively. The line of sight corresponding to the target face in the fifth image is the first line of sight, and the line of sight corresponding to the target face in the fourth image is the second line of sight.
[0200] For example, see Figure 17E , Figure 17EThis is a schematic diagram of the gaze corresponding to a target face provided in an embodiment of this application. The line connecting the center point of the pupil and the center point of the eye is the optical axis. Here, the optical axis refers to the direction in which light enters the pupil of the user corresponding to the target face. The user's gaze usually refers to the visual axis direction, where the visual axis represents the path of light from an object after refraction through the cornea and lens and focusing on the retina. However, in gaze correction, the user's visual axis cannot be easily determined. There is usually a deviation angle between the visual axis and the optical axis. This deviation angle is called the Kappa angle. The Kappa angle is small and varies due to differences in the eye structure between different people. In the folding screen shooting method provided in this embodiment of the application, the gaze correction process does not require high precision. Therefore, the optical axis direction can be approximated as the gaze direction. Figure 17E As shown, the center point of the eye can be connected to the center point of the pupil and extended outwards from the eyeball as a line of sight; in other possible implementations, the Kappa angle can be considered, and the first and second lines of sight can be determined more accurately based on eye feature points, which is not limited in this application embodiment.
[0201] S604: Generate a line-of-sight transformation matrix based on the first and second lines of sight.
[0202] Specifically, a gaze transformation matrix for the fifth image is generated based on the first gaze, the second gaze, and the eye mask of the target face. For a description of the transformation matrix, please refer to the relevant description in step S306 above.
[0203] S309: Perform completion processing on the fifth image based on the fourth image to obtain the target captured image with facial content completion.
[0204] Specifically, the fifth image can be completed using deep learning methods, for example, by generating a target image with facial content completion based on the fourth and fifth images using a model.
[0205] In one possible implementation, image completion processing can be performed using a generative network model. Figure 18A This is a schematic diagram of a fifth image completion process provided in an embodiment of this application. Figure 18B This is another schematic diagram of the fifth image completion process provided in the embodiments of this application. The following is in conjunction with... Figure 18A and Figure 18B This describes a possible process for fifth-image completion based on a generative adversarial network, which may include steps S701-S705:
[0206] S701: Generate the target face corresponding to the fifth image based on the fourth and fifth images.
[0207] For example, such as Figure 18BAs shown, the fourth and fifth images can be input into the generator, which then generates the corresponding completed target face in the fifth image. The fourth and fifth images can be preprocessed before being input into the generator. This preprocessing may include adjusting image size and normalizing pixel values. Normalizing pixel values means normalizing the pixel values of each pixel in the fourth and fifth images; for example, the pixel values in the fourth and fifth images can be normalized to between 0 and 1. This reduces the size difference of the input data, helps the generator converge faster, reduces the computational burden, and improves the generator's performance.
[0208] Furthermore, the generator generates the target face corresponding to the fifth image, that is, it only generates the target face region in the fifth image. This reduces the amount of data that the generator needs to generate, reduces the computational burden, and avoids erroneously generating other regions in the fifth image, which would cause distortion of the target image. Therefore, when generating the target face corresponding to the fifth image, the generator can also determine the region of the target face in the fifth image through facial feature point detection, and generate the target face mask of the fifth image based on the region. The relevant description of the mask can be referred to the relevant description in step S306 above, and will not be repeated here.
[0209] S702: Extract the target face from the first image.
[0210] Specifically, in order to iteratively optimize the target face corresponding to the fifth image generated in step S701 to obtain the target face that best matches the user's pose in the fifth image, it is necessary to provide a face acquired by the first camera with the same pose as the target face acquired by the second camera as prompt information. Here, pose refers to the overall pose of the target face. The first image is an image acquired by the first camera with the same or similar image specifications as the fifth image, and the pose of the target face in the image is exactly the same as that in the fifth image. Face extraction can be used to extract a face from the first image with the same pose as the target face in the fifth image. At the same time, the extracted target face has a pose similar or approximately the same as the target face in the second image (and the fourth image), which can be used as prompt information for correcting the target face in step S703.
[0211] S703: Correct the target face corresponding to the completion of the fifth image based on the loss function.
[0212] For example, such as Figure 18BAs shown, the difference between the target face in the fifth image generated by the generator and the target face in the extracted first image can be measured by the loss function. The generator minimizes the loss function to correct the target face in the fifth image and make it more realistic and natural.
[0213] S704: Replace the target face in the fifth image.
[0214] For example, the target face corresponding to the corrected fifth image is replaced back in the fifth image to obtain the fifth image after the target face is replaced.
[0215] S705: The discriminator performs discrimination and feedback processing on the replaced fifth image to obtain the target image.
[0216] For example, such as Figure 18B As shown, the discriminator can be used to distinguish the fifth image after the target face is replaced. The discriminator can determine whether the fifth image after the target face is replaced is a real image and generate feedback information to guide the generator to perform feedback processing. For example, it can guide the generator to continuously iterate and optimize the fifth image after the target face is replaced, and finally obtain the target image. Figure 16E This is a schematic diagram of another target image provided in an embodiment of this application. Figure 16E The target image shown is obtained by capturing images of the target through... Figure 16C The fifth image shown was obtained by completing the image. Figure 16E and Figure 16C In contrast, the right side of the target face is complete, and Figure 16E The pose of the human face in the target image shown is referenced from the fourth and fifth images.
[0217] It should be noted that, Figure 16E This is an exemplary demonstration of a target image provided in an embodiment of this application, and does not constitute a specific limitation on the examples in this application.
[0218] S310: Based on the gaze transformation matrix, perform gaze correction on the target face in the fifth image to obtain the target image.
[0219] Specifically, the gaze corresponding to the target face in the fifth image is corrected based on the gaze transformation matrix and the eye mask of the target face to obtain the target image.
[0220] Figure 15F This is a schematic diagram of a target image captured according to an embodiment of this application. Figure 15F The target image shown is based on Figure 15E The target face eye mask shown Figure 15DThe fifth image shown is obtained by applying the corresponding gaze transformation matrix, which is based on... Figure 15C The second line of sight in the fourth image shown and Figure 15D The first line of sight in the fifth image shown is obtained as follows: Figure 15F As shown, with Figure 15D Compared to the fifth image shown, Figure 15F In the image shown, the gaze of the target face in the captured image changes from a slanted gaze to a rightward gaze, and then to a direct gaze. Figure 15B Compared to the fourth image shown, Figure 15F The line of sight corresponding to the target face in the target image shown is aligned with the fourth image, resulting in a target image with high image quality where the target face pose is aligned with the fourth image obtained by the front camera (with a small skew angle).
[0221] S311: Save the target image to the gallery and display a thumbnail of the target image on the second screen.
[0222] Specifically, after acquiring the target image, the electronic device saves the target image to the gallery application and displays a thumbnail of the target image on the second screen 1092.
[0223] For example, Figure 2B This is a schematic diagram of the second screen during the shooting process of the electronic device provided in this application embodiment. Figure 2B The user interface 54 is displayed on the second screen 1092. When the user inputs a second operation to the electronic device through the second screen 1092, the electronic device saves the target image to the electronic device and displays a thumbnail of the saved target image in the "View Image" control in the operation menu bar shown on the second screen. The user can click the "View Image" control to view the captured and saved target image.
[0224] It should be noted that the electronic device can save the captured image to the local device or to an application used by the electronic device to view the image. Users can view the image through the application or access or view the image directly in the local file. The "image library" application mentioned above is just a general term for applications in the electronic device that allow users to view images. In other possible implementations, the application may have other possible names, and this application embodiment does not limit this.
[0225] Figure 19 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0226] like Figure 19As shown, the electronic device 100 may include a processor 101, a memory 102, a wireless communication module 103, a mobile communication module 104, an antenna 103A, an antenna 104A, a power switch 105, a sensor module 106, a focusing motor 107, a camera 108, a display screen 109, etc. The sensor module 106 may include a gyroscope sensor 106A, an accelerometer sensor 106B, an ambient light sensor 106C, an image sensor 106D, a proximity sensor 106E, etc. The wireless communication module 103 may include a WLAN communication module, a Bluetooth communication module, etc. All of the above components can transmit data via a bus.
[0227] Processor 101 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc.
[0228] The GPU, or Graphics Processing Unit, also known as a display core, visual processor, or display chip, is a miniature processor specifically designed to perform image and graphics-related computations. It is responsible for performing complex mathematical and geometric calculations to render images, videos, and other graphical content. Through its highly parallel architecture and numerous computing units, the GPU excels at handling large-scale parallel computing tasks, especially in graphics rendering. This design allows the GPU to provide tens or even hundreds of times the performance of the CPU in areas such as floating-point operations and parallel computing. The GPU can be a standalone device or integrated into the processor 101.
[0229] The processor 101 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 101 is a cache memory. This memory can store instructions or data that the processor 101 has just used or that are used repeatedly. If the processor 101 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 101, and thus improves the efficiency of the system.
[0230] In some embodiments, the processor 101 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0231] Memory 102 can be used to store computer executable program code, which may include instructions. Processor 101 executes various functional applications and data processing of electronic device 100 by running the instructions stored in memory 102. Memory 102 may include a program storage area and a data storage area. In specific implementations, memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices.
[0232] The wireless communication function of the electronic device 100 can be implemented through antenna 103A, antenna 104A, mobile communication module 104, wireless communication module 103, modem processor, and baseband processor.
[0233] Antennas 103A and 104A can be used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0234] The mobile communication module 104 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on electronic devices 100.
[0235] The wireless communication module 103 can provide solutions for wireless communication applications on electronic devices 100, including wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR).
[0236] The gyroscope sensor 106A can be used to determine the motion attitude of the electronic device 100.
[0237] Accelerometer 106B can detect the magnitude of acceleration of electronic device 100 in various directions (generally three axes).
[0238] The 106C ambient light sensor can detect the intensity of ambient light and automatically adjust the screen brightness.
[0239] The image sensor 106D can capture optical images (including visible light, infrared light, etc.) and convert them into electrical signals for processing, display, or storage.
[0240] The 106E distance sensor can measure the distance between an object and the sensor.
[0241] Electronic device 100 can implement display functions through GPU, display screen 109, and AP, etc. GPU is a microprocessor for image processing, connected to display screen 109 and AP. Processor 101 may include one or more GPUs, which execute program instructions to generate or change display information.
[0242] The display screen 109 is used to display images, videos, etc. The display screen 109 includes a display panel. The display panel can be a liquid crystal display (LCD). The display panel can also be manufactured using organic light-emitting diodes (OLEDs), active-matrix organic light-emitting diodes (AMOLEDs), flexible light-emitting diodes (FLEDs), miniled, microled, micro-OLEDs, quantum dot light-emitting diodes (QLEDs), etc. In some embodiments, the electronic device may include one or N displays 109, where N is a positive integer greater than 1.
[0243] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0244] Figure 20 This is a schematic diagram of the system architecture of an electronic device provided in an embodiment of the present invention.
[0245] A layered architecture divides the system into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom: application layer, application framework layer, hardware abstraction layer, driver layer, and hardware layer.
[0246] The application layer can include a series of application packages. For example, it can include a camera, a gallery, etc.
[0247] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 20 As shown, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, etc.
[0248] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0249] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0250] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0251] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0252] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog-style notifications on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0253] In this embodiment, the camera application can utilize the interfaces and services (such as content providers) provided by the application framework layer to effectively access media files (including pictures, videos, etc.) stored on the device, and at the same time use the view system to build a user interface to display these media contents in an intuitive way.
[0254] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In this embodiment, the hardware abstraction layer may include a camera hardware abstraction layer, a camera algorithm library, and a sensor control center.
[0255] The camera hardware abstraction layer can provide virtual hardware for camera device 1, camera device 2, camera device 3, or more camera devices; the camera algorithm library includes various components related to image processing algorithms and technologies, which can improve the quality of photos and videos and enrich shooting effects; the sensor control center can include a main control chip that processes data from various sensors and makes corresponding control decisions based on this data. The sensor control center can also control the operation of the camera through the camera module, including turning the camera on or off, adjusting parameters such as focus and exposure, to ensure that the camera can make precise adjustments according to shooting needs, thereby capturing high-quality images and videos. In addition, the sensor control center can also include sensor nodes, which act as a bridge connecting upper-layer applications and lower-layer hardware, encapsulating and managing sensor interfaces related to imaging, and providing key environmental information to help the camera shoot more accurately.
[0256] The driver layer is the layer between hardware and software. It includes drivers for various hardware components, such as camera drivers, image processor drivers, display drivers, sensor drivers, and digital signal processor drivers.
[0257] The camera device driver is used to drive the camera sensor to acquire images and to drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process images. The image processor driver is used to drive the graphics processor to process images.
[0258] The hardware layer is the most fundamental layer in a computer system or embedded system, directly involving the existence and operation of physical hardware devices. The hardware layer is the foundation upon which software can run, including all physically tangible and visible computer components, as well as those invisible but equally crucial components such as integrated circuits and circuit boards. The hardware layer can include sensors, image signal processors, digital signal processors, and image processors. Sensors can include multispectral sensors, light sensors, and time-of-flight (TOF) sensors. A multispectral sensor is a sensor capable of simultaneously acquiring information from multiple optical spectrum bands, typically including visible light and bands extending into infrared and ultraviolet light, capturing richer spectral information to improve the color reproduction and accuracy of the phone. A light sensor is a sensor capable of detecting the intensity of ambient light. A TOF sensor is a 3D imaging technology that calculates the distance to an object by measuring the time difference between the emission and reflection of a light pulse, providing more accurate spatial positioning and depth information.
[0259] It should be noted that the specific process of the foldable screen shooting method described in the embodiments of this application can be found in the above. Figures 1A-18BThe relevant descriptions in the application embodiments described herein will not be repeated here.
[0260] This application also provides a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the method executed by the electronic device in any of the above embodiments.
[0261] This application also provides a computer storage medium storing a computer program (also referred to as code or instructions). When the computer program is run, it causes the computer to perform the method executed by the electronic device in any of the above embodiments.
[0262] This application also provides a computer program product including instructions that, when executed by a computer, enable the computer to perform the method executed by the electronic device in any of the above embodiments.
[0263] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0264] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0265] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0266] In summary, the above description is merely an embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the disclosure of this application should be included within the scope of protection of this application.
Claims
1. A method for taking photos with a foldable screen, characterized in that, The method is applied to an electronic device, the electronic device including a foldable screen, the foldable screen including a first folding body and a second folding body, the first folding body including a first screen and a first camera disposed on different sides, the second folding body including a second screen and a second camera disposed on the same side, the method including: Upon receiving the user's first operation, the second screen displays a shooting preview interface captured by the first camera; Upon receiving a second user operation, if the inward folding angle between the first folding body and the second folding body is within a preset range, then a first image is acquired through the first camera and a second image is acquired through the second camera; the second operation is used to instruct the image displayed in the current shooting preview interface to be saved as the target shooting image; The first image and the second image are respectively subjected to a first image processing to obtain a third image and a fourth image, wherein the third image and the fourth image have the same scale and field of view. Determine whether the third image and the fourth image contain a complete target face. If they do, perform second image processing on the third image based on the fourth image to obtain a fifth image. The pose of the target face in the fifth image is the same as the pose of the target face in the third image. Based on the fourth and fifth images, the target image is saved.
2. The method according to claim 1, characterized in that, The step of determining whether the third image and the fourth image contain a complete target face includes: Facial feature point detection is performed on the third and fourth images; Based on the results of the facial feature point detection, it is determined whether the third image and the fourth image contain the complete target face.
3. The method according to claim 1 or 2, characterized in that, The step of performing a first image processing on the first image and the second image respectively to obtain the third image and the fourth image includes: The first image and the second image are aligned in terms of field of view to obtain the first image and the second image after field of view alignment. Based on the size of the first image after field-of-view alignment, the first image and the second image after field-of-view alignment are scale-aligned to obtain the third image and the fourth image.
4. The method according to claim 3, characterized in that, The process of aligning the field of view of the first image and the second image includes: Obtain the first parameter of the first camera and the second parameter of the second camera; Based on the first parameter and the second parameter, a first cropping box corresponding to the first image and a second cropping box corresponding to the second image are generated. The first image is cropped based on the first cropping frame to obtain the first image after the field of view is aligned; The second image is cropped based on the second cropping frame to obtain the second image after the field of view is aligned.
5. The method according to claim 4, characterized in that, The step of generating the first cropping box corresponding to the first image and the second cropping box corresponding to the second image includes: Calculate the first field of view of the first camera based on the first parameter, and calculate the second field of view of the second camera based on the second parameter; Obtain the current zoom ratio of the first camera, and calculate the target field of view of the first camera based on the current zoom ratio; Based on the target field of view and the first field of view, a first cropping box for the first image is generated, and based on the target field of view and the second field of view, a second cropping box for the second image is generated.
6. The method according to any one of claims 1-5, characterized in that, The step of performing a second image processing on the third image based on the fourth image to obtain a fifth image includes: Calculate the first pose and the second pose of the target face in the third image and the fourth image, respectively; The pose of the first pose in the third image is corrected based on the second pose to obtain the fifth image.
7. The method according to any one of claims 1-6, characterized in that, After performing a second image processing on the third image based on the fourth image to obtain the fifth image, the method further includes: By comparing the third image and the fifth image, it is determined whether the target face in the fifth image is missing facial content. If so, the fifth image is completed to obtain the target image.
8. The method according to claim 7, characterized in that, The completion process for the fifth image includes: A complete face is generated based on the target face in the fourth and fifth images; The target face in the fifth image is replaced with the completed face to obtain a target captured image with completed facial content.
9. The method according to any one of claims 1-8, characterized in that, Saving the target image based on the fourth and fifth images includes: Based on the fourth image, the fifth image is processed by a third image to obtain a target image, wherein the target image has the same gaze direction as the target face in the fourth image; The target image is saved to the gallery, and a thumbnail of the target image is displayed on the shooting preview interface on the second screen.
10. The method according to claim 9, characterized in that, The step of performing a third image processing on the fifth image based on the fourth image to obtain the target image includes: Based on the fourth image and the fifth image, calculate the gaze transition matrix corresponding to the target face in the fifth image; Based on the gaze transformation matrix, the gaze of the target face in the fifth image is corrected to obtain the target image.
11. The method according to any one of claims 1-10, characterized in that, The first camera is a rear camera, and the second camera is a front camera; The first image is acquired through the first camera, while the second image is acquired through the second camera.
12. An electronic device, characterized in that, The electronic device includes a memory and one or more processors; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-11.
13. A chip system applied to an electronic device, characterized in that, The chip system includes one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-11.
14. A computer storage medium, characterized in that, The computer storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-11.
15. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a computer, enable the computer to perform the method described in any one of claims 1-11.