Image processing method and related device

The image processing method addresses perspective distortion and shake-induced blurring by correcting cropped image regions, improving video quality by eliminating the Jello effect and maintaining frame-to-frame consistency.

JP7708206B2Active Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023559812
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-29
Filing Date
2022-03-25
Publication Date
2025-07-15
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing image processing technologies fail to effectively correct perspective distortion and shake-induced blurring in images captured by wide-angle and medium-focus cameras, leading to low-quality video output with noticeable Jello phenomena.

Method used

An image processing method that performs shake correction and perspective distortion correction on cropped regions of images, ensuring frame-to-frame consistency by maintaining the same degree of transformation for adjacent frames, thereby reducing the Jello effect.

Benefits of technology

The method improves video quality by eliminating the Jello phenomenon and maintaining frame-to-frame consistency, ensuring that the same object appears with minimal shape difference between frames, thus enhancing the overall video output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708206000003
    Figure 0007708206000003
  • Figure 0007708206000004
    Figure 0007708206000004
  • Figure 0007708206000005
    Figure 0007708206000005
Patent Text Reader

Abstract

The present application provides an image processing method, comprising: performing stabilization on two adjacent image frames of a video stream acquired by a camera to obtain a cropped region; and performing perspective correction on the cropped region acquired through stabilization. The cropped regions output after stabilization is performed on the two adjacent frames contain the same or essentially the same content, so that the subsequently generated video can be guaranteed to be smooth and stable. In the present application, for perspective correction, the positional relationship between the cropped regions acquired through stabilization is further mapped, so that the degree or effect of perspective correction of adjacent frames of the subsequently generated video is the same, in other words, the degree of transformation of the same object between the images is the same. In this way, the inter-frame consistency of the subsequently generated video is maintained, obvious Jello phenomenon is prevented, and the quality of the captured video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application was filed with the China National Intellectual Property Administration on March 29, 2021, and claims priority to Chinese Patent Application No. 202110333994.9, entitled "IMAGE PROCESSING METHOD AND RELATED DEVICE", which is hereby incorporated by reference in its entirety.

[0002] This application relates to the field of image processing, and in particular, to an image processing method and related device.

Background Art

[0003] When the camera is in the mode where the wide-angle camera acquires an image or the medium-focus camera acquires an image, the acquired original image has obvious perspective distortion. Specifically, perspective distortion (also referred to as 3D distortion) is the transformation in 3D object imaging (e.g., horizontal elongation, radial elongation, and combinations thereof) caused by different magnifications due to the difference in the depth of field of the subject for shooting when the three-dimensional shape in space is mapped onto the image plane. Therefore, perspective distortion correction needs to be performed on the image.

[0004] When the user holds the terminal for shooting, the user cannot keep the posture of holding the terminal stable, so the terminal shakes in the direction of the image plane. Specifically speaking, when the user hopes to hold the terminal to shoot a target area, since the user's posture when holding the terminal is unstable, the target area has a large offset between adjacent image frames acquired by the camera, and as a result, the video acquired by the camera contains blurs.

[0005] However, the fact that the terminal can shake is a common problem, and the existing technologies for performing perspective distortion correction are not mature. As a result, the corrected video still has low quality.

Summary of the Invention

[0006] Embodiments of the present application provide an image display method, as a result of which an obvious Jello phenomenon does not occur, thereby improving the quality of video output.

[0007] According to a first aspect, the present application provides an image processing method. The method is used by a terminal to acquire a video stream in real time, and the method includes: performing shake correction on the acquired first image to obtain a first cropped area; performing perspective distortion correction on the first cropped area to obtain a third cropped area; obtaining a third image based on the first image and the third cropped area; and performing shake correction on the acquired second image to obtain a second cropped area The second cropped area is related to the first cropped area and shake information, and the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image; for shake correction, the offset shift of the photo in the viewfinder frame of the latter's original image frame from the photo of the former's original image frame (including the offset direction and the offset distance) needs to be determined. In one implementation, the offset shift may be determined based on the shake that occurs in the process of the terminal acquiring two adjacent original image frames. Specifically, the offset shift may be a posture change that occurs in the process of the terminal acquiring two adjacent original image frames. The posture change may be obtained based on the photographed posture corresponding to the viewfinder frame in the previous original image frame and the posture when the terminal photographs the current original image frame.

[0008] The method includes: performing perspective distortion correction on the second cropped area to obtain a fourth cropped area; and obtaining a fourth image based on the second image and the fourth cropped area further comprising, wherein the third image and the fourth image are used to generate a target video, the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain.

[0009] Different from performing perspective distortion correction on all regions of the original image according to the existing implementation, in this embodiment of the present application, the perspective distortion correction is performed on the first cropped region. When the terminal device shakes, content misalignment usually occurs between two adjacent original image frames acquired by the camera. Therefore, the positions of the cropped regions to be output after performing shake correction on the two image frames are different in the original image, that is, the distances between the centers of the cropped regions and the original image to be output after performing shake correction on the two image frames are different. However, the sizes of the cropped regions to be output after performing shake correction on the two frames are the same. Therefore, when perspective distortion correction is performed on the cropped regions obtained through shake correction, the degree of perspective distortion correction performed on adjacent frames is the same (because the distances between each sub-region in each of the cropped regions output through shake correction for the two frames and the center points of the cropped regions output through shake correction are the same). In addition, the cropped regions output through shake correction for two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object in the images is the same, that is, frame-to-frame consistency is maintained. Therefore, an obvious Jello phenomenon does not occur, thereby improving the quality of the video output. The frame-to-frame consistency in this specification can also be referred to as time-domain consistency, and it should be understood that it indicates that the processing results for regions with the same content in adjacent frames are the same.

[0010] Therefore, in this embodiment of the present application, perspective distortion correction is not performed on the first image, but perspective distortion correction is performed on the first cropped area obtained through shake correction, and a third cropped area is obtained. The same process is further performed for a frame (second image) adjacent to the first image, that is, perspective distortion correction is performed on the second cropped area obtained through shake correction, and a fourth cropped area is obtained. In this way, the degree of conversion correction of the same object in the third cropped area and the fourth cropped area is the same. As a result, frame-to-frame consistency is maintained, and no obvious Jello phenomenon occurs, thereby improving the quality of the video output.

[0011] In this embodiment of the present application, the same target object (for example, a human object) may be included in the cropped areas of the former frame and the latter frame. Since the object for perspective distortion correction is the cropped area obtained through shake correction, for the target object, the difference between the degrees of conversion of the target object in the former frame and the latter frame due to perspective distortion correction is very small. The fact that the difference is very small can be understood as that it is difficult for the human naked eye to distinguish the shape difference, or it is difficult for the human naked eye to discover the Jello phenomenon between the former frame and the latter frame.

[0012] The image processing method according to the first aspect can be further understood as follows: For the processing of a single image (e.g., the second image) in the acquired video stream, specifically, hand shake correction may be performed on the acquired second image to obtain a second cropped region, where the second cropped region is related to the first cropped region and shake information; the first cropped region is adjacent to the second image and the hand shake correction is obtained after being performed on the image frame that is adjacent to the second image and before the second image in the time domain, where the shake information indicates the shake that occurs in the process of the terminal acquiring the second image and the image frame that is adjacent to the second image and before the second image in the time domain; perspective distortion correction is performed on the second cropped region to obtain a fourth cropped region, the fourth image is obtained based on the second image and the fourth cropped region, and the target video is generated based on the fourth image. The second image may be any image that is not the first in the acquired video stream, and it should be understood that the image that is not the first refers to the image frame that is not the first in the video stream in the time domain. The target video may be obtained by processing a plurality of images in the video stream in the same manner as the way of processing the second image.

[0013] In a possible implementation, the direction of the offset of the position of the second cropped region in the second image with respect to the position of the first cropped region in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal acquiring the first image and the second image.

[0014] Performing shake correction is to eliminate the shake in the acquired video caused by the change in the posture when the user holds the terminal device. Specifically, performing shake correction is to crop the original image acquired by the camera in order to obtain the output in the viewfinder frame. The output in the viewfinder frame is the output of shake correction. When the terminal has a large posture change within a very short time, there is a large offset of the photo in the viewfinder frame of the latter image frame from the photo of the former image frame. For example, when the terminal acquires the (n + 1)-th frame compared to the case where the n-th frame is acquired, there is shake in the direction A within the direction of the image plane (it should be understood that the shake in the direction A may be the shake of the main optical axis of the camera on the terminal in the direction A or the shake of the position of the optical center of the camera on the terminal in the direction A within the direction of the image plane). In this case, the photo in the viewfinder frame of the (n + 1)-th image frame is shifted in the direction opposite to the direction A in the n-th image frame.

[0015] In a possible implementation, the first cropped region represents the first sub-region in the first image, and the second cropped region represents the second sub-region in the second image.

[0016] The similarity between the first image content corresponding to the first sub-region in the first image and the second image content corresponding to the second sub-region in the second image is higher than the similarity between the first image and the second image. The similarity between image contents can be understood as the similarity between scenarios in the image, the similarity between backgrounds in the image, or the similarity between target subjects in the image.

[0017] In a possible implementation, the similarity between images may be understood as the overall similarity between the image contents within the image regions at the same position. Due to the shaking of the terminal, the image contents of the first image and the second image are displaced, and the positions of the same image content in the images are different. In other words, the similarity within the image content in the image regions at the same position in the first image and the second image is lower than the similarity between the image contents in the image regions at the same position in the first sub-region and the second sub-region.

[0018] In a possible implementation, alternatively, the similarity between images may be understood as the similarity between the image contents included in the images. Due to the shaking of the terminal, there are portions of image content in the edge region of the first image that do not exist in the second image, and similarly, there are portions of image content in the edge region of the second image that do not exist in the first image. In other words, both the first image and the second image contain image content that does not exist in each other, and the similarity between the image contents included in the first image and the second image is lower than the similarity between the image contents included in the first sub-region and the second sub-region.

[0019] In addition, when pixels are used as the granularity for determining similarity, the similarity may be represented by the overall similarity between the pixel values of the pixels at the same position within the image.

[0020] In a possible implementation, before performing perspective distortion correction on the first cropped region, the method: further includes the steps of detecting that the terminal meets the distortion correction conditions and enabling the distortion correction function. It can be further understood that the current shooting environment, shooting parameters, or shooting content of the terminal meet the conditions under which perspective distortion correction needs to be performed.

[0021] The point in time when it is detected that the terminal satisfies the distortion correction condition and the action to enable the distortion correction function is executed may be before the stage of performing shake correction on the acquired first image, or after the shake correction has been performed on the acquired first image and before the perspective distortion correction is performed on the first cropped area. It should be understood that this is not limited in the present application.

[0022] In one implementation, the perspective distortion correction needs to be enabled, that is, the perspective distortion correction is executed on the first cropped area only when it is detected that the perspective distortion correction is currently in the enabled state.

[0023] In a possible implementation, detecting that the terminal satisfies the distortion correction condition includes one or more of the following cases, but is not limited thereto:

[0024] Case 1: It is detected that the zoom ratio of the terminal is lower than a first preset threshold.

[0025] In one implementation, whether the perspective distortion correction is enabled may be determined based on the zoom ratio of the terminal. When the zoom ratio of the terminal is large (for example, within the full or partial zoom ratio range of the telephoto shooting mode, or within the full or partial zoom ratio range of the medium focal length shooting mode), the degree of perspective distortion in the photo of the image acquired by the terminal is low. The fact that the degree of perspective distortion is low can be understood as that almost no perspective distortion in the photo of the image acquired by the terminal can be identified by the human eye. Therefore, when the zoom ratio of the terminal is large, the perspective distortion correction does not need to be enabled.

[0026] In a possible implementation, when the zoom ratio ranges from a to 1, the terminal acquires a video stream in real time by using the front wide-angle camera or the rear wide-angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or When the zoom ratio ranges from 1 to b, the terminal acquires a video stream in real time by using the rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0027] Case 2: A first activation operation of the user to enable the perspective distortion correction function is detected.

[0028] In a possible implementation, the first activation operation includes a first operation on a first control in the shooting interface on the terminal. The first control is used to instruct to enable or disable the perspective distortion correction, and the first operation is used to instruct to enable the perspective distortion correction.

[0029] In one implementation, the control (referred to as the first control in this embodiment of the present application) used to instruct to enable or disable the perspective distortion correction may be included in the shooting interface. The user may trigger to enable the perspective distortion correction through the first operation on the first control, that is, trigger to enable the perspective distortion correction through the first operation on the first control. In this way, the terminal may detect the first activation operation of the user to enable the perspective distortion correction, and the first activation operation includes the first operation on the first control in the shooting interface on the terminal.

[0030] In one implementation, the control (referred to as the first control in this embodiment of the present application) used to instruct to enable or disable the perspective distortion correction is displayed in the shooting interface only when it is detected that the zoom ratio of the terminal is lower than the first preset threshold.

[0031] Case 3: A human face is identified in the shooting scenario; A human face is recognized in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or A human face is recognized in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio.

[0032] When a human face exists in a shooting scenario, the degree of transformation of the human face due to perspective distortion correction is more visually obvious, and the distance between the human face and the terminal is smaller or the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher, then the size of the human face in the image is larger. Therefore, it should be understood that the degree of transformation of the human face due to perspective distortion correction is even more visually obvious. Thus, perspective distortion correction needs to be enabled in the aforementioned scenarios. In this embodiment, whether to enable perspective distortion correction is determined by determining the aforementioned conditions related to the human face in the shooting scenario. As a result, the shooting scenarios in which perspective distortion occurs can be accurately determined, and perspective distortion correction is executed for the shooting scenarios in which perspective distortion occurs. In shooting scenarios where no perspective distortion occurs, perspective distortion correction is not executed. In this way, the image signal is accurately processed and power consumption is reduced.

[0033] The shooting scenario can be understood as an image acquired before the camera acquires the first image, or if the first image is the first image frame of the video acquired by the camera, the shooting scenario should be understood to be the first image, or an image close to the first image in the time domain. Whether there is a human face in the shooting scenario, whether the distance between the human face and the terminal is smaller than a preset value, and the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario may be determined by using a neural network or in another way. The preset value may range from 0m to 10m. For example, the preset value may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. This is not limited in this specification. The ratio of the pixels may range from 30% to 100%. For example, the ratio of the pixels may be 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100%. This is not limited in this specification.

[0034] In a possible implementation, before performing shake correction on the acquired first image, the method: The stage of detecting that the terminal meets the shake correction condition and enabling the shake correction function is further provided. It can be further understood that the current shooting environment, shooting parameters, or shooting content of the terminal meet the conditions under which shake correction needs to be performed.

[0035] In a possible implementation, detecting that the terminal meets the shake correction condition includes one or more of the following cases, but is not limited thereto:

[0036] Case 1: It is detected that the zoom ratio of the terminal is higher than a second preset threshold.

[0037] In one implementation, the difference between the second preset threshold and the fixed zoom ratio of the target camera on the terminal device is within a preset range, where the target camera is the camera on the terminal with the minimum fixed zoom ratio.

[0038] In one implementation, whether to enable the shake correction function may be determined based on the zoom ratio of the terminal. When the zoom ratio of the terminal is small (for example, within the full or partial zoom ratio range of the wide-angle shooting mode), the size of the physical area corresponding to the photo displayed in the viewfinder frame in the shooting interface is already very small. If further shake correction is performed, that is, if the image is further cropped, the size of the physical area corresponding to the photo displayed in the viewfinder frame will become smaller. Therefore, the shake correction function may be enabled only when the zoom ratio of the terminal is higher than a specific preset threshold.

[0039] The target camera may be a front wide-angle camera, and the preset range may be 0 to 0.3. For example, when the fixed zoom ratio of the front wide-angle camera is 0.5, the second preset threshold may be 0.5, 0.6, 0.7, or 0.8.

[0040] Case 2: A second activation operation by the user to enable the shake correction function is detected.

[0041] In one implementation, the second activation operation includes a second operation on a second control in the shooting interface on the terminal, the second control being used to indicate enabling or disabling the shake correction function, and the second operation being used to indicate enabling the shake correction function.

[0042] In one implementation, a control (referred to as the second control in this embodiment of the present application) used to instruct enabling or disabling of the hand shake correction function may be included within the shooting interface. The user may enable the hand shake correction function through a second operation on the second control, that is, trigger enabling of the hand shake correction function through the second operation on the second control. In this way, the terminal may detect a second enabling operation of the user to enable the hand shake correction function, and the second enabling operation includes the second operation on the second control in the shooting interface on the terminal.

[0043] In one implementation, the control (referred to as the second control in this embodiment of the present application) used to instruct enabling or disabling of the hand shake correction function is displayed in the shooting interface only when it is detected that the zoom ratio of the terminal is higher than or equal to a second preset threshold.

[0044] In a possible implementation, the step of obtaining the fourth image based on the second image and the fourth cropped area includes: obtaining a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates the mapping relationship between each position point in the second cropped area and the corresponding position point in the second image, and the second mapping relationship indicates the mapping relationship between each position point in the fourth cropped area and the corresponding position point in the second cropped area; and obtaining the fourth image based on the second image, the first mapping relationship, and the second mapping relationship.

[0045] In a possible implementation, the step of obtaining the fourth image based on the second image, the first mapping relationship, and the second mapping relationship includes: combining the first mapping relationship with the second mapping relationship to determine a target mapping relationship, where the target mapping relationship includes the mapping relationship between each position point in the fourth cropped area and the corresponding position point in the second image; and Determining a fourth image based on the second image and the target mapping relationship including.

[0046] That is, the index tables for the first mapping relationship and the second mapping relationship are combined. By combining the index tables, after the cropped region is obtained through hand shake correction, it is not necessary to perform a warp operation and then perform the warp operation again after perspective distortion correction. It is only necessary to perform a warp operation once after the output is obtained based on the combined index table after perspective distortion correction, thereby reducing the overhead of the warp operation. The warp operation may refer to an affine transformation operation on the image. For specific implementations, refer to existing warp techniques.

[0047] In a possible implementation, the first mapping relationship is different from the second mapping relationship.

[0048] In a possible implementation, the step of performing perspective distortion correction on the first cropped region includes: performing optical distortion correction on the first cropped region to obtain a corrected first cropped region; and performing perspective distortion correction on the corrected first cropped region. The step of obtaining a third image based on the first image and the third cropped region includes: performing optical distortion correction on the third cropped region to obtain a corrected third cropped region; and obtaining a third image based on the first image and the corrected third cropped region. The step of performing perspective distortion correction on the second cropped region includes: performing optical distortion correction on the second cropped region to obtain a corrected second cropped region; and performing perspective distortion correction on the corrected second cropped region; or The step of obtaining the fourth image based on the second image and the fourth cropped region includes: performing optical distortion correction on the fourth cropped region to obtain a corrected fourth cropped region; and obtaining the fourth image based on the second image and the corrected fourth cropped region.

[0049] Compared with combining the optical distortion correction index table and the perspective distortion correction index table by using the distortion correction module, optical distortion is an important cause of the jelly phenomenon, and the effectiveness of optical distortion correction and perspective distortion correction is controlled in the video shake correction backend. Therefore, this embodiment can solve the problem that when the video shake correction module invalidates the optical distortion correction, the perspective distortion correction module cannot control the jelly phenomenon.

[0050] According to a second aspect, the present application provides an image processing apparatus. The apparatus is used by a terminal to obtain a video stream in real time, and the apparatus includes: A shake correction module configured to perform shake correction on the obtained first image to obtain a first cropped region, and further configured to perform shake correction on the obtained second image to obtain a second cropped region, where the second cropped region is related to the first cropped region and shake information, and the shake information indicates the shake that occurs in the process of the terminal obtaining the first image and the second image, and the first image and the second image are a pair of original images obtained by the terminal and adjacent to each other in the time domain; A perspective distortion correction module configured to perform perspective distortion correction on the first cropped region to obtain a third cropped region, and further configured to perform perspective distortion correction on the second cropped region to obtain a fourth cropped region; and configured to obtain a third image based on the first image and the third cropped region, obtain a fourth image based on the second image and the fourth cropped region, and further configured to generate a target video based on the third image and the fourth image comprising

[0051] In a possible implementation, the direction of the offset of the position of the second cropped region in the second image with respect to the position of the first cropped region in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal obtaining the first image and the second image.

[0052] In a possible implementation, the first cropped region represents a first sub-region in the first image, and the second cropped region represents a second sub-region in the second image.

[0053] The similarity between the first image content corresponding to the first sub-region in the first image and the second image content corresponding to the second sub-region in the second image is higher than the similarity between the first image and the second image.

[0054] In a possible implementation, the apparatus is: a first detection module configured to detect that the terminal meets the distortion correction condition and enable the distortion correction function before the hand shake correction module performs hand shake correction on the obtained first image further comprising

[0055] In a possible implementation, detecting that the terminal meets the distortion correction condition is: detecting that the zoom ratio of the terminal is lower than a first preset threshold but not limited thereto.

[0056] In a possible implementation, when the zoom ratio ranges from a to 1, the terminal obtains a video stream in real time by using the front wide-angle camera or the rear wide-angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or When the zoom ratio ranges from 1 to b, the terminal obtains a video stream in real time by using the rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0057] In a possible implementation, for the terminal to detect that it meets the distortion correction condition: Detecting a first activation operation of the user to activate the perspective distortion correction function, where the first activation operation includes a first operation on a first control in the shooting interface on the terminal, the first control is used to instruct to activate or deactivate the perspective distortion correction, and the first operation is used to instruct to activate the perspective distortion correction but is not limited thereto.

[0058] In a possible implementation, for the terminal to detect that it meets the distortion correction condition: Identifying a human face in a shooting scenario; Identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or Identifying a human face in a shooting scenario, where the proportion of pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset proportion but is not limited thereto.

[0059] In a possible implementation, the device is: Before the hand shake correction module performs hand shake correction on the acquired first image, a second detection module configured to detect that the terminal satisfies the hand shake correction condition and enable the hand shake correction function is further provided.

[0060] In a possible implementation, detecting that the terminal satisfies the hand shake correction condition includes: detecting that the zoom ratio of the terminal is higher than the fixed zoom ratio of the camera with the minimum zoom ratio on the terminal but is not limited thereto.

[0061] In a possible implementation, detecting that the terminal satisfies the hand shake correction condition includes: detecting a second activation operation of the user to enable the hand shake correction function, where the second activation operation includes a second operation on a second control in the shooting interface on the terminal, the second control is used to instruct to enable or disable the hand shake correction function, and the second operation is used to instruct to enable the hand shake correction function but is not limited thereto.

[0062] In a possible implementation, the image generation module is specifically configured to: acquire a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates the mapping relationship between each position point in the second cropped area and the corresponding position point in the second image, and the second mapping relationship indicates the mapping relationship between each position point in the fourth cropped area and the corresponding position point in the second cropped area; and acquire a fourth image based on the second image, the first mapping relationship, and the second mapping relationship is specifically configured to perform.

[0063] In a possible implementation, the perspective distortion correction module is: Determining a target mapping relationship based on a second mapping relationship for a first mapping relationship, where the target mapping relationship includes a mapping relationship between each position point in a fourth cropped region and a corresponding position point in a second image; and Determining a fourth image based on the target mapping relationship and the second image is specifically configured to perform.

[0064] In a possible implementation, the first mapping relationship is different from the second mapping relationship.

[0065] In a possible implementation, the perspective distortion correction module is specifically configured to: perform optical distortion correction on a first cropped region to obtain a corrected first cropped region; and perform perspective distortion correction on the corrected first cropped region; The image generation module is specifically configured to: perform optical distortion correction on a third cropped region to obtain a corrected third cropped region; and obtain a third image based on a first image and the corrected third cropped region; The perspective distortion correction module is specifically configured to: perform optical distortion correction on a second cropped region to obtain a corrected second cropped region; and perform perspective distortion correction on the corrected second cropped region; or The image generation module is specifically configured to: perform optical distortion correction on a fourth cropped region to obtain a corrected fourth cropped region; and obtain a fourth image based on a second image and the corrected fourth cropped region.

[0066] According to a third aspect, the present application provides an image processing method. The method is used by a terminal to obtain a video stream in real time, and the method includes: Determine the target object in the acquired first image, and obtain a first cropped region including the target object; Perform perspective distortion correction on the first cropped region to obtain a third cropped region; Obtain a third image based on the first image and the third cropped region; Determine the target object in the acquired second image, and obtain a second cropped region including the target object, where the second cropped region is related to target movement information, and the target movement information indicates the movement of the target object in the process of the terminal acquiring the first image and the second image; Perform perspective distortion correction on the second cropped region to obtain a fourth cropped region; Obtain a fourth image based on the second image and the fourth cropped region; and Generate a target video based on the third image and the fourth image where the first image and the second image are pairs of original images acquired by the terminal and adjacent to each other in the time domain.

[0067] The image processing method according to the third aspect can be further understood as follows: For the processing of a single image (e.g., the second image) in the acquired video stream, specifically, a target object in the acquired second image is determined, and a second cropped region including the target object is acquired, where the second cropped region is related to target movement information, and the target movement information indicates the movement of the target object in the process where the terminal acquires the second image and an image frame adjacent to the second image and before the second image in the time domain; perspective distortion correction is performed on the second cropped region to acquire a fourth cropped region; a fourth image is acquired based on the second image and the fourth cropped region; a target video is generated based on the fourth image. The second image may be any image that is not the first in the acquired video stream, and it should be understood that the image that is not the first refers to an image frame that is not the first in the video stream in the time domain. The target video may be acquired by processing a plurality of images in the video stream in the same process as the process of processing the second image.

[0068] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.

[0069] Unlike performing perspective distortion correction for all regions of the original image according to existing implementations, in this embodiment of the present application, perspective distortion correction is performed on the cropped regions. Since the target object moves, the positions of the cropped regions including the human object in the former image frame and the latter image frame are different, and the shifts for perspective distortion correction for the position points or pixels in the two output image frames are different. However, the sizes of the cropped regions of the two frames determined after target identification is performed are the same. Therefore, when perspective distortion correction is performed for each of the cropped regions, the degree of perspective distortion correction performed for adjacent frames is the same (because the distance between each position point or pixel in the cropped region of each of the two frames and the center point of the cropped region is the same). In addition, the cropped regions of two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object in the images is the same, that is, frame-to-frame consistency is maintained, so that an obvious Jello phenomenon does not occur, thereby improving the quality of the video output.

[0070] In a possible implementation, before performing perspective distortion correction on the first cropped region, the method: detecting that the terminal meets the distortion correction condition, and enabling the distortion correction function is further provided.

[0071] In a possible implementation, detecting that the terminal meets the distortion correction condition includes, but is not limited to, one or more of the following cases: Case 1: It is detected that the zoom ratio of the terminal is lower than a first preset threshold.

[0072] In one implementation, whether the perspective distortion correction is enabled may be determined based on the zoom ratio of the terminal. When the zoom ratio of the terminal is large (for example, within the full or partial zoom ratio range of the telephoto shooting mode, or within the full or partial zoom ratio range of the medium-focus shooting mode), the degree of perspective distortion in the photo of the image acquired by the terminal is low. The fact that the degree of perspective distortion is low can be understood as that with the human eye, almost no perspective distortion in the photo of the image acquired by the terminal can be distinguished. Therefore, when the zoom ratio of the terminal is large, the perspective distortion correction does not need to be enabled.

[0073] In a possible implementation, when the zoom ratio ranges from a to 1, the terminal acquires a video stream in real time by using the front wide-angle camera or the rear wide-angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or when the zoom ratio ranges from 1 to b, the terminal acquires a video stream in real time by using the rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0074] Case 2: The first activation operation of the user to enable the perspective distortion correction function is detected.

[0075] In a possible implementation, the first activation operation includes the first operation on the first control in the shooting interface on the terminal, the first control is used to instruct to enable or disable the perspective distortion correction, and the first operation is used to instruct to enable the perspective distortion correction.

[0076] In one implementation, a control (referred to as the first control in this embodiment of the present application) used to instruct enabling or disabling of perspective distortion correction may be included within the shooting interface. The user may trigger enabling perspective distortion correction through a first operation on the first control, that is, trigger enabling perspective distortion correction through the first operation on the first control. In this way, the terminal may detect the user's first enabling operation for enabling perspective distortion correction, and the first enabling operation includes the first operation on the first control in the shooting interface on the terminal.

[0077] In one implementation, a control (referred to as the first control in this embodiment of the present application) used to instruct enabling or disabling of perspective distortion correction is displayed in the shooting interface only when it is detected that the zoom ratio of the terminal is lower than a first preset threshold.

[0078] Case 3: A human face is identified in the shooting scenario; A human face is identified in the shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or A human face is identified in the shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio.

[0079] When a human face exists in a shooting scenario, the degree of conversion of the human face due to perspective distortion correction is more visually apparent, and when the distance between the human face and the terminal is smaller or the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher, the size of the human face in the image is larger, so it should be understood that the degree of conversion of the human face due to perspective distortion correction is even more visually apparent. Therefore, perspective distortion correction needs to be enabled in the aforementioned scenario. In this embodiment, whether to enable perspective distortion correction is determined by determining the aforementioned conditions related to the human face in the shooting scenario, and as a result, the shooting scenario in which perspective distortion occurs can be accurately determined, and perspective distortion correction is executed for the shooting scenario in which perspective distortion occurs. In a shooting scenario where no perspective distortion occurs, perspective distortion correction is not executed. In this way, the image signal is accurately processed and power consumption is reduced.

[0080] According to a fourth aspect, the present application provides an image processing apparatus. The apparatus is used by a terminal to acquire a video stream in real time, and the apparatus: is provided with an object determination module configured to determine a target object in the acquired first image and acquire a first cropped area including the target object, and further configured to determine a target object in the acquired second image and acquire a second cropped area including the target object, where the second cropped area is related to target movement information, and the target movement information indicates the movement of the target object in the process of the terminal acquiring the first image and the second image, and the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain; a perspective distortion correction module configured to perform perspective distortion correction on the first cropped area to acquire a third cropped area, and further configured to perform perspective distortion correction on the second cropped area to acquire a fourth cropped area; and configured to obtain a third image based on the first image and the third cropped region, obtain a fourth image based on the second image and the fourth cropped region, and further configured to generate a target video based on the third image and the fourth image comprising.

[0081] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.

[0082] In a possible implementation, the apparatus:[[]] a first detection module configured to detect that the terminal meets the distortion correction condition and enable the distortion correction function before the perspective distortion correction module performs perspective distortion correction on the first cropped region further comprising.

[0083] In a possible implementation, detecting that the terminal meets the distortion correction condition is:[[]] detecting that the zoom ratio of the terminal is lower than a first preset threshold but not limited thereto.

[0084] In a possible implementation, when the zoom ratio ranges from a to 1, the terminal obtains a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or when the zoom ratio ranges from 1 to b, the terminal obtains a video stream in real time by using a rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0085] In a possible implementation, detecting that the terminal meets the distortion correction condition is:[[]] Detecting a first activation operation of a user to activate a perspective distortion correction function, where the first activation operation includes a first operation on a first control in a shooting interface on the terminal, and the first control is used to indicate enabling or disabling perspective distortion correction, and the first operation is used to indicate enabling perspective distortion correction but not limited thereto.

[0086] In a possible implementation, detecting that the terminal meets the distortion correction condition may include: Identifying a human face in a shooting scenario; Identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or Identifying a human face in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio but not limited thereto.

[0087] According to a fifth aspect, the present application provides an image processing device including a processor, a memory, a camera, and a bus. The processor, the memory, and the camera are connected through the bus.

[0088] The camera is configured to acquire video in real time.

[0089] The memory is configured to store a computer program or instructions.

[0090] The processor is configured to call or execute the program or instructions stored in the memory, and is further configured to call the camera to implement the steps in the first aspect or any possible implementation of the first aspect and the steps in the third aspect or any possible implementation of the third aspect.

[0091] According to a sixth aspect, the present application provides a computer storage medium comprising computer instructions. When the computer instructions are executed on an electronic device or a server, the steps in the first aspect or any possible implementation of the first aspect and the steps in the third aspect or any possible implementation of the third aspect are executed.

[0092] According to a seventh aspect, the present application provides a computer program product. When the computer program product is executed on an electronic device or a server, the steps in the first aspect or any possible implementation of the first aspect and the steps in the third aspect or any possible implementation of the third aspect are executed.

[0093] According to an eighth aspect, the present application provides a chip system. The chip system includes a processor configured to support an execution device or a training device when implementing functions in the foregoing aspects, for example, the function of transmitting or processing data or information in the foregoing method. In a possible design, the chip system further includes a memory. The memory is configured to store program instructions and data required for the execution device or the training device. The chip system may include a chip, or may include a chip and another discrete device.

[0094] This application relates to an image processing method, which is used by a terminal to acquire a video stream in real time, and includes the following steps: performing shake correction on the acquired first image to obtain a first cropped area; performing perspective distortion correction on the first cropped area to obtain a third cropped area; obtaining a third image based on the first image and the third cropped area; performing shake correction on the acquired second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image; performing perspective distortion correction on the second cropped area to obtain a fourth cropped area; and obtaining a fourth image based on the second image and the fourth cropped area, where the third image and the fourth image are used to generate a target video, the first image and the second image are acquired by the terminal, and are a pair of original images adjacent to each other in the time domain. An image display method is provided. In addition, the outputs of the shake correction for two adjacent frames are images including the same or basically the same content, and the degree of transformation of the same object in the images is the same, that is, the inter-frame consistency is maintained, so that no obvious Jello phenomenon occurs, thereby improving the quality of video display.

Brief Description of the Drawings

[0095]

Figure 1

[0096]

Figure 2

[0097]

Figure 3

[0098]

Figure 4a

[0099]

Figure 4b-1

Figure 4b-2

Figure 4b-3

Figure 4b-4

[0100]

Figure 4c

[0101]

Figure 4d

[0102]

Figure 4e

[0103]

Figure 5a

[0104]

Figure 5b

[0105]

Figure 6

[0106]

Figure 7

[0107]

Figure 8

[0108]

Figure 9

[0109]

Figure 10

[0110]

Figure 11

[0111]

Figure 12

[0112]

Figure 13

[0113]

Figure 14

[0114]

Figure 15

[0115]

Figure 16

[0116]

Figure 17

[0117]

Figure 18

[0118]

Figure 19

[0119]

Figure 20

[0120]

Figure 21

[0121]

Figure 22a

[0122]

Figure 22b

[0123]

Figure 22c

[0124]

Figure 22d

[0125]

Figure 22e

[0126]

Figure 22f

[0127]

Figure 23A

Figure 23B

[0128]

Figure 24a

[0129]

Figure 24b

[0130]

Figure 25

[0131]

Figure 26

[0132]

Figure 27

[0133]

Figure 28

[0134]

Figure 29

[0135]

Figure 30

[0136]

Figure 31

[0137]

Figure 32

[0138]

Figure 33

[0139]

Figure 34

[0140]

Figure 35

Embodiments for Carrying Out the Invention

[0141] Embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. The terms used in the implementation of the present invention are only intended to describe specific embodiments of the present invention and are not intended to limit the present invention.

[0142] Embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art can understand that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0143] In the description, claims, and attached drawings of this application, terms such as "first", "second", etc. are used to distinguish between similar objects and are not necessarily intended to describe a specific order or sequence. Terms used in such a way are interchangeable in a particular situation, and it should be understood that this is merely a discriminative method for describing objects having the same attributes in the embodiments of this application. In addition, the terms "include", "contain", and any other variations thereof are intended to cover non-exclusive inclusion. As a result, a process, method, system, product, or device comprising a series of units is not necessarily limited to those units and may include other units that are not explicitly listed or are not inherent in such a process, method, system, product, or device.

[0144] For ease of understanding, the structure of the terminal 100 provided in an embodiment of this application is described below using an example. Please refer to FIG. 1. FIG. 1 is a schematic diagram of the structure of a terminal device according to an embodiment of this application.

[0145] As shown in FIG. 1, the terminal 100 may include a processor 110, an external storage interface 120, an internal storage 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, an optical proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, a peripheral light sensor 180L, a bone conduction sensor 180M, etc.

[0146] It can be understood that the structure shown in this embodiment of the present invention does not constitute a specific limitation to the terminal 100. In some other embodiments of the present application, the terminal 100 may include more or fewer components than those shown in the figure, some components may be combined, some components may be divided, or it may have different component arrangements. The components shown in the figure may be implemented by hardware, software, or a combination of software and hardware.

[0147] Processor 110 may include one or more processing units. For example, processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent components or may be integrated into one or more processors.

[0148] The controller may generate an operation control signal based on the instruction operation code and the time sequence signal to complete the control of instruction reading and instruction execution.

[0149] A memory may be further disposed within processor 110 and is configured to store instructions and data. In some embodiments, the memory in processor 110 is a cache memory. The memory may store instructions or data that have just been used or are periodically used by processor 110. When processor 110 needs to reuse an instruction or data, the processor may directly call the instruction or data from the memory. This avoids repeated accesses, reduces the waiting time for processor 110, and improves system efficiency.

[0150] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, a universal serial bus (USB) interface, etc.

[0151] The I2C interface is a bidirectional synchronous serial bus and includes one serial data line (SDA) and one serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 may be individually coupled to the touch sensor 180K, a charger, a flashlight, a camera 193, etc. through different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K through the I2C interface. As a result, the processor 110 communicates with the touch sensor 180K through the I2C bus interface to implement the touch function of the terminal 100.

[0152] The I2S interface may be used for audio communication. In some embodiments, the processor 110 may include a plurality of groups of I2S buses. The processor 110 may be coupled to the audio module 170 through an I2S bus to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of responding to a call by using a Bluetooth (registered trademark) headset.

[0153] The PCM interface may also be used for audio communication by using analog signal sampling, quantization, and encoding. In some embodiments, the audio module 170 may be coupled to the wireless communication module 160 through a PCM bus interface. In some embodiments, alternatively, the audio module 170 may transmit an audio signal to the wireless communication module 160 through the PCM interface to implement the function of responding to a call by using a Bluetooth (registered trademark) headset. Both the I2S interface and the PCM interface may be used for audio communication.

[0154] The UART interface is a general-purpose serial data bus and is used for asynchronous communication. The bus may be a bidirectional communication bus. The bus converts between serial communication and parallel communication for the data to be transmitted. In some embodiments, the UART interface is typically used to connect the processor 110 to the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 may transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music by using a Bluetooth headset.

[0155] The MIPI interface may be used to connect the processor 110 to peripheral components, such as the display screen 194 or the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 communicates with the camera 193 through the CSI to implement the shooting function of the terminal 100. The processor 110 communicates with the display screen 194 through the DSI to implement the display function of the terminal 100.

[0156] The GPIO interface may be configured by using software. The GPIO interface may be configured using control signals or may be configured using data signals. In some embodiments, the GPIO interface may be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface may be further configured as an I2C interface, an I2S interface, a UART interface, MIPI, etc.

[0157] Specifically, the video acquired by the camera 193 (for example, including an image frame sequence including the first image and the second image in the present application) may be transferred to the processor 110 through the interface described above (for example, the CSI interface or the GPIO interface) that connects the camera 193 to the processor 110, although it is not limited thereto.

[0158] The processor 110 may obtain an instruction from the memory and, according to the obtained instruction, perform video processing (for example, hand shake correction or perspective distortion correction in the present application) on the video acquired by the camera 193 to obtain a processed image (for example, the third image and the fourth image in the present application).

[0159] The processor 110 may transfer the processed image to the display screen 194 through the interface described above (for example, the DSI interface or the GPIO interface) that connects the display screen 194 and the processor 110, although it is not limited thereto, and as a result, the display screen 194 may display the video.

[0160] The USB interface 130 is an interface that conforms to the use of the USB standard. Specifically, it may be a mini-USB interface, a micro-USB interface, a USB Type-C interface, etc. The USB interface 130 may be used to connect to a charger to charge the terminal 100, or may be used to transmit data between the terminal 100 and peripheral devices, or may be used to connect to a headset for playing audio through the headset. The interface may be further used to connect to another electronic device, such as an AR device.

[0161] It can be understood that the interface connection relationship between the modules shown in this embodiment of the present invention is only an example for explanation and does not constitute a limitation on the structure of the terminal 100. In some other embodiments of the present application, the terminal 100 may alternatively use an interface connection method different from the method in the foregoing embodiment, or may use a combination of multiple interface connection methods.

[0162] The charging management module 140 is configured to receive a charging input from a charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive a charging input from a wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 may receive a wireless charging input through the wireless charging coil of the terminal 100. The charging management module 140 supplies power to the electronic device by using the power management module 141 while charging the battery 142.

[0163] The power management module 141 is configured to be connected to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal storage 121, the display screen 194, the camera module 193, the wireless communication module 160, etc. The power management module 141 may be further configured to monitor parameters such as battery capacity, battery cycle count, and battery health status (electrical leakage or impedance). In some other embodiments, the power management module 141 may alternatively be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may alternatively be disposed in the same component.

[0164] The wireless communication function of the terminal 100 may be implemented by using the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, the baseband processor, etc.

[0165] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic wave signals. Each antenna in the terminal 100 may be used to cover one or more communication frequency bands. Different antennas may be further reused to improve antenna utilization. For example, the antenna 1 may be reused as a diversity antenna in a wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.

[0166] The mobile communication module 150 may provide a wireless communication solution including 2G / 3G / 4G / 5G applicable to the terminal 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 may receive electromagnetic waves through the antenna 1, perform processing, such as filtering or amplification, on the received electromagnetic waves, and transmit the electromagnetic waves to the modem processor for demodulation. The mobile communication module 150 may further amplify the signal modulated by the modem processor and convert the signal into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least a part of the functional modules of the mobile communication module 150 may be arranged in the processor 110. In some embodiments, at least a part of the functional modules of the mobile communication module 150 and at least a part of the modules of the processor 110 may be arranged in the same component.

[0167] The modem processor may include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Next, the demodulator transmits the low-frequency baseband signal obtained through demodulation to the baseband processor for processing. The baseband processor processes the low-frequency baseband signal, which is then transferred to the application processor. The application processor outputs a sound signal through an audio device (not limited to the loudspeaker 170A, receiver 170B, etc.) or displays an image or video by using the display screen 194. In some embodiments, the modem processor may be an independent component. In some other embodiments, the modem processor may be independent of the processor 110 and arranged together with the mobile communication module 150 or another functional module in the same device.

[0168] The wireless communication module 160 is applied to the terminal 100 and may provide a wireless communication solution including a wireless local area network (WLAN) (for example, a wireless fidelity (Wi-Fi (registered trademark)) network), Bluetooth (registered trademark) (Bluetooth (registered trademark), BT), a global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC) technology, infrared (IR) technology, and the like. The wireless communication module 160 may be one or more components that integrate at least one communication processor module. The wireless communication module 160 receives electromagnetic waves through the antenna 2, performs frequency modulation and filtering on the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 may further receive a signal to be transmitted from the processor 110, perform frequency modulation and amplification on the signal, and convert the signal into an electromagnetic wave for radiation through the antenna 2.

[0169] In some embodiments, the antenna 1 and the mobile communication module 150 on the terminal 100 are coupled, and the antenna 2 and the wireless communication module 160 in the terminal device are coupled. As a result, the terminal 100 can communicate with the network and other devices by using wireless communication technologies. The wireless communication technologies may include global system for mobile communication (GSM (registered trademark)), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA (registered trademark)), time-division-synchronous code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, IR, etc. GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), BeiDou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0170] The terminal 100 may implement a display function by using a GPU, a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing and is connected to the display screen 194 and the application processor. The GPU is configured to execute mathematical and geometric calculations and render images. The processor 110 may include one or more GPUs that execute program instructions for generating or changing display information. Specifically, one or more GPUs of the processor 110 may implement an image rendering task (for example, a rendering task related to an image that needs to be displayed in the present application, where the rendering result is transferred to an application processor or another display driver, and the application processor or another display driver triggers the display screen 194 to display a video).

[0171] The display screen 194 is configured to display images, videos, etc. The display screen 194 includes a display panel. The display panel may use a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini LED, a micro LED, a micro OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal 100 may include one or N display screens 194, where N is a positive integer greater than 1. The display screen 194 may display a target video in the embodiments of the present application. In one implementation, the terminal 100 may execute an application related to photography. When the terminal starts an application related to photography, the display screen 194 may display a photography interface. The photography interface may include a viewfinder frame, and the target video may be displayed in the viewfinder frame.

[0172] The terminal 100 may implement a photography function by using an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.

[0173] The ISP is configured to process the data fed back by the camera 193. For example, during shooting, the shutter is pressed, and light is transmitted through the lens to the photosensitive element of the camera. The optical signal is converted into an electrical signal, and the photosensitive element of the camera transmits the electrical signal to the ISP for processing and converts the electrical signal into a visible image. The ISP may further perform algorithm optimization for image noise, brightness, and color. The ISP may further optimize parameters such as exposure and color temperature in the shooting scenario. In some embodiments, the ISP may be disposed in the camera 193.

[0174] The camera 193 is configured to capture still images or videos. The optical image of the object is generated through the lens and projected onto the photosensitive element. The photosensitive element may be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) photoelectric transistor. The photosensitive element converts the optical signal into an electrical signal and then transmits the electrical signal to the ISP to convert the electrical signal into a digital image signal for processing. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard format, such as an RGB or YUV image signal. In some embodiments, the terminal 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0175] The DSP converts the digital image signal into a standard format, such as an RGB or YUV image signal, to obtain the original image (for example, the first image and the second image in the embodiments of the present application). The processor 110 may further perform image processing on the original image. The image processing includes, but is not limited to, shake correction, perspective distortion correction, optical distortion correction, and cropping to fit the size of the display screen 194. The processed image (for example, the third image and the fourth image in the embodiments of the present application) may be displayed in the viewfinder frame in the shooting interface displayed on the display screen 194.

[0176] In this embodiment of the present application, at least two cameras 193 may be present on the terminal 100. For example, two cameras are present, where one is a front camera and the other is a rear camera; for example, three cameras are present, where one is a front camera and the other two are rear cameras; or, for example, four cameras are present, where one is a front camera and the other three are rear cameras. It should be noted that the camera 193 may be one or more of a wide-angle camera, a main camera, or a telephoto camera.

[0177] For example, two cameras are present, where the front camera may be a wide-angle camera and the rear camera may be a main camera. In this case, the field of view of the image obtained by the rear camera is larger and there is more image information.

[0178] For example, three cameras are present, where the front camera may be a wide-angle camera and the rear cameras may be a wide-angle camera and a main camera.

[0179] For example, four cameras are present, where the front camera may be a wide-angle camera and the rear cameras may be a wide-angle camera, a main camera, and a telephoto camera.

[0180] The digital signal processor is configured to process digital signals and may process another digital signal in addition to the digital image signal. For example, when the terminal 100 selects a frequency, the digital signal processor is configured to perform a Fourier transform on the frequency energy.

[0181] The video codec is configured to compress or decompress digital video. The terminal 100 may support one or more types of video codecs. Therefore, the terminal 100 may play or record videos in multiple encoding formats, for example, Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-3, and MPEG-4 videos.

[0182] The NPU is a neural-network (NN) computing processor that quickly processes input information by referring to the structure of biological neural networks, for example, by referring to the transmission mode between human brain nerve cells, and may further continuously perform self-learning. Applications such as the intelligent cognition of the terminal 100, for example, image recognition, face recognition, voice recognition, and text understanding, may be implemented by using the NPU.

[0183] The external storage interface 120 may be used to connect to an external storage card, for example, a micro SD card, to expand the storage capacity of the terminal 100. The external storage card communicates with the processor 110 through the external storage interface 120 to implement a data storage function. For example, files such as music and videos are stored in the external storage card.

[0184] The internal storage 121 may be configured to store computer-executable program code. The executable program code includes instructions. The internal storage 121 may include a program storage area and a data storage area. The program storage area may store an operating system, applications required for at least one function (such as a sound playback function or an image playback function), etc. The data storage area may store data created during the use of the terminal 100 (such as audio data and an address book), etc. In addition, the internal storage 121 may include a high-speed random access memory and may further include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or a universal flash storage (UFS). The processor 110 executes instructions stored in the internal storage 121 and / or instructions stored in a memory disposed in the processor to implement various functional applications and data processing of the terminal 100.

[0185] The terminal 100 may implement an audio function, such as music playback or recording, by using an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, an application processor, etc.

[0186] The audio module 170 is configured to convert digital audio information into an analog audio signal for output and is further configured to convert an analog audio input into a digital audio signal. The audio module 170 may be further configured to encode and decode audio signals. In some embodiments, the audio module 170 may be disposed in the processor 110, or some of the functional modules of the audio module 170 may be disposed in the processor 110.

[0187] The loudspeaker 170A, also referred to as a "speaker", is configured to convert an electrical audio signal into a sound signal. The terminal 100 may receive music or a hands-free call by using the loudspeaker 170A.

[0188] The receiver 170B, also referred to as an "earpiece", is configured to convert an electrical audio signal into a sound signal. When a call or audio information is received by the terminal 100, the receiver 170B can be placed near a human ear to listen to the voice.

[0189] The microphone 170C, also referred to as a "mike" or "mic", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user may move their mouth near the microphone 170C to input a sound signal into the microphone 170C. At least one microphone 170C may be disposed in the terminal 100. In some other embodiments, two microphones 170C may be disposed in the terminal 100 to acquire a sound signal and further reduce noise. In some other embodiments, alternatively, three, four, or more microphones 170C may be disposed in the terminal 100 to acquire a sound signal, reduce noise, and identify a sound source, such as for implementing a directional sound recording function.

[0190] The headset jack 170D is configured to connect to a wired headset. The headset jack 170D may be a USB interface 130, or may be a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunication industry association (CTIA) of the USA standard interface.

[0191] The pressure sensor 180A is configured to detect a pressure signal and may convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194. There are multiple types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. The capacitive pressure sensor may include at least two parallel plates made of a conductive material. When pressure is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The terminal 100 determines the pressure intensity based on the change in capacitance. When a touch operation is performed on the display screen 194, the terminal 100 detects the intensity of the touch operation by using the pressure sensor 180A. The terminal 100 may calculate the touch position based on the detection signal from the pressure sensor 180A. In some embodiments, touch operations that are performed at the same touch position but have different touch operation intensities may correspond to different operation instructions. For example, when a touch operation with a touch operation intensity smaller than a first pressure threshold is performed on a message icon, an instruction for viewing an SMS message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold is performed on a message icon, an instruction for creating a new SMS message is executed.

[0192] The shooting interface displayed on the display screen 194 may include a first control and a second control. The first control is used to enable or disable the hand shake correction function, and the second control is used to enable or disable the perspective distortion correction function. For example, the user may perform an enabling operation on the display screen to enable the hand shake correction function. The enabling operation may be a click operation on the first control. The terminal 100 may determine, based on the detection signal of the pressure sensor 180A, that the position of the click on the display screen is the position of the first control, and then generate an operation instruction for enabling the hand shake correction function. In this way, the hand shake correction function is enabled according to the operation instruction for enabling the hand shake correction function. For example, the user may perform an enabling operation on the display screen to enable the perspective distortion correction function. The enabling operation may be a click operation on the second control. The terminal 100 may determine, based on the detection signal of the pressure sensor 180A, that the position of the click on the display screen is the position of the second control, and then generate an operation instruction for enabling the perspective distortion correction function. In this way, the perspective distortion correction function is enabled according to the operation instruction for enabling the perspective distortion correction function.

[0193] The gyroscope sensor 180B may be configured to determine the motion posture of the terminal 100. In some embodiments, the angular velocity of the terminal 100 around three axes (i.e., axes x, y, and z) may be determined by using the gyroscope sensor 180B. The gyroscope sensor 180B may be configured to implement hand shake correction during shooting. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle when the terminal 100 shakes, calculates the distance that the lens module needs to compensate based on the angle, and enables the lens to cancel out the shake of the terminal 100 through reverse motion, thereby implementing hand shake correction. The gyroscope sensor 180B may be further used in navigation and motion sensing game scenarios.

[0194] The barometric pressure sensor 180C is configured to measure barometric pressure. In some embodiments, the terminal 100 calculates altitude based on the value of the barometric pressure measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0195] The magnetic sensor 180D includes a Hall sensor. The terminal 100 may detect the opening and closing of the flip cover by using the magnetic sensor 180D. In some embodiments, when the terminal 100 is a folding phone, the terminal 100 may detect the opening and closing of the flip cover by using the magnetic sensor 180D. Further, based on the detected open or closed state of the flip cover, functions when the flip cover is open, such as automatic unlocking, are set.

[0196] The acceleration sensor 180E may detect the acceleration of the terminal 100 in various directions (usually on three axes). When the terminal 100 is stationary, the magnitude and direction of gravity may be detected. The acceleration sensor 180E may be further configured to identify the posture of the electronic device, for example, used for switching between landscape mode and portrait mode or for a pedometer.

[0197] The distance sensor 180F is configured to measure distance. The terminal 100 may measure distance using infrared or laser. In some embodiments, in a shooting scenario, the terminal 100 may measure distance by using the distance sensor 180F to implement quick focusing.

[0198] The optical proximity sensor 180G may include a light-emitting diode (LED) and an optical detector, such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The terminal 100 emits infrared light through the light-emitting diode. The terminal 100 detects infrared light reflected from an object in the vicinity by using the photodiode. If sufficient reflected light is detected, the terminal 100 may determine that there is an object near the terminal 100. If sufficient reflected light is not detected, the terminal 100 may determine that there is no object near the terminal 100. The terminal 100 may detect that the user is holding the terminal 100 near the ear when making a call by using the optical proximity sensor 180G so as to automatically execute screen-off for power saving. The optical proximity sensor 180G may be further used in the leather case mode or the pocket mode to automatically unlock or lock the screen.

[0199] The ambient light sensor 180L is configured to detect the ambient light luminance. The terminal 100 may adaptively adjust the luminance of the display screen 194 based on the detected ambient light luminance. The ambient light sensor 180L may be further configured to automatically adjust the white balance during photography. The ambient light sensor 180L may further cooperate with the optical proximity sensor 180G to detect whether the terminal 100 is in the pocket to avoid accidental touches.

[0200] The fingerprint sensor 180H is configured to acquire fingerprints. The terminal 100 may implement fingerprint-based unlocking, application lock access, fingerprint-based photography, fingerprint-based call response, etc. by using the characteristics of the acquired fingerprints.

[0201] The temperature sensor 180J is configured to detect temperature. In some embodiments, the terminal 100 executes a temperature processing policy based on the temperature detected by the temperature sensor 180J. For example, if the temperature reported by the temperature sensor 180J exceeds a threshold, the terminal 100 reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and provide thermal protection. In some other embodiments, if the temperature is lower than another threshold, the terminal 100 heats the battery 142 to prevent the terminal 100 from shutting down abnormally due to low temperature. In some other embodiments, if the temperature is lower than yet another threshold, the terminal 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown due to low temperature.

[0202] The touch sensor 180K may also be referred to as a "touch screen device". The touch sensor 180K may be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, which is also referred to as a "touch screen". The touch sensor 180K is configured to detect touch operations performed on or near the touch sensor. The touch sensor may transfer the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation may be provided through the display screen 194. In some other embodiments, the touch sensor 180K may alternatively be disposed at a position on the surface of the terminal 100 that is different from the position of the display screen 194.

[0203] The bone conduction sensor 180M may acquire vibration signals. In some embodiments, the bone conduction sensor 180M may acquire vibration signals of the vibrating bones in the vocal cord region of a human. The bone conduction sensor 180M may further come into contact with a human pulse and receive blood pressure and pulse signals. In some embodiments, alternatively, the bone conduction sensor 180M may be disposed on a headset to form a bone conduction headset. The audio module 170 may be for the vibrating bones in the vocal cord region to implement a voice function, and may acquire a voice signal through analysis based on the vibration signals acquired by the bone conduction sensor 180M. The application processor may analyze heart rate information based on the blood pressure and pulse signals acquired by the bone conduction sensor 180M to implement a heart rate detection function.

[0204] The button 190 includes a power button, a volume button, etc. The button 190 may be a mechanical button or a touch button. The terminal 100 may receive button inputs and generate key signal inputs related to user settings and function controls on the terminal 100.

[0205] The motor 191 may generate vibration prompts. The motor 191 may be configured to generate an incoming call vibration prompt and may be configured to provide touch vibration feedback. For example, touch operations performed for different applications (such as photo taking and audio playback) may correspond to different vibration feedback effects. For touch operations performed on different regions of the display screen 194, the motor 191 may also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, information reception, alarm clock, and game) may correspond to different vibration feedback effects. The touch vibration feedback effect may be further customized.

[0206] The indicator 192 may be an indicator light and may be configured to indicate a charging status and power change, and may be configured to indicate messages, missed calls, notifications, etc.

[0207] The SIM card interface 195 is used to connect a SIM card. The SIM card may be inserted into the SIM card interface 195 or removed from the SIM card interface 195 to implement contact with or separation from the terminal 100. The terminal 100 may support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 may support a nano SIM card, a micro SIM card, a SIM card, etc. A plurality of cards may be inserted into the same SIM card interface 195 simultaneously. The plurality of cards may be of the same type or different types. The SIM card interface 195 is compatible with different types of SIM cards. The SIM card interface 195 is also compatible with an external storage card. The terminal 100 interacts with a network by using the SIM card to implement functions such as calls and data communication. In some embodiments, the terminal 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card may be embedded in the terminal 100 and cannot be separated from the terminal 100.

[0208] The software system of the terminal 100 may use a hierarchical architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In this embodiment of the present invention, an Android (registered trademark) system having a hierarchical architecture is used as an example to illustrate the software structure of the terminal 100.

[0209] FIG. 2 is a block diagram of the software structure of the terminal 100 according to an embodiment of the present disclosure.

[0210] In a hierarchical architecture, software is divided into several layers, and each layer has a clear role and task. The layers communicate with each other through software interfaces. In some embodiments, the Android (registered trademark) system is divided into four layers: from top to bottom, an application layer, an application framework layer, an Android runtime and native libraries, and a kernel layer.

[0211] The application layer may include a series of application packages.

[0212] As shown in FIG. 2, the application packages may include applications such as a camera, a gallery, a calendar, a call, a map, navigation, WLAN, Bluetooth (registered trademark), music, video, and messages.

[0213] The application framework layer provides application programming interfaces (APIs) and programming frameworks to the applications in the application layer. The application framework layer includes several predefined functions.

[0214] As shown in FIG. 2, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.

[0215] The window manager is used to manage window programs. The window manager may obtain the size of the display screen, determine whether a status bar exists, lock the screen, take a screenshot, etc.

[0216] The content provider: stores and retrieves data and is used to enable the data to be accessed by an application. The data may include videos, images, audio, outgoing and incoming calls, browsing history and bookmarks, contacts, etc.

[0217] The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system may be used to build an application. The display interface may include one or more views. For example, a display interface including a message notification icon may include a text display view and a photo display view.

[0218] The phone manager is used to provide management of the communication functions of the terminal 100, such as management of call status (including answering, ending a call, etc.).

[0219] The resource manager provides an application with various resources such as localized strings, icons, photos, layout files, and video files.

[0220] The notification manager enables an application to display notification information in the status bar and may be used to send notification type messages. The displayed information may automatically disappear after a short pause without user interaction. For example, the notification manager is used to notify of download completion, provide message notifications, etc. The notification manager may alternatively be a notification that appears in the status bar at the top of the system in the form of a graph or scroll bar text, such as a notification of an application running in the background, or a notification that appears on the screen in the form of a dialog window. For example, text information is prompted in the status bar, a prompt tone is played, the electronic device vibrates, or an indicator light blinks.

[0221] The Android (registered trademark) runtime includes a core library and a virtual machine. The Android (registered trademark) runtime is responsible for the scheduling and management of the Android (registered trademark) system.

[0222] The core library includes two parts. One of these parts is the execution functions that need to be called in the Java (registered trademark) language, and the other part is the core library of Android (registered trademark).

[0223] The application layer and the application framework layer are executed on the virtual machine. The virtual machine executes Java (registered trademark) files as binary files in the application layer and the application framework layer. The virtual machine is configured to implement functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0224] The native library may include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (e.g., OpenGL ES), and a 2D graphics engine (e.g., SGL).

[0225] The surface manager manages the display subsystem and is used to provide the fusion of 2D and 3D layers for multiple applications.

[0226] The media library supports the playback and recording of multiple commonly used audio and video formats, still image files, etc. The media library may support multiple audio and video encoding formats, such as MPEG-4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0227] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, composition, layer processing, etc.

[0228] The 2D graphics engine is a drawing engine for 2D drawing.

[0229] The kernel layer is a layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.

[0230] An example of the operation process of the software and hardware on the terminal 100 will be described below with reference to a shooting scenario.

[0231] When the touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation to obtain the original input event (information, including, for example, the touch coordinates and time stamp of the touch operation). The original input event is stored in the kernel layer. The application framework layer obtains the original input event from the kernel layer and recognizes the control corresponding to the input event. For example, the touch operation is a touch-and-click operation, and the control corresponding to this click operation is the control of the icon of the camera application. The camera application calls the interface in the application framework layer to start the camera application, activates the camera driver by calling the kernel layer, and captures a still image or a video by using the camera 193. The captured video may be the first image and the second image in the embodiments of the present application.

[0232] For ease of understanding, the image processing method provided in an embodiment of the present application will be described in detail with reference to the accompanying drawings and application scenarios.

[0233] This embodiment of the present application may be applied to scenarios such as real-time shooting, video post-processing, and target tracking. Descriptions regarding each of the scenarios are provided below.

[0234] Example 1: Real-time shooting

[0235] In this embodiment of the present application, in the scenario of real-time shooting or recording of a terminal, the camera on the terminal may acquire a video stream in real time and display a preview photo generated based on the video stream acquired by the camera in a shooting interface.

[0236] The video stream acquired by the camera may include a first image and a second image. The first image and the second image are two original image frames acquired by the terminal and adjacent to each other in the time domain. Specifically, the video acquired by the terminal in real time is a sequence of original images sorted sequentially in the time domain. The second image is an original image adjacent to the first image in the time domain and after the first image. For example, the video acquired by the terminal in real time includes a 0th image frame, a 1st image frame, a 2nd image frame,..., and an Xth image frame sorted in the time domain. Here, the first image is the nth image frame, the second image is the (n + 1)th image frame, and n is an integer greater than or equal to 0 and less than X.

[0237] The first image and the second image are original images acquired by the camera on the terminal. The concept of the original image is explained below.

[0238] The terminal may open the shutter for shooting, and light may be transferred through the lens to the image sensor of the camera. The image sensor of the camera may convert the optical signal into an electrical signal and transfer the electrical signal to an image signal processor (ISP), a digital signal processor (DSP), etc. for processing. As a result, an image may be obtained through the conversion. The image may be referred to as the original image acquired by the camera. The first image and the second image described in this embodiment of the present application may be the aforementioned original images. According to the image processing method provided in this embodiment of the present application, image processing may be performed on the original image based on the original image in order to obtain the processed target images (the third image and the fourth image). As a result, a preview photo including the third image and the fourth image may be displayed in the viewfinder frame 402. It should be understood that the first image and the second image may alternatively be images obtained by cropping the images acquired after being processed by the ISP and the DSP. Cropping may be performed to fit the size of the display screen of the terminal or the size of the viewfinder frame in the shooting interface.

[0239] In this embodiment of the present application, shake correction and perspective distortion correction may be performed on the video stream acquired by the camera in real time. The first image and the second image in this embodiment of the present application are used as an example to illustrate how to perform shake correction and perspective distortion correction on the first image and the second image.

[0240] In some scenarios, the user may hold the terminal device to shoot the target area. The target area may include a moving object, or the target area may be a stationary area, that is, it may not include a moving object.

[0241] When hand shake correction and perspective distortion correction are performed on a video stream, a jello effect (or jello phenomenon) may occur in the preview photo displayed in the viewfinder frame. The jello phenomenon (jello effect) means that the photo is transformed and changed like jello. The cause of the jello effect is explained below.

[0242] Perspective distortion and correction of perspective distortion of an image are first introduced.

[0243] When the camera is in the wide-angle shooting mode or the main camera shooting mode, the acquired original image has obvious perspective distortion. Specifically, perspective distortion (also referred to as 3D distortion) is the transformation (e.g., horizontal elongation, radial elongation, and combinations thereof) in 3D object imaging caused by different magnifications due to the difference in the depth of field of the subject for shooting when the 3D shape in space is mapped onto the image plane. Therefore, perspective distortion correction needs to be performed on the image.

[0244] It should be understood that the transformation in 3D object imaging can be understood as an image transformation caused by the transformation of the object in the image compared to how the actual object appears to the human eye.

[0245] In an image with perspective distortion, compared to an object closer to the center of the image, an object farther from the center of the image has more obvious perspective distortion, that is, it has a larger imaging transformation in length (a larger amplitude of horizontal elongation, radial elongation, and combinations thereof). Therefore, in order to obtain an image without the problem of perspective distortion, during perspective distortion correction of the image, for position points or pixels at different distances from the center of the image, the shift that needs to be corrected is different, and for position points or pixels farther from the center of the image, the shift that needs to be corrected is larger.

[0246] Please refer to FIG. 4a. FIG. 4a is a schematic diagram of an image with perspective distortion. As shown in FIG. 4a, a person in the edge region of the image has a specific 3D transformation.

[0247] In some scenarios, shake may occur in the process of the user holding the terminal device for shooting.

[0248] When the user holds the terminal for shooting, since the user cannot stably maintain the posture of holding the terminal, the terminal shakes in the direction of the image plane (that is, the terminal has large posture changes within a very short time). The image plane may refer to the plane where the imaging plane of the camera is located, specifically, it may be the plane where the photosensitive surface of the photosensitive element of the camera is located. From the user's perspective, the image plane may be basically the same as or exactly the same as the plane where the display screen is located. The shake in the direction of the image plane may include shakes in each direction on the image plane, for example, left, right, up, down, upper left, lower left, upper right, and lower right shakes.

[0249] When the user holds the terminal to shoot the target area, if the user holds the terminal and has a shake in direction A, in the image content that is in the original image acquired by the camera on the terminal and includes the target area, an offset in the direction opposite to direction A is included. For example, when the user holds the terminal and has a shake in the upper left direction, in the image content that is in the original image acquired by the camera on the terminal and includes the target area, an offset in the direction opposite to the upper left direction is included.

[0250] It should be understood that the shake in direction A that occurs when the user holds the terminal can be understood as the terminal shaking in direction A from the perspective of the user holding the terminal.

[0251] Please refer to FIGS. 4b-1 to 4b-4. FIGS. 4b-1 to 4b-4 are schematic diagrams of a scenario where the terminal shakes. As shown in FIGS. 4b-1 to 4b-4, at time point A1, the user is holding the terminal device to photograph object 1. In this case, object 1 may be displayed in the photographing interface. Between time points A1 and A2, the user holds the terminal device and has a shake in the lower left direction. In this case, object 1 in the original image acquired by the camera has an offset in the upper right direction.

[0252] When the user holds the terminal to photograph the target area, since the user's posture when holding the terminal is unstable, the target area has a large offset between adjacent image frames acquired by the camera (in other words, there is a large content misalignment between two adjacent image frames acquired by the camera). When the terminal device shakes, usually, there is a large content misalignment between two adjacent image frames acquired by the camera, and as a result, the video acquired by the camera includes blurring. Compared with using a wide-angle camera, when the terminal shakes, the misalignment between the images acquired by using the main camera becomes more serious. It should be understood that compared with using the main camera, when the terminal shakes, the misalignment between the images acquired by using the telephoto camera becomes more serious.

[0253] During shake correction, after two adjacent image frames are aligned, the edge portions are cropped to obtain a clear image (the output of shake correction or the shake correction viewfinder frame, for example, in this embodiment of the present application, it may be referred to as the first cropped area and the third cropped area).

[0254] Performing hand shake correction is to eliminate the shake in the acquired video caused by the change in the user's posture when holding the terminal device. Specifically, performing hand shake correction is to crop the original image acquired by the camera in order to obtain the output in the viewfinder frame. The output in the viewfinder frame is the output of hand shake correction. When the terminal has a large posture change within a very short time, there is a large offset in the photo of the viewfinder frame of the latter image frame from the photo of the former image frame. For example, when the terminal acquires the (n + 1)-th frame as compared with the case where the n-th frame is acquired, there is a shake in direction A within the direction of the image plane (it should be understood that the shake in direction A may be the shake of the main optical axis of the camera on the terminal towards direction A or the shake of the position of the optical center of the camera on the terminal towards direction A within the direction of the image plane). In this case, the photo in the viewfinder frame of the (n + 1)-th image frame is shifted in the direction opposite to direction A in the n-th image frame.

[0255] Figures 4b-1 to 4b-4 are used as an example. Between time points A1 and A2, the user holds the terminal device and has a shake in the lower left direction. In this case, object 1 in the original image acquired by the camera has an offset towards the upper right direction, and thus the cropped area (the dashed box shown in Figures 4b-1 to 4b-4) that needs to be cropped for hand shake correction has an offset towards the upper right direction accordingly.

[0256] For hand shake correction, it is necessary to determine the offset shift (including the offset direction and offset distance) of the photo in the viewfinder frame of the latter original image frame from the photo of the former original image frame. In one implementation, the offset shift may be determined based on the shake that occurs in the process where the terminal acquires two adjacent original image frames. Specifically, the offset shift may be a posture change that occurs in the process where the terminal acquires two adjacent original image frames. The posture change may be acquired based on the photographed posture corresponding to the viewfinder frame in the immediately previous original image frame and the posture when the terminal photographs the current original image frame.

[0257] After hand shake correction, when the user holds the terminal device and has a large shake, the photos of adjacent original image frames have a large difference, but the photos in the viewfinder frames acquired after the adjacent original image frames are cropped have no large difference or no difference. As a result, the photos displayed in the viewfinder frame in the shooting interface do not show a large shake or blur.

[0258] Referring to FIG. 4c, a schematic diagram of the detailed process of hand shake correction is provided below. Please refer to FIG. 4c. For the nth original image frame acquired by the camera, the posture parameters of the handheld terminal when the nth image frame is photographed may be acquired. The posture parameters of the handheld terminal may indicate the posture of the terminal when the nth frame is acquired, and the posture parameters of the handheld terminal may be acquired based on the information acquired by the sensors on the terminal. The sensors on the terminal may be, but are not limited to, gyroscopes, and the information acquired may be, but is not limited to, the rotational angular velocity of the terminal on the x-axis, the rotational angular velocity of the terminal on the y-axis, and the rotational angular velocity of the terminal on the z-axis.

[0259] In addition, the stable pose parameters of the immediately preceding image frame, i.e., the pose parameters of the (n - 1)-th image frame, may be further obtained. The stable pose parameters of the immediately preceding image frame may refer to the shooting pose corresponding to the viewfinder frame in the (n - 1)-th image frame. The shake direction and amplitude of the viewfinder frame in the original image on the image plane (the shake direction and amplitude may also be referred to as shake shift) may be calculated based on the stable pose parameters of the immediately preceding image frame and the pose parameters of the handheld terminal when the n-th image frame is obtained. The size of the viewfinder frame is preset. When n is 0, i.e., when the first image frame is obtained, the position of the viewfinder frame may be the center of the original image. After the shake direction and amplitude of the viewfinder frame in the original image on the image plane are obtained, the position of the viewfinder frame in the n-th original image frame may be determined, and it may be determined whether the boundary of the viewfinder frame exceeds the range of the input image (i.e., the n-th original image frame). If the boundary of the viewfinder frame exceeds the range of the input image due to an overly large shake amplitude, the stable pose parameters of the n-th image frame are adjusted so that the boundary of the viewfinder frame remains within the boundary of the input original image. As shown in FIGS. 4b-1 to 4b-4, if the boundary of the viewfinder frame exceeds the range of the input image in the upper right direction, the stable pose parameters of the n-th image frame are adjusted so that the position of the viewfinder frame moves toward the lower left, and the boundary of the viewfinder frame may be within the boundary of the input original image. Further, a hand shake correction index table is output. The output hand shake correction index table specifically shows the pose correction shift of each pixel or position point in the image after the shake is removed, and the output of the viewfinder frame may be determined based on the output hand shake correction index table and the n-th original image frame. When the shake amplitude is not large or no shake occurs, and thus the boundary of the viewfinder frame does not exceed the range of the input image, the hand shake correction index table may be directly output.

[0260] Similarly, when shake correction is performed on the (n + 1)-th image frame, the pose parameters of the n-th image frame may be obtained. The pose parameters of the n-th image frame may indicate the shooting pose corresponding to the viewfinder frame in the n-th image frame. The shake direction and amplitude of the viewfinder frame in the original image on the image plane (which may also be referred to as shake shift) may be calculated based on the stable pose parameters of the n-th image frame and the pose parameters of the handheld terminal when the (n + 1)-th image frame is obtained. The size of the viewfinder frame is preset. Next, the position of the viewfinder frame in the (n + 1)-th original image frame may be determined, and it may be determined whether the boundary of the viewfinder frame exceeds the range of the input image (i.e., the (n + 1)-th original image frame). If the boundary of the viewfinder frame exceeds the range of the input image due to an overly large shake amplitude, the stable pose parameters of the (n + 1)-th image frame are adjusted so that the boundary of the viewfinder frame remains within the boundary of the input original image. Further, a shake correction index table is output. If the shake amplitude is not large or no shake occurs, and thus the boundary of the viewfinder frame does not exceed the range of the input image, the shake correction index table may be directly output.

[0261] To more clearly explain the above-described process of shake correction, an example where n is 0, i.e., the first image frame is obtained, is used to explain the process of shake correction.

[0262] Refer to FIG. 4d. FIG. 4d is a schematic diagram of a process of hand shake correction according to an embodiment of the present application. As shown in FIG. 4d, the central portion of the image of the 0th frame is cropped as the output of the hand shake correction of the 0th frame (i.e., the output of the viewfinder frame of the 0th frame). Next, since the user has a shake when holding the terminal and the posture of the terminal changes, the captured video photo has a shake. The boundary of the viewfinder frame of the 1st frame may be calculated based on the posture parameter of the handheld terminal of the 1st frame and the stable posture parameter of the posture in the 0th image frame. As shown in FIG. 4d, it is determined that the position of the viewfinder frame in the 1st original image frame has a rightward offset. Assume that there is a very large shake when the viewfinder frame of the 1st frame is output and the 2nd image frame is captured. If the boundary of the viewfinder frame of the 2nd frame calculated based on the posture parameter of the handheld terminal of the 2nd frame and the posture parameter of the 1st image frame after the posture is corrected exceeds the boundary of the 2nd input original image frame, the stable posture parameter is adjusted. As a result, in order to obtain the output of the viewfinder frame of the 2nd original image frame, the boundary of the viewfinder frame of the 2nd original image frame remains within the boundary of the 2nd original image frame.

[0263] In some shooting scenarios, performing hand shake correction and perspective distortion correction on the acquired original image may cause a Jello effect in the photo. The reason is explained below.

[0264] In some existing implementations, in order to perform both shake correction and perspective distortion correction on an image simultaneously, the shake correction viewfinder frame needs to be confirmed in the original image, and perspective distortion correction needs to be performed on all regions of the original image. In this case, this is equivalent to perspective distortion correction being performed on the image region in the shake correction viewfinder frame and the image region in the shake correction viewfinder frame being output. However, the aforementioned method has the following problems: When the terminal device shakes, poor content alignment usually occurs between two adjacent original image frames acquired by the camera. Therefore, the positions of the image regions that will be output after shake correction is performed on the two image frames are different in the original image. That is, the distances between the centers of the image regions that will be output after shake correction is performed on the two image frames and the original image are different. As a result, the shifts for perspective distortion correction for the position points or pixels in the two image frames that will be output are different. The outputs of shake correction for two adjacent frames are images containing the same or basically the same content, but the degrees of transformation of the same object in the images are different. That is, the frame-to-frame consistency is lost, resulting in the Jello phenomenon and a decrease in the quality of the video output.

[0265] More specifically, the cause of the Jello phenomenon can be shown in FIG. 4e. There is an obvious shake between the first frame and the second frame, and the shake correction viewfinder frame moves from the center of the photo to the edge of the photo. The content within the smaller frame is the output of shake correction, and the larger frame is the original image. When the position of the content represented by the checkerboard pattern in the original image changes, the change in the interval between the squares in the former frame and the latter frame is different due to different shifts for perspective distortion correction. As a result, there is an obvious elongation of the squares (i.e., the Jello phenomenon) in the outputs of shake correction for the former frame and the latter frame.

[0266] One embodiment of the present application provides an image processing method. As a result, when blur correction and perspective distortion correction are simultaneously performed on an image, frame - to - frame consistency can be guaranteed, the jello phenomenon between images can be avoided, or the degree of conversion caused by the jello phenomenon between images can be reduced.

[0267] How to guarantee frame - to - frame consistency when blur correction and perspective distortion correction are simultaneously performed on an image will be described in detail below. This embodiment of the present application is described by using an example in which blur correction and perspective distortion correction are performed on a first image and a second image. The first image and the second image are two original image frames acquired by a terminal and adjacent to each other in the time domain. Specifically, a video acquired by a terminal in real time is a sequence of original images sorted sequentially in the time domain. The second image is an original image adjacent to the first image and after the first image in the time domain. For example, a video acquired by a terminal in real time includes a 0th image frame, a 1st image frame, a 2nd image frame,..., and an Xth image frame sorted in the time domain. Here, the first image is the nth image frame, the second image is the (n + 1)th image frame, and n is an integer greater than or equal to 0 and less than X.

[0268] In this embodiment of the present application, blur correction may be performed on the acquired first image to obtain a first cropped region. The first image is one original image frame of a video acquired by a camera on a terminal in real time. For the process of acquiring an original image by using a camera, refer to the description in the foregoing embodiment. Details will not be described again in this specification.

[0269] In one implementation, the first image is the first original image frame of the video acquired by the camera on the terminal in real time, and the first cropped area is in the central area of the first image. For details, refer to the description of how to perform shake correction on the nth frame in the foregoing embodiments. The details will not be described again in this specification.

[0270] In one implementation, the first image is the original image after the first original image frame of the video acquired by the camera on the terminal in real time, and the first cropped area is determined based on the shake that occurs in the process where the terminal acquires the first image, the original image adjacent to the first image, and the original image before the first image. For details, refer to the description of how to perform shake correction on the nth (n is not equal to 0) frame in the foregoing embodiments. The details will not be described again in this specification.

[0271] In one implementation, the shake correction may be performed on the first image by default. By default, it means that the shake correction is directly performed on the acquired original image without the need to enable the shake correction function. For example, the shake correction is performed on the acquired video when the shooting interface is opened. The shake correction function may be enabled based on the user's enabling operation. Alternatively, whether to enable the shake correction function is determined by analyzing the shooting parameters of the current shooting scenario and the content of the acquired video.

[0272] In one implementation, it may be detected that the terminal meets the shake correction condition, and the shake correction function is enabled to perform shake correction on the first image.

[0273] In a specific implementation process, detecting that the terminal meets the shake correction condition includes, but is not limited to, one or more of the following cases:

[0274] Case 1: It is detected that the zoom ratio of the terminal is higher than a second preset threshold value.

[0275] In one implementation, whether to enable the shake correction function may be determined based on the zoom ratio of the terminal. When the zoom ratio of the terminal is small (for example, within the full or partial zoom ratio range of the wide-angle shooting mode), the size of the physical area corresponding to the photo displayed in the viewfinder frame in the shooting interface is already very small. If further shake correction is performed, that is, if the image is further cropped, the size of the physical area corresponding to the photo displayed in the viewfinder frame will become smaller. Therefore, the shake correction function may be enabled only when the zoom ratio of the terminal is higher than a specific preset threshold value.

[0276] Specifically, before shake correction is performed on the first image, it is necessary to detect that the zoom ratio of the terminal is higher than or equal to a second preset threshold value. The second preset threshold value is the minimum zoom ratio for enabling the shake correction function. When the zoom ratio of the terminal ranges from a to 1, the terminal acquires a video stream in real time by using the front wide-angle camera or the rear wide-angle camera, where a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9. In this case, the second preset threshold value may be a zoom ratio value that is higher than or equal to a and lower than or equal to 1. For example, the second preset threshold value may be a. In this case, the shake correction function is not enabled only when the zoom ratio of the terminal is lower than a. For example, the second preset threshold value may be 0.6, 0.7, or 0.8. This is not limited in this specification.

[0277] Case 2: A second activation operation of the user for enabling the shake correction function is detected.

[0278] In one implementation, a control (referred to as the second control in this embodiment of the present application) used to instruct enabling or disabling of the hand shake correction function may be included within the shooting interface. The user may trigger enabling of the hand shake correction function through a second operation on the second control, i.e., trigger enabling of the hand shake correction function through the second operation on the second control. In this way, the terminal may detect a second enabling operation of the user to enable the hand shake correction function, and the second enabling operation includes the second operation on the second control in the shooting interface on the terminal.

[0279] In one implementation, the control (referred to as the second control in this embodiment of the present application) used to instruct enabling or disabling of the hand shake correction function is displayed in the shooting interface only when it is detected that the zoom ratio of the terminal is higher than or equal to a second preset threshold. In one implementation, optical distortion correction may be further performed on the first cropped region to obtain a corrected first cropped region.

[0280] In one implementation, perspective distortion correction may be performed on the first cropped region by default. By default means that perspective distortion correction is directly performed on the first cropped region without the need to enable perspective distortion correction. For example, perspective distortion correction is performed on the first cropped region obtained after hand shake correction when the shooting interface is opened. Perspective distortion correction may be enabled based on a user enabling operation. Alternatively, whether to enable perspective distortion correction is determined by analyzing the shooting parameters of the current shooting scenario and the content of the acquired video.

[0281] In one implementation, it may be detected that the terminal satisfies the distortion correction condition, and the distortion correction function is enabled to perform perspective distortion correction on the first cropped region.

[0282] It should be understood that the point in time when the terminal detects that it satisfies the distortion correction condition and an action to activate the distortion correction function is executed may be before the stage of performing shake correction on the acquired first image, or after the shake correction has been performed on the acquired first image and before the perspective distortion correction is performed on the first cropped area. This is not limited in the present application.

[0283] In a specific implementation process, detecting that the terminal satisfies the distortion correction condition includes, but is not limited to, one or more of the following cases:

[0284] Case 1: It is detected that the zoom ratio of the terminal is lower than a first preset threshold.

[0285] In one implementation, whether to activate the perspective distortion correction may be determined based on the zoom ratio of the terminal. When the zoom ratio of the terminal is large (for example, within the full or partial zoom ratio range of the telephoto shooting mode, or within the full or partial zoom ratio range of the medium-focus shooting mode), the degree of perspective distortion in the photo of the image acquired by the terminal is low. The fact that the degree of perspective distortion is low can be understood as that almost no perspective distortion in the photo of the image acquired by the terminal can be identified by the human eye. Therefore, when the zoom ratio of the terminal is large, the perspective distortion correction does not need to be activated.

[0286] Specifically, before the perspective distortion correction is performed on the first cropped area by default, it is necessary to detect that the zoom ratio of the terminal is lower than the first preset threshold. The first preset threshold is the maximum zoom ratio value for activating the perspective distortion correction. In other words, when the zoom ratio of the terminal is lower than the first preset threshold, the perspective distortion correction can be activated.

[0287] When the zoom ratio of the terminal ranges from a to 1, the terminal acquires a video stream in real time by using the front wide-angle camera or the rear wide-angle camera. Here, the zoom ratio corresponding to the image captured by the wide-angle camera may range from a to 1. The minimum value a may be the fixed zoom ratio of the wide-angle camera, and the value of a may range from 0.5 to 0.9. For example, a may be, but is not limited to, 0.5, 0.6, 0.7, 0.8, or 0.9. It should be understood that the wide-angle camera may include the front wide-angle camera and the rear wide-angle camera. The range of the zoom ratio corresponding to the front wide-angle camera and the rear wide-angle camera may be the same or different. For example, the zoom ratio corresponding to the front wide-angle camera may range from a1 to 1. Here, the minimum value a1 may be the fixed zoom ratio of the front wide-angle camera, and the value of a1 may range from 0.5 to 0.9. For example, a1 may be, but is not limited to, 0.5, 0.6, 0.7, 0.8, or 0.9. For example, the zoom ratio corresponding to the rear wide-angle camera may range from a2 to 1. Here, the minimum value a2 may be the fixed zoom ratio of the rear wide-angle camera, and the value of a2 may range from 0.5 to 0.9. For example, a2 may be, but is not limited to, 0.5, 0.6, 0.7, 0.8, or 0.9. When the terminal uses the front wide-angle camera or the rear wide-angle camera to acquire a video stream in real time, the first preset threshold is higher than 1. For example, the first preset threshold may be 2, 3, 4, 5, 6, 7, or 8. This is not limited in this specification.

[0288] When the zoom ratio of the terminal ranges from 1 to b, the terminal acquires a video stream in real time by using the rear main camera. When a rear telephoto camera is further integrated on the terminal, b may be the fixed zoom ratio of the rear telephoto camera, and the value of b may range from 3 to 15. For example, b may be, but is not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. When a rear telephoto camera is not integrated on the terminal, the value of b may be the maximum zoom of the terminal, and the value of b may range from 3 to 15. For example, b may be, but is not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. When the terminal acquires a video stream in real time by using the rear main camera, the first preset threshold is from 1 to 15. For example, the first preset threshold may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. This is not limited in this specification.

[0289] It should be understood that the range of the first preset threshold (1 to 15) may include the zoom ratios of the two endpoints (1 and 15), may not include the zoom ratios of the two endpoints (1 and 15), or may include one of the zoom ratios of the two endpoints.

[0290] For example, the range of the first preset threshold may be the interval [1, 15].

[0291] For example, the range of the first preset threshold may be the interval (1, 15].

[0292] For example, the range of the first preset threshold may be the interval [1, 15).

[0293] For example, the range of the first preset threshold may be the interval (1, 15).

[0294] Case 2: A first activation operation of the user to activate the perspective distortion correction function is detected.

[0295] In one implementation, a control (referred to as the first control in this embodiment of the present application) used to instruct enabling or disabling of perspective distortion correction may be included within a shooting interface. The user may trigger enabling perspective distortion correction through a first operation on the first control, that is, trigger enabling perspective distortion correction through the first operation on the first control. In this way, the terminal may detect a first enabling operation of the user for enabling perspective distortion correction, and the first enabling operation includes the first operation on the first control in the shooting interface on the terminal.

[0296] In one implementation, a control (referred to as the first control in this embodiment of the present application) used to instruct enabling or disabling of perspective distortion correction is displayed in the shooting interface only when it is detected that the zoom ratio of the terminal is lower than a first preset threshold.

[0297] Case 3: To determine to enable perspective distortion correction, a human face is identified in the shooting scenario; a human face is identified in the shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or a human face is identified in the shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio.

[0298] When a human face exists in a shooting scenario, the degree of conversion of the human face due to perspective distortion correction is more visually obvious, and when the distance between the human face and the terminal is smaller or the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher, the size of the human face in the image is larger. Therefore, it should be understood that the degree of conversion of the human face due to perspective distortion correction is even more visually obvious. Therefore, perspective distortion correction needs to be enabled in the above-described scenario. In this embodiment, whether to enable perspective distortion correction is determined by determining the above-described conditions related to the human face in the shooting scenario. As a result, the shooting scenario in which perspective distortion occurs can be accurately determined, and perspective distortion correction is executed for the shooting scenario in which perspective distortion occurs. In a shooting scenario where no perspective distortion occurs, perspective distortion correction is not executed. In this way, the image signal is accurately processed and power consumption is reduced.

[0299] In the above-described method, when the hand shake correction function is enabled and the perspective distortion correction function is enabled, the following steps 301 to 306 may be executed. Specifically, please refer to FIG. 3. FIG. 3 is a schematic diagram of an embodiment of an image processing method according to an embodiment of the present application. As shown in FIG. 3, the image processing method provided in the present application includes the following steps.

[0300] 301: Perform hand shake correction on the acquired first image to obtain a first cropped area.

[0301] In this embodiment of the present application, hand shake correction may be performed on the acquired first image to obtain a first cropped area. The first image is one original image frame of the video acquired by the camera on the terminal in real time. For the process of acquiring the original image by using the camera, please refer to the description in the above-described embodiment. Details will not be described again in this specification.

[0302] In one implementation, the first image is the first original image frame of the video acquired by the camera on the terminal in real time, and the first cropped region is in the central region of the first image. For details, refer to the description of how to perform shake correction on the 0th frame in the foregoing embodiment. The details will not be described again in this specification.

[0303] In one implementation, the first image is the original image after the first original image frame of the video acquired by the camera on the terminal in real time, and the first cropped region is determined based on the shake that occurs in the process where the terminal acquires the first image, the original image adjacent to the first image, and the original image before the first image. For details, refer to the description of how to perform shake correction on the nth (n is not equal to 0) frame in the foregoing embodiment. The details will not be described again in this specification.

[0304] In one implementation, optical distortion correction may be further performed on the first cropped region to obtain a corrected first cropped region.

[0305] 302: Perform perspective distortion correction on the first cropped region to obtain a third cropped region.

[0306] In this embodiment of the present application, after the first cropped region is obtained, perspective distortion correction may be performed on the first cropped region to obtain a third cropped region.

[0307] The shooting scenario can be understood as an image acquired before the camera acquires the first image, or when the first image is the first image frame of the video acquired by the camera, it should be understood that the shooting scenario may be the first image or an image close to the first image in the time domain. Whether there is a human face in the shooting scenario, whether the distance between the human face and the terminal is smaller than the preset value, and the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario may be determined by using a neural network or in another way. The preset value may range from 0 m to 10 m. For example, the preset value may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. This is not limited in this specification. The ratio of the pixels may range from 30% to 100%. For example, the ratio of the pixels may be 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100%. This is not limited in this specification.

[0308] How to enable the perspective distortion correction is described above. How to perform the perspective distortion correction is described below.

[0309] Unlike performing perspective distortion correction for all regions of the original image according to existing implementations, in this embodiment of the present application, perspective distortion correction is performed on the first cropped region. When the terminal device shakes, content misalignment usually occurs between two adjacent original image frames acquired by the camera. Therefore, the positions of the cropped regions that will be output after performing shake correction on the two image frames are different in the original image, that is, the distances between the centers of the cropped regions and the original image that will be output after performing shake correction on the two image frames are different. However, the sizes of the cropped regions that will be output after performing shake correction on the two frames are the same. Therefore, when perspective distortion correction is performed on the cropped regions obtained through shake correction, the degree of perspective distortion correction performed on adjacent frames is the same (because the distances between each sub-region in each of the cropped regions output through shake correction for the two frames and the center point of the cropped region output through shake correction are the same). In addition, the cropped regions output through shake correction for two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object in the images is the same, that is, frame-to-frame consistency is maintained. Therefore, no obvious Jello phenomenon occurs, thereby improving the quality of the video output. Frame-to-frame consistency in this specification may also be referred to as time-domain consistency, and it should be understood that it indicates that the processing results for regions having the same content in adjacent frames are the same.

[0310] Therefore, in this embodiment of the present application, perspective distortion correction is not performed on the first image, but rather perspective distortion correction is performed on the first cropped region obtained through shake correction, and a third cropped region is obtained. The same process is further performed for a frame (second image) adjacent to the first image, that is, perspective distortion correction is performed on the second cropped region obtained through shake correction, and a fourth cropped region is obtained. In this way, the degree of conversion correction of the same object in the third cropped region and the fourth cropped region is the same. As a result, frame-to-frame consistency is maintained, and no obvious Jello phenomenon occurs, thereby improving the quality of the video output.

[0311] In this embodiment of the present application, the same target object (for example, a human object) may be included within the cropped regions of the former frame and the latter frame. Since the object for perspective distortion correction is the cropped region obtained through shake correction, for the target object, the difference in the degree of conversion of the target object in the former frame and the latter frame due to perspective distortion correction is very small. The fact that the difference is very small can be understood as it being difficult for the human naked eye to distinguish the shape difference, or it being difficult for the human naked eye to detect the Jello phenomenon between the former frame and the latter frame.

[0312] Specifically, refer to FIG. 22a. An obvious Jello phenomenon exists in the output obtained after perspective distortion correction is performed on all regions of the original image, and no obvious Jello phenomenon exists in the output obtained after perspective distortion correction is performed on the cropped region obtained through shake correction.

[0313] In one implementation, optical distortion correction may be further performed on the third cropped region, and a corrected third cropped region may be obtained.

[0314] 303: Obtain a third image based on the first image and the third cropped region.

[0315] In this embodiment of the present application, the third cropped region may indicate the output of perspective distortion correction. Specifically, the third cropped region may indicate how to crop the first image and the shift of pixels or position points in order to obtain the image region that needs to be output.

[0316] In one implementation, the third cropped region may indicate the pixels or position points in the first image that need to be mapped or selected after shake correction and perspective distortion correction are performed, and the required shift of the pixels or position points in the first image that need to be mapped or selected.

[0317] Specifically, the third cropped region may include the mapping relationship between each pixel or position point in the output of perspective distortion correction and the pixels or position points in the first image. In this way, the third image may be obtained based on the first image and the third cropped region, and the third image may be displayed in the viewfinder frame in the shooting interface. The third image may be obtained based on the first image and the third cropped region through a warp operation, where it should be understood that the warp operation refers to an affine transformation on the image.

[0318] 304: Perform shake correction on the obtained second image to obtain a second cropped region, where the second cropped region is related to the first cropped region and shake information, and the shake information indicates the shake that occurs in the process of the terminal obtaining the first image and the second image.

[0319] The shake of the terminal indicated by the shake information may be a shake of the terminal that occurs within the period from time point T1 to time point T2. T1 may be the time point when the terminal acquires the first image, or a time point having a specific deviation from the time point when the terminal acquires the first image. T2 is the time point when the terminal acquires the second image, or a time point having a specific deviation from the time point when the terminal acquires the second image.

[0320] Specifically, the shake of the terminal indicated by the shake information may be a shake of the terminal that occurs within the period from a time point before the time point when the terminal acquires the first image to the time point when the terminal acquires the second image, or a shake of the terminal that occurs within the period from a time point after (and before the time point when the terminal acquires the second image) the time point when the terminal acquires the first image to the time point when the terminal acquires the second image, or a shake of the terminal that occurs within the period from a time point after (and before the time point when the terminal acquires the second image) the time point when the terminal acquires the first image to a time point after the time point when the terminal acquires the second image, or a shake of the terminal that occurs within the period from a time point after the time point when the terminal acquires the first image to a time point before the time point when the terminal acquires the second image, or a shake of the terminal that occurs within the period from the time point when the terminal acquires the first image to a time point before the time point when the terminal acquires the second image, or a shake of the terminal that occurs within the period from the time point when the terminal acquires the first image to a time point after the time point when the terminal acquires the second image. In a possible implementation, the time point before the time point when the terminal acquires the first image may be the time point when the terminal acquires an image before acquiring the first image, and the time point after the time point when the terminal acquires the second image may be the time point when the terminal acquires an image after acquiring the second image.

[0321] In one implementation, the direction of the offset of the position of the second cropped region in the second image with respect to the position of the first cropped region in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal acquiring the first image and the second image.

[0322] For further explanation regarding step 304, please refer to the description of how to perform shake correction on the acquired first image in the foregoing embodiment. Details will not be described again in this specification.

[0323] It should be understood that there is no strict time sequence limit between step 304, step 302, and step 303. Step 304 may be executed before step 302 and after step 301, or step 304 may be executed before step 303 and after step 302, or step 304 may be executed after step 303. This is not limited in this application.

[0324] 305: Perform perspective distortion correction on the second cropped region to obtain a fourth cropped region.

[0325] In this embodiment of the present application, distortion correction may be performed on the second cropped region to obtain a second mapping relationship. The second mapping relationship includes the mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second cropped region; the fourth cropped region is determined based on the second cropped region and the second mapping relationship.

[0326] In this embodiment of the present application, the first mapping relationship may be determined based on shake information and the second image. The first mapping relationship includes the mapping relationship between each position point in the second cropped region and the corresponding position point in the second image. Next, the second cropped region may be determined based on the first mapping relationship and the second image. Distortion correction is performed on the second cropped region to obtain a second mapping relationship. The second mapping relationship includes the mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second cropped region; the fourth cropped region is determined based on the second cropped region and the second mapping relationship.

[0327] In one implementation, the first mapping relationship may be further combined with the second mapping relationship to determine a target mapping relationship. The target mapping relationship includes a mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second image; the fourth image is determined based on the target mapping relationship and the second image. That is, the index tables for the first mapping relationship and the second mapping relationship are combined. By combining the index tables, it is not necessary to perform a warp operation after the cropped region is obtained through hand shake correction and then perform a warp operation again after perspective distortion correction. It is only necessary to perform a warp operation once after the output is obtained based on the combined index table after perspective distortion correction, thereby reducing the overhead of the warp operation.

[0328] The first mapping relationship may be referred to as a hand shake correction index table, and the second mapping relationship may be referred to as a perspective distortion correction index table. A detailed description will be provided in detail below. First, how to determine the correction index table (or referred to as the second mapping relationship) for perspective distortion correction will be described below. The mapping relationship between the second cropped region and the fourth cropped region output through hand shake correction may be as follows:

[0329]

Equation

[0330] In the formula, (x0, y0) are the normalized coordinates of the pixel or position point in the second cropped region output through hand shake correction, and the relationship between (x0, y0) and the pixel coordinates (xi, yi), the image width W, and the image height H may be as follows: x0 = (xi - 0.5×W) / W; and y0 = (yi - 0.5×H) / H

[0331] Where (x, y) are the normalized coordinates of the corresponding position point or pixel in the fourth cropped area, and (x, y) need to be converted to the target pixel coordinates (xo, yo) according to the following formula: xo = x × W + 0.5 × W; and yo = y × H + 0.5 × H

[0332] Where K01, K02, K11, and K12 are relationship parameters with specific values determined by the field of view of the distorted image. r x is the distortion distance of (x0, y0) in the x direction, and r y is the distortion distance of (x0, y0) in the y direction. The two distances may be determined by using the following relational expressions:

[0333]

Equation

[0334] Where α x = 1, and α y , β x , and β y are relationship coefficients each having a value ranging from 0.0 to 1.0. When the values of αx, αy, βx, and βy are 1.0, 1.0, 1.0, and 1.0, r x = r y which is the distance from the pixel (x0, y0) to the center of the image. When these values are 1.0, 0.0, 0.0, and 1.0, r x is the distance from the pixel (x0, y0) to the central axis of the image in the x direction, and r y is the distance from the pixel (x0, y0) to the central axis of the image in the y direction. The coefficients may be adjusted based on the effect of perspective distortion correction and the degree of bending of the background straight line in order to achieve a correction effect having a balance between the effect of perspective distortion correction and the degree of bending of the background straight line.

[0335] How to perform shake correction and perspective distortion correction on the first image based on the index table (the first mapping relationship and the second mapping relationship) will be described below.

[0336] In one implementation, the optical distortion correction index table T1 may be determined based on relevant parameters of the imaging module of the camera, and the optical distortion correction index table T1 includes points mapped from positions in the image obtained through optical distortion correction in the first image. The granularity of the optical distortion correction index table T1 may be a pixel or a position point. Next, the hand shake correction index table T2 may be generated, and the hand shake correction index table T2 includes points mapped from positions in the image obtained through shake compensation in the first image. Next, the optical distortion and hand shake correction index table T3 may be generated, and the optical and hand shake correction index table T3 may include points mapped from positions in the image obtained through optical distortion correction and shake compensation in the first image (the optical distortion correction mentioned above is optional, that is, the optical distortion correction may not be performed, and it should be understood that the hand shake correction may be directly performed on the first image to obtain the hand shake correction index table T2). Next, the perspective distortion correction index table T4 is determined based on information regarding the image output through hand shake correction (including the position of the image obtained through optical distortion correction and shake compensation), and the perspective distortion correction index table T4 may include points mapped from the second cropped region output through hand shake correction from the fourth cropped region. Based on T3 and T4, the combined index table T5 may be generated in a coordinate system switching table lookup manner. Refer to FIG. 22b. Specifically, starting from the fourth cropped region, the table may be looked up inversely twice to establish the mapping relationship between the fourth cropped region and the first image. The position point a in the fourth cropped region is used as an example. For the position point a currently being looked up in the fourth cropped region, the mapped position b of point a in the hand shake correction output coordinate system may be determined based on the index according to the perspective distortion correction index table T4.Furthermore, in the optical distortion and camera shake correction output index table T3, the optical distortion correction and camera shake correction shifts delta1, delta2, ..., and deltan of the mesh points c0, c1, ..., and cn adjacent to point b may be obtained based on the mapping relationship between the adjacent mesh points and the coordinates d0, d1, ..., and dn in the first image. Finally, after the optical and camera shake correction shift of point b is calculated based on delta1, delta2, ..., and deltan according to the interpolation algorithm, the mapped position d of point b in the first image is obtained, and as a result, the mapping relationship between the position point a in the fourth cropped region and the mapped position d in the first image is established. For each position point in the fourth cropped region, the combined index table T5 may be obtained by executing the above-described procedure.

[0337] Next, an image warping interpolation algorithm may be executed based on the combined index table T5 and the second image, and a fourth image may be output.

[0338] For example, the second image is processed. For a specific flowchart, refer to FIGS. 22c, 22e, and 22f. FIG. 22c illustrates how to create an index table and the warping process based on the index table T5 that combines the second image and camera shake correction, perspective distortion, and optical distortion. FIGS. 22e, 22f, and 22c differ in that the camera shake correction index table and the optical distortion correction index table are combined in FIG. 22c, the perspective distortion correction index table and the optical distortion correction index table are combined in FIG. 22e, and all distortion correction processes (including optical distortion correction and perspective distortion correction) are combined in FIG. 22f by using a distortion correction module in the video camera shake correction backend. FIG. 22d shows the coordinate system switching table lookup process. In one implementation, it should be understood that the perspective distortion correction index table and the camera shake correction index table may first be combined, and then the optical distortion correction index table is combined. The specific procedure is as follows: Based on the camera shake correction index table T2 and the perspective distortion correction index table T3, an index table T4 that combines camera shake correction and perspective distortion is generated in a coordinate system switching table lookup manner. Next, starting from the output coordinates of the perspective distortion correction, the table is looked up inversely twice to establish the mapping relationship between the output coordinates of the perspective distortion correction and the input coordinates of the camera shake correction. Specifically, referring to FIGS. 23A and 23B, for the point a currently being looked up in the perspective distortion correction index table T3, the mapped position b of the point a in the camera shake correction output (perspective correction input) coordinate system may be determined based on the index. In the camera shake correction index table T2, the camera shake correction shifts delta1', delta2',..., and deltan' of the adjacent mesh points c0, c1,..., and cn of the point b may be obtained based on the mapping relationship between the adjacent mesh points and the camera shake correction input coordinates d0, d1,..., and dn.After the hand shake correction shift of point b is calculated according to the interpolation algorithm, the mapped position d of point b in the hand shake correction input coordinate system is obtained, and as a result, the mapping relationship between the perspective distortion correction output point a and the hand shake correction input coordinate point d is established. For each point in the perspective distortion correction index table T3, the aforementioned procedure is executed to obtain an index table T4 that combines perspective distortion correction and hand shake correction.

[0339] Next, based on T1 and T4, an index table T5 that combines hand shake correction, perspective distortion, and optical distortion may be further generated in a coordinate system switching table lookup manner. Starting from the fourth cropped region, the table is looked up twice in reverse to establish the coordinate mapping relationship between the fourth cropped region and the second image. Specifically, for the currently looked-up position point a in the fourth cropped region, the mapped position d of position point a in the hand shake correction input (optical distortion correction output) coordinate system may be determined based on the index according to the index table T4 that combines perspective distortion correction and hand shake correction. Further, in the optical distortion correction index table T1, the optical distortion correction shifts delta1'', delta2'',..., and deltan'' of the adjacent mesh points d0, d1,..., and dn of point d may be obtained based on the mapping relationship between the adjacent mesh points and the original image coordinates e0, e1,..., and en. After the optical distortion correction shift of point d is calculated according to the interpolation algorithm, the mapped position e of point d in the coordinate system of the second image is obtained, and as a result, the mapping relationship between position point a and the mapped position e in the second image is established. For each position point in the fourth cropped region, the index table T5 that combines hand shake correction, perspective distortion, and optical distortion may be obtained by executing the aforementioned procedure.

[0340] By using the distortion correction module, compared with combining the optical distortion correction index table and the perspective distortion correction index table, optical distortion is an important cause of the jelly phenomenon, and the effectiveness of optical distortion correction and perspective distortion correction is controlled in the video shake correction backend. Therefore, this embodiment can solve the problem that when the video shake correction module invalidates the optical distortion correction, the perspective distortion correction module cannot control the jelly phenomenon.

[0341] In one implementation, in order to strictly maintain the frame-to-frame consistency between the former frame and the latter frame, the correction index table (or referred to as the second mapping relationship) used for perspective distortion correction for the former frame and the latter frame may be the same. As a result, the frame-to-frame consistency between the former frame and the latter frame can be strictly maintained.

[0342] For a further description of step 305, please refer to the description of how to perform perspective distortion correction on the first cropped region in the foregoing embodiment. Similar points will not be described again in this specification.

[0343] 306: Obtain a fourth image based on the second image and the fourth cropped region.

[0344] In one implementation, the first mapping relationship and the second mapping relationship may be obtained, where the first mapping relationship indicates the mapping relationship between each position point in the second cropped region and the corresponding position point in the second image, and the second mapping relationship indicates the mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second cropped region; the fourth image is obtained based on the second image, the first mapping relationship, and the second mapping relationship.

[0345] For further explanation regarding stage 306, refer to the description of obtaining the third image based on the first image and the third cropped region in the foregoing embodiments. Details will not be explained again in this specification.

[0346] 307: Generate a target video based on the third image and the fourth image.

[0347] In one implementation, the target video may be generated based on the third image and the fourth image, and the target video is displayed. Specifically, for any two image frames of a video stream obtained in real time by a terminal device, the output images for which shake correction and perspective distortion correction have been performed may be obtained according to the foregoing stages 301 to 306. Next, a video for which shake correction and perspective distortion correction have been performed (i.e., the target video) may be obtained. For example, the first image may be the nth frame of the video stream, and the second image may be the (n + 1)th frame. In this case, the (n - 2)th frame and the (n - 1)th frame may be further used as the first image and the second image, respectively, for the processing in stages 301 to 306, the (n - 1)th frame and the nth frame may be further used as the first image and the second image, respectively, for the processing in stages 301 to 306, and the (n + 1)th frame and the (n + 2)th frame may be further used as the first image and the second image, respectively, for the processing in stages 301 to 306. The rest can be inferred from the above description, and further details and examples will not be explained in this specification.

[0348] In one implementation, after the target video is obtained, the target video may be displayed in the viewfinder frame in the shooting interface on the camera.

[0349] The shooting interface may be a preview interface in the shooting mode, and it should be understood that the target video is a preview photo in the shooting mode. The user may trigger the camera shutter by clicking on the shooting control in the shooting interface or in another way of the trigger to obtain the captured image.

[0350] The shooting interface may be a preview interface before video recording starts in the video recording mode, and it should be understood that the target video is a preview photo before video recording starts in the video recording mode. The user may trigger the camera to start recording by clicking on the shooting control in the shooting interface or in another way of the trigger.

[0351] The shooting interface may be a preview interface after video recording starts in the video recording mode, and it should be understood that the target video is a preview photo after video recording starts in the video recording mode. The user may trigger the camera to stop recording by clicking on the shooting control in the shooting interface or in another way of the trigger and save the video captured during recording.

[0352] This application provides an image processing method, which is used by a terminal to obtain a video stream in real time, and includes the following steps: performing shake correction on the acquired first image to obtain a first cropped area; performing perspective distortion correction on the first cropped area to obtain a third cropped area; obtaining a third image based on the first image and the third cropped area; performing shake correction on the acquired second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image; performing perspective distortion correction on the second cropped area to obtain a fourth cropped area; and obtaining a fourth image based on the second image and the fourth cropped area, where the third image and the fourth image are used to generate a target video, the first image and the second image are acquired by the terminal, and are a pair of original images adjacent to each other in the time domain. In addition, the outputs of shake correction for two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object in the images is the same, that is, frame-to-frame consistency is maintained, so that no obvious Jello phenomenon occurs, thereby improving the quality of video display.

[0353] The image processing method provided in an embodiment of this application will be described below with reference to the interaction with the terminal side.

[0354] Please refer to FIG. 5a. FIG. 5a is a schematic flowchart of an image processing method according to an embodiment of this application. As shown in FIG. 5a, the image processing method provided in this embodiment of this application includes the following steps.

[0355] 501: Start the camera and display the shooting interface.

[0356] This embodiment of the present application may be applied to a shooting scenario using a terminal. The shooting may include image shooting and video recording. Specifically, the user may start an application related to shooting on the terminal (which may also be referred to as a camera application in the present application), and the terminal may display a shooting interface for shooting.

[0357] In this embodiment of the present application, the shooting interface may be a preview interface in the shooting mode, a preview interface before video recording is started in the video recording mode, or a preview interface after video recording is started in the video recording mode.

[0358] The process in which the terminal starts the camera application and opens the shooting interface will first be described with reference to the interaction between the terminal and the user.

[0359] For example, the user may instruct the terminal to start the camera application by touching a specific control on the screen of the mobile phone, pressing a specific physical button or a set of buttons, inputting voice, performing a gesture without touching, or in another way. In response to receiving an instruction from the user to start the camera, the terminal may start the camera and display the shooting interface.

[0360] For example, as shown in FIG. 5b, the user may click on the application icon 401 "Camera" on the home screen of the mobile phone to instruct the mobile phone to start the camera application, and the mobile phone may display the shooting interface as shown in FIG. 6.

[0361] For another example, when the mobile phone is in the locked screen state, the user may instruct the mobile phone to start the camera application by using a gesture of swiping right on the screen of the mobile phone, and the mobile phone may further display a shooting interface as shown in FIG. 6.

[0362] Alternatively, when the mobile phone is in the locked screen state, the user may click on the shortcut icon of the "Camera" application in the lock screen interface to instruct the mobile phone to start the camera application, and the mobile phone may further display a shooting interface as shown in FIG. 6.

[0363] For another example, when another application is running on the mobile phone, the user may click on the corresponding control to enable the mobile phone to start the camera application for shooting. For example, when the user is using an instant messaging application (e.g., the application WeChat (registered trademark)), the user may alternatively instruct the mobile phone to start the camera application for shooting and video recording by selecting the control of the camera function.

[0364] As shown in FIG. 6, a viewfinder frame 402, a shooting control, and other function controls (such as "wide aperture", "portrait", "photo", "video", etc.) are generally included within the shooting interface on the camera. The viewfinder frame may be used to display a preview photo generated based on the video acquired by the camera. The user may determine the time to instruct the terminal to perform a shooting operation based on the preview photo in the viewfinder frame. The shooting operation to be performed when the user instructs the terminal to perform it may be, for example, the operation of the user clicking on the shooting control or the operation of the user pressing the volume button.

[0365] In one implementation, the user may trigger the camera to enter the image capture mode by clicking on the "Photo" control. When the camera is in the image capture mode, the viewfinder frame in the capture interface may be used to display a preview photo generated based on the video acquired by the camera. The user may determine when to instruct the terminal to perform the capture operation based on the preview photo in the viewfinder frame, and the user may click the capture button to acquire an image.

[0366] In one implementation, the user may trigger the camera to enter the video recording mode by clicking on the "Video" control. When the camera is in the video recording mode, the viewfinder frame in the capture interface may be used to display a preview photo generated based on the video acquired by the camera. The user may determine when to instruct the terminal to perform the video recording based on the preview photo in the viewfinder frame, and the user may click the capture button to start the video recording. In this case, the preview photo during the video recording may be displayed in the viewfinder frame, and then the user may click the capture button to stop the video recording.

[0367] In this embodiment of the present application, the third image and the fourth image may be displayed in the preview photo in the image capture mode, in the preview photo before the video recording is started in the video recording mode, or in the preview photo after the video recording is started in the video recording mode.

[0368] In some embodiments, the zoom ratio indication 403 may be further included in the shooting interface. The user may adjust the zoom ratio of the terminal based on the zoom ratio indication 403. The zoom ratio of the terminal may be described as the zoom ratio of the preview photo in the viewfinder frame in the shooting interface.

[0369] It should be understood that the user may alternatively adjust the zoom ratio of the preview photo in the viewfinder frame in the shooting interface through other interactions (e.g., based on the volume button on the terminal). This is not limited in the present application.

[0370] The zoom ratio of the terminal will be described in detail below.

[0371] When a plurality of cameras are integrated on the terminal, each of the cameras has a fixed zoom ratio. The fixed zoom ratio of the camera can be understood as being equivalent to the focal length of the camera shrinking / expanding by a magnification of the reference focal length. The reference focal length is usually the focal length of the main camera on the terminal.

[0372] In some embodiments, the user adjusts the zoom ratio of the preview photo in the viewfinder frame in the shooting interface, and as a result, the terminal may perform digital zoom on the image captured by the camera. Specifically, the ISP or another processor of the terminal is used to expand the area occupied by each pixel of the image captured by the camera with a fixed zoom ratio and a narrow viewfinder coverage corresponding thereto, and as a result, the processed image is equivalent to the image captured by the camera using another zoom ratio. In the present application, the zoom ratio of the terminal can be understood as the zoom ratio of the image presented in the processed image.

[0373] The images captured by each of the cameras may correspond to a zoom ratio range. Different types of cameras are described below.

[0374] In one implementation, one or more of a front short focus (wide angle) camera, a rear short focus (wide angle) camera, a rear medium focus camera, and a rear long focus camera may be integrated on the terminal.

[0375] An example where a short focus (wide angle) camera (including a front wide angle camera and a rear wide angle camera), a rear medium focus camera, and a rear long focus camera are integrated on the terminal is used for illustration. When the position of the terminal relative to the position of the object to be photographed remains unchanged, the short focus (wide angle) camera has a minimum focal length and a maximum field of view, and the object in the image captured by the camera has a minimum size. The medium focus camera has a focal length larger than that of the short focus (wide angle) camera and a field of view smaller than that of the short focus (wide angle) camera, and the object in the image captured by the medium focus camera has a size larger than that of the object in the image captured by the short focus (wide angle) camera. The long focus camera has a maximum focal length and a minimum field of view, and the object in the image captured by the camera has a maximum size.

[0376] The field of view indicates the maximum range of angles that can be photographed through the camera during image capture of a mobile phone. That is, if the object to be photographed is within the range of angles, the object to be photographed can be acquired by the mobile phone. If the object to be photographed is outside the range of angles, the object to be photographed cannot be acquired by the mobile phone. Usually, a larger field of view of the camera indicates a larger range that can be photographed through the camera. A smaller field of view of the camera indicates a smaller range that can be photographed through the camera. It can be understood that the "field of view" may alternatively be another term, such as "field range", "view range", "field of vision", "imaging range", or "imaging field".

[0377] Typically, the user uses the mid-focus camera in most scenarios. Therefore, the mid-focus camera is usually set as the main camera. The focal length of the main camera is set to the reference focal length, and the fixed zoom ratio of the main camera is usually 1×. In some embodiments, digital zoom may be performed on the image captured by the main camera. Specifically, the ISP of the mobile phone or another processor is used to enlarge the area occupied by each pixel of the "1×" image captured by the main camera and the narrow viewfinder coverage corresponding thereto, and as a result, the processed image is equivalent to the image captured by the main camera using another zoom ratio (for example, "2×"). In other words, the image captured by the main camera may correspond to a zoom ratio range, for example, "1×" to "5×". When the rear telephoto camera is further integrated on the terminal, the zoom ratio range corresponding to the image captured by the main camera may be 1 to b, and it should be understood that the maximum zoom ratio b within the zoom ratio range may be the fixed zoom ratio of the rear telephoto camera. The value of b may range from 3 to 15. For example, b may be, but is not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0378] When the rear telephoto camera is not integrated on the terminal, the maximum zoom ratio b within the zoom ratio range corresponding to the image captured by the main camera may be the maximum zoom of the terminal, and it should be understood that the value of b may range from 3 to 15. For example, b may be, but is not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0379] It should be understood that the zoom ratio range (1 to b) corresponding to the image captured by the main camera may include the zoom ratios of the two endpoints (1 and b), may not include the zoom ratios of the two endpoints (1 and b), or may include one of the zoom ratios of the two endpoints.

[0380] For example, the zoom ratio range corresponding to the image captured by the main camera may be the interval [1, b].

[0381] For example, the zoom ratio range corresponding to the image captured by the main camera may be the interval (1, b].

[0382] For example, the zoom ratio range corresponding to the image captured by the main camera may be the interval [1, b).

[0383] For example, the zoom ratio range corresponding to the image captured by the main camera may be the interval (1, b).

[0384] Similarly, how many times greater the focal length of the telephoto camera is than the focal length of the main camera may be used as the zoom ratio of the telephoto camera, which is also referred to as the fixed zoom ratio of the telephoto camera. For example, the focal length of the telephoto camera may be 5 times the focal length of the main camera, that is, the fixed zoom ratio of the telephoto camera is "5×". Similarly, digital zoom may be further performed on the image captured by the telephoto camera. That is, the image captured by the telephoto camera may correspond to a zoom ratio range, for example, "5×" to "50×". The zoom ratio range corresponding to the image captured by the telephoto camera may extend over the range of b to c, and it should be understood that the maximum zoom ratio c within the zoom ratio range corresponding to the image captured by the telephoto camera may extend over the range of 3 to 50.

[0385] Similarly, how many times smaller the focal length of the wide-angle camera is than the focal length of the main camera may be used as the zoom ratio of the wide-angle camera. For example, the focal length of the wide-angle camera may be 0.5 times the focal length of the main camera, that is, the fixed zoom ratio of the wide-angle camera is "0.5×". Similarly, digital zoom may be further performed on the image captured by the wide-angle camera. That is, the image captured by the wide-angle camera may correspond to a zoom ratio range, for example, "0.5×" to "1×".

[0386] The zoom ratio corresponding to an image captured by the wide-angle camera may range from a to 1. The minimum value a may be the fixed zoom ratio of the wide-angle camera, and the value of a may range from 0.5 to 0.9. For example, a may be, but is not limited to, 0.5, 0.6, 0.7, 0.8, or 0.9.

[0387] It should be understood that the wide-angle camera may include a front wide-angle camera and a rear wide-angle camera. The range of the zoom ratio corresponding to the front wide-angle camera and the rear wide-angle camera may be the same or different. For example, the zoom ratio corresponding to the front wide-angle camera may range from a1 to 1, where the minimum value a1 may be the fixed zoom ratio of the front wide-angle camera, and the value of a1 may range from 0.5 to 0.9. For example, a1 may be, but is not limited to, 0.5, 0.6, 0.7, 0.8, or 0.9. For example, the zoom ratio corresponding to the rear wide-angle camera may range from a2 to 1, where the minimum value a2 may be the fixed zoom ratio of the rear wide-angle camera, and the value of a2 may range from 0.5 to 0.9. For example, a2 may be, but is not limited to, 0.5, 0.6, 0.7, 0.8, or 0.9.

[0388] It should be understood that the zoom ratio range (a to 1) corresponding to an image captured by the wide-angle camera may include the zoom ratios of the two endpoints (a and 1), may not include the zoom ratios of the two endpoints (a and 1), or may include only one of the zoom ratios of the two endpoints.

[0389] For example, the zoom ratio range corresponding to an image captured by the wide-angle camera may be the interval [a, 1].

[0390] For example, the zoom ratio range corresponding to an image captured by the main camera may be the interval (a, 1].

[0391] For example, the zoom ratio range corresponding to an image captured by the main camera may be the interval [a, 1).

[0392] For example, the zoom ratio range corresponding to the image captured by the main camera may be the interval (a, 1).

[0393] 502: Detect a first activation operation of the user to activate the perspective distortion correction function, where the first activation operation includes a first operation on a first control in the shooting interface on the terminal, the first control is used to instruct to activate or deactivate the perspective distortion correction, and the first operation is used to instruct to activate the perspective distortion correction.

[0394] 503: Detect a second activation operation of the user to activate the shake correction function, where the second activation operation includes a second operation on a second control in the shooting interface on the terminal, the second control is used to instruct to activate or deactivate the shake correction function, and the second operation is used to instruct to activate the shake correction function.

[0395] First, how to adjust the zoom ratio of the terminal for shooting will be described.

[0396] In some embodiments, the user may manually adjust the zoom ratio of the terminal for shooting.

[0397] For example, as shown in FIG. 6, the user may adjust the zoom ratio of the terminal by operating on the zoom ratio indication 403 in the shooting interface. For example, when the current zoom ratio of the terminal is "1×", the user may change the zoom ratio of the terminal to "5×" by clicking the zoom ratio indication 403 once or multiple times.

[0398] For another example, the user may decrease the zoom ratio used by the terminal by using a pinch gesture with two fingers (or three fingers) in the shooting interface, or may increase the zoom ratio used by the terminal by using an outward swipe gesture (in a direction opposite to the pinch direction) with two fingers (or three fingers).

[0399] For another example, alternatively, the user may change the zoom ratio of the terminal by dragging the zoom scale in the shooting interface.

[0400] For another example, alternatively, the user may change the zoom ratio of the terminal by switching the currently used camera in the shooting interface or the shooting settings interface. For example, if the user selects to switch to a telephoto camera, the zoom ratio of the terminal will automatically increase.

[0401] For another example, alternatively, the user may change the zoom ratio of the terminal by selecting controls for a telephoto shooting scenario, controls for a long-distance shooting scenario, etc. in the shooting interface or the shooting settings interface.

[0402] In some other embodiments, alternatively, the terminal may automatically identify a specific scenario of an image captured by the camera and automatically adjust the zoom ratio based on the identified specific scenario. For example, if the terminal identifies that the image captured by the camera includes a scene with a large view range, such as the sea, mountains, or a forest, the zoom ratio may automatically decrease. For another example, if the terminal identifies that the image captured by the camera is a distant object, such as a distant bird or an athlete on a sports ground, the zoom ratio may automatically increase. This is not limited in the present application.

[0403] In one implementation, the zoom ratio of the terminal may be detected in real time. When the zoom ratio of the terminal is higher than a second preset threshold, the shake correction function is enabled, and a second control or prompt related to enabling the shake correction function is displayed in the shooting interface.

[0404] In one implementation, the zoom ratio of the terminal may be detected in real time. When the zoom ratio of the terminal is lower than a first preset threshold, the perspective distortion correction is enabled, and a first control or prompt related to the perspective distortion correction is displayed in the shooting interface.

[0405] In one implementation, the controls for triggering to enable the shake correction and the perspective distortion correction may be displayed in the shooting interface, and the shake correction function and the perspective distortion correction function are enabled through interaction with the user, or in the aforementioned mode, the enabling of the shake correction function and the perspective distortion correction function is automatically triggered.

[0406] Alternatively, the controls for triggering to enable the shake correction and the perspective distortion correction are displayed in the shooting parameter adjustment interface on the camera, and the shake correction function and the perspective distortion correction function are enabled through interaction with the user, or in the aforementioned mode, the enabling of all or part of the shake correction function and the perspective distortion correction function is automatically triggered.

[0407] Details are provided below.

[0408] 1. The second control used to instruct to enable or disable the shake correction and the first control used to instruct to enable the perspective distortion correction are included within the shooting interface.

[0409] For example, the camera is in the main camera shooting mode, and the shooting interface on the camera is the preview interface in the shooting mode. Refer to FIG. 7. In this case, a second control 405 (control VS (video shake correction) shown in FIG. 7) used to instruct to enable shake correction and a first control 404 (control DC (distortion correction) shown in FIG. 7) used to instruct to enable perspective distortion correction may be displayed in the shooting interface. Since neither the shake correction nor the perspective distortion correction is enabled at the current time, the display content of the second control may prompt the user that the shake correction is not enabled at the current time, and the display content of the first control may prompt the user that the perspective distortion correction is not enabled at the current time. Refer to FIG. 8. If the user desires to enable the shake correction function, the user may click on the second control 405. In response to the user's click operation on the second control, the interface content shown in FIG. 9 may be displayed. Since the shake correction function is enabled at the current time, the display content of the second control 405 may prompt the user that the shake correction is enabled at the current time. If the user desires to enable the perspective distortion correction function, the user may click on the first control 404. In response to the user's click operation on the first control, the perspective distortion correction function is enabled at the current time, and the display content of the first control 404 may prompt the user that the perspective distortion correction is enabled at the current time.

[0410] If the user wishes to disable the hand shake correction function, the user may click on the second control 405. In response to the user's click operation on the second control, the hand shake correction function is disabled at the current time, and the display content of the second control 405 may prompt the user that the hand shake correction is disabled at the current time. If the user wishes to disable the perspective distortion correction function, the user may click on the first control 404. In this case, the perspective distortion correction function is disabled, and the display content of the first control 404 may prompt the user that the perspective distortion correction is disabled at the current time.

[0411] In one implementation, when the camera is in the wide-angle shooting mode, the second control used to instruct to enable the hand shake correction and the first control used to instruct to enable the perspective distortion correction may be included within the shooting interface on the camera.

[0412] Please refer to FIG. 10. The user may adjust the zoom ratio used by the terminal by operating on the zoom ratio indication 403 in the shooting interface to switch the camera to the wide-angle shooting mode, and as a result, display a shooting interface as shown in FIG. 11. If the user wishes to enable the hand shake correction function, the user may click on the second control 405. In response to the user's click operation on the second control, the hand shake correction function is enabled, and the display content of the second control 405 may prompt the user that the hand shake correction is enabled at the current time. If the user wishes to enable the perspective distortion correction function, the user may click on the first control 404. In response to the user's click operation on the second control, the perspective distortion correction function is enabled at the current time, and the display content of the first control 404 may prompt the user that the perspective distortion correction is enabled at the current time.

[0413] If the user desires to disable the shake correction function, the user may click on the second control 405. In response to the user's click operation on the second control, the shake correction function is disabled at the current time, and the display content of the second control 405 may prompt the user that the shake correction is disabled at the current time. If the user desires to disable the perspective distortion correction function, the user may click on the first control 404. In this case, the perspective distortion correction function is disabled, and the display content of the first control 404 may prompt the user that the perspective distortion correction is disabled at the current time.

[0414] It should be understood that in the telephoto shooting mode, the perspective distortion is not obvious and the perspective distortion correction may not be performed on the image. Therefore, in the shooting interface in the telephoto shooting mode, the second control used to instruct to enable the shake correction may be displayed, and the first control used to instruct to enable the perspective distortion correction may not be displayed, or a prompt indicating that the user has disabled the perspective distortion correction may be displayed, indicating that the perspective distortion correction is disabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling the perspective distortion correction. For example, refer to FIG. 12. The user may adjust the zoom ratio used by the mobile phone by operating on the zoom ratio indication 403 in the shooting interface to switch to the telephoto shooting mode of the camera, and as a result, display a shooting interface as shown in FIG. 13. The first control used to instruct to enable the perspective distortion correction may not be included in the shooting interface shown in FIG. 13.

[0415] The second control used to instruct to enable shake correction and the first control used to trigger to enable perspective distortion correction are displayed in the shooting parameter adjustment interface on the camera.

[0416] In one implementation, when the camera is in the wide-angle shooting mode, the ultra-wide-angle shooting mode, and the medium-focus shooting mode, the second control used to instruct to enable shake correction and the first control used to trigger to enable perspective distortion correction may be displayed in the shooting parameter adjustment interface on the camera.

[0417] In this embodiment of the present application, refer to FIG. 14. The first control may be included within the shooting interface on the camera shown in FIG. 14. The first control is used to instruct to open the shooting parameter adjustment interface. The user may click on the first control, and the terminal device may receive the user's third action on the first control, and in response to the third action, may open the shooting parameter adjustment interface as shown in FIG. 15. As shown in FIG. 15, a second control used to instruct to enable shake correction and a first control used to instruct to enable perspective distortion correction are included within the shooting parameter adjustment interface. Since neither the shake correction nor the perspective distortion correction is enabled at the current time, the display content of the second control may prompt the user that the shake correction is not enabled at the current time, and the display content of the first control may prompt the user that the perspective distortion correction is not enabled at the current time. Refer to FIG. 15. If the user desires to enable the shake correction function, the user may click on the second control. In response to the user's click action on the second control, the interface content shown in FIG. 16 may be displayed. Since the shake correction function is enabled at the current time, the display content of the second control may prompt the user that the shake correction is enabled at the current time. If the user desires to enable the perspective distortion correction function, the user may click on the first control. In response to the user's click action on the second control, the perspective distortion correction function is enabled at the current time, and the display content of the first control may prompt the user that the perspective distortion correction is enabled at the current time.

[0418] If the user wishes to disable the shake correction function, the user may click on the second control. In response to the user's click operation on the second control, the shake correction function is disabled at the current time, and the display content of the second control may prompt the user that the shake correction is disabled at the current time. If the user wishes to disable the perspective distortion correction function, the user may click on the first control. In this case, the perspective distortion correction function is disabled, and the display content of the first control may prompt the user that the perspective distortion correction is disabled at the current time.

[0419] It should be understood that after the user returns to the shooting interface, a prompt indicating whether the shake correction and perspective distortion correction are currently enabled may be further displayed on the shooting interface.

[0420] In one implementation, in the telephoto shooting mode, the perspective distortion is not obvious, and the perspective distortion correction may not be performed on the image. Therefore, in the shooting parameter adjustment interface on the camera in the telephoto shooting mode, the second control used to indicate enabling the shake correction may be displayed, the first control used to indicate enabling the perspective distortion correction may not be displayed, or a prompt indicating that the user has disabled the perspective distortion correction may be displayed, indicating that the perspective distortion correction is disabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling the perspective distortion correction. For example, refer to FIG. 17. The first control used to indicate enabling the perspective distortion correction is not included in the shooting parameter adjustment interface shown in FIG. 17, and a prompt indicating that the user has disabled the perspective distortion correction is included, indicating that the perspective distortion correction is disabled by default.

[0421] A second control used to instruct to enable shake correction is displayed, and perspective distortion correction is enabled by default.

[0422] In one implementation, when the camera is in a wide-angle shooting mode, an ultra-wide-angle shooting mode, or a medium-focus shooting mode, the second control used to instruct to enable shake correction may be displayed in the shooting interface, the first control used to instruct to enable perspective distortion correction may not be displayed, or a prompt indicating that the user has enabled perspective distortion correction is displayed, indicating that perspective distortion correction is enabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling perspective distortion correction.

[0423] In one implementation, when the camera is in a wide-angle shooting mode, the second control used to instruct to enable shake correction may be included within the shooting interface on the camera, and perspective distortion correction is enabled by default.

[0424] In one implementation, when the camera is in an ultra-wide-angle shooting mode, the second control used to instruct to enable shake correction may be included within the shooting interface on the camera, and perspective distortion correction is enabled by default.

[0425] In one implementation, when the camera is in a medium-focus shooting mode, the second control used to instruct to enable shake correction may be included within the shooting interface on the camera, and perspective distortion correction is enabled by default.

[0426] Refer to FIG. 18. The second control 405 used to instruct enabling of shake correction may be displayed on the shooting interface on the camera, and the first control used to instruct enabling of perspective distortion correction may not be displayed. Alternatively, refer to FIG. 19. A prompt 404 indicating that the user has enabled perspective distortion correction is displayed, indicating that perspective distortion correction is enabled by default. This prompt 404 cannot be used for interaction with the user to switch between enabling and disabling perspective distortion correction. Since the shake correction function is not enabled at the current time, the display content of the second control may prompt the user that shake correction is not enabled at the current time. If the user desires to enable the shake correction function, the user may click on the second control. In response to the user's click operation on the second control, the shake correction function is enabled, and the display content of the second control may prompt the user that shake correction is enabled at the current time.

[0427] 4. The second control used to instruct enabling of shake correction is displayed on the shooting parameter adjustment interface on the camera, and perspective distortion correction is enabled by default.

[0428] In one implementation, when the camera is in the wide-angle shooting mode, the ultra-wide-angle shooting mode, or the medium-focus shooting mode, the second control used to instruct enabling of shake correction may be displayed on the shooting parameter adjustment interface on the camera, the first control used to instruct enabling of perspective distortion correction may not be displayed, or a prompt indicating that the user has enabled perspective distortion correction is displayed, indicating that perspective distortion correction is enabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling perspective distortion correction.

[0429] In one implementation, a third image acquired by the camera may be acquired. The third image is an image frame acquired before the first image is acquired. Perspective distortion correction is performed on the third image to obtain a fourth image, which is displayed on the shooting interface on the camera. In this embodiment, the shake correction is not enabled on the camera, and the perspective distortion correction is enabled by default. In this case, the image displayed on the shooting interface is an image (the fourth image) on which the perspective distortion correction is performed but the shake correction is not performed.

[0430] A first control used to instruct to enable the perspective distortion correction is displayed, and the shake correction is enabled by default.

[0431] In one implementation, when the camera is in the wide-angle shooting mode, the ultra-wide-angle shooting mode, or the medium-focus shooting mode, the first control used to instruct to enable the perspective distortion correction may be displayed on the shooting interface, and the second control used to instruct to enable the shake correction may not be displayed, or a prompt indicating that the user has enabled the shake correction is displayed, indicating that the shake correction is enabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling the shake correction.

[0432] In one implementation, when the camera is in the wide-angle shooting mode, the first control used to instruct to enable the perspective distortion correction may be included in the shooting interface on the camera, and the shake correction is enabled by default.

[0433] In one implementation, when the camera is in the ultra-wide-angle shooting mode, the first control used to instruct to enable the perspective distortion correction may be included in the shooting interface on the camera, and the shake correction is enabled by default.

[0434] In one implementation, when the camera is in the central focus shooting mode, the first control used to instruct to enable perspective distortion correction may be included within the shooting interface on the camera, and shake correction is enabled by default.

[0435] Refer to FIG. 20. The first control 404 used to instruct to enable perspective distortion correction may be displayed in the shooting interface on the camera, and the second control used to instruct to enable shake correction may not be displayed. Alternatively, refer to FIG. 21. A prompt 405 indicating that the user has enabled shake correction is displayed, indicating that shake correction is enabled by default. The prompt 405 cannot be used for interaction with the user to switch between enabling and disabling shake correction. Since perspective distortion correction is not enabled at the current time, the display content of the first control may prompt the user that perspective distortion correction is not enabled at the current time. If the user wishes to enable the perspective distortion correction function, the user may click on the first control. In response to the user's click operation on the first control, the perspective distortion correction function is enabled, and the display content of the first control may prompt the user that perspective distortion correction is enabled at the current time.

[0436] 6. The first control used to instruct to enable perspective distortion correction is displayed in the shooting parameter adjustment interface on the camera, and shake correction is enabled by default.

[0437] In one implementation, when the camera is in a wide-angle shooting mode, an ultra-wide-angle shooting mode, or a medium-focus shooting mode, the first control used to instruct enabling perspective distortion correction may be displayed in the shooting parameter adjustment interface on the camera, the second control used to instruct enabling shake correction may not be displayed, or a prompt indicating that the user has enabled shake correction may be displayed, indicating that shake correction is enabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling shake correction.

[0438] In one implementation, a fourth image acquired by the camera may be acquired. The fourth image is an image frame acquired before the first image is acquired. Shake correction is performed on the fourth image, a second output of the shake correction is acquired, and the image indicated by the second output of the shake correction is displayed in the shooting interface on the camera. In this embodiment, when perspective distortion correction is not enabled on the camera, shake correction is enabled by default. In this case, the image displayed in the shooting interface is an image on which shake correction is performed but perspective distortion correction is not performed (the image indicated by the second output of the shake correction).

[0439] 7. Perspective distortion correction and shake correction are enabled by default.

[0440] In one implementation, when the camera is in the wide-angle shooting mode, ultra-wide-angle shooting mode, and medium focal length shooting mode, the first control used to instruct enabling perspective distortion correction may not be displayed in the shooting interface, and the second control used to instruct enabling shake correction may not be displayed. The perspective distortion correction and shake correction functions are enabled by default. Alternatively, a prompt indicating that the user has enabled shake correction is displayed, indicating that shake correction is enabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling shake correction; a prompt indicating that the user has enabled perspective distortion correction is displayed, indicating that perspective distortion correction is enabled by default. This prompt cannot be used for interaction with the user to switch between enabling and disabling perspective distortion correction.

[0441] The foregoing description is provided by using, as an example, a shooting interface on the camera before shooting is started. In a possible implementation, the shooting interface may be the shooting interface before the video recording function is enabled, or the shooting interface after the video recording function is enabled.

[0442] In another embodiment, some controls in the shooting interface may be further hidden on the mobile phone to prevent the image from being blocked by the controls as much as possible, thereby improving the user's visual experience.

[0443] For example, when the user has not been operating in the shooting interface for a long time, the "second control" and "first control" in the foregoing embodiment may not be displayed any more. After the user's click operation on the touch screen is detected, the hidden "second control" and "first control" may be displayed again.

[0444] Regarding how to perform shake correction on the obtained first image to obtain the first cropped region, and how to perform perspective distortion correction on the first cropped region to obtain the third cropped region, refer to the descriptions of steps 301 and 302 in the foregoing embodiment. Details will not be described again in this specification.

[0445] 504: Perform shake correction on the obtained first image to obtain the first cropped region; perform perspective distortion correction on the first cropped region to obtain the third cropped region; obtain a third image based on the first image and the third cropped region; perform shake correction on the obtained second image to obtain a second cropped region, where the second cropped region is related to the first cropped region and shake information, and the shake information indicates the shake that occurs in the process of the terminal obtaining the first image and the second image; perform perspective distortion correction on the second cropped region to obtain a fourth cropped region; obtain a fourth image based on the second image and the fourth cropped region.

[0446] Regarding the description of step 504, refer to the descriptions of steps 301 to 305. Details will not be described again in this specification.

[0447] 505: Generate a target video based on the third image and the fourth image, and display the target video.

[0448] In one implementation, the target video may be generated based on the third image and the fourth image, and the target video is displayed. Specifically, for any two image frames of the video stream acquired by the terminal device in real time, the output images after performing shake correction and perspective distortion correction may be acquired according to the aforementioned steps 301 to 306. Next, a video after performing shake correction and perspective distortion correction (i.e., the target video) may be acquired. For example, the first image may be the nth frame of the video stream, and the second image may be the (n + 1)th frame. In this case, the (n - 2)th frame and the (n - 1)th frame may be further used as the first image and the second image respectively for the processing in steps 301 to 306, the (n - 1)th frame and the nth frame may be further used as the first image and the second image respectively for the processing in steps 301 to 306, and the (n + 1)th frame and the (n + 2)th frame may be further used as the first image and the second image respectively for the processing in steps 301 to 306.

[0449] In one implementation, after the target video is acquired, the target video may be displayed in the viewfinder frame in the shooting interface on the camera.

[0450] The shooting interface may be the preview interface in the shooting mode, and it should be understood that the target video is the preview photo in the shooting mode. The user may trigger the camera to take a photo by clicking the shooting control in the shooting interface or in another way of the trigger, and acquire the taken image.

[0451] The shooting interface may be a preview interface before video recording starts in the video recording mode, and it should be understood that the target video is a preview photo before video recording starts in the video recording mode. The user may trigger the camera to start recording by clicking on the shooting control in the shooting interface or in another way of triggering the trigger.

[0452] The shooting interface may be a preview interface after video recording starts in the video recording mode, and it should be understood that the target video is a preview photo after video recording starts in the video recording mode. The user may trigger the camera to stop recording by clicking on the shooting control in the shooting interface or in another way of triggering the trigger, and save the video captured during recording.

[0453] Example 2: Video post-processing

[0454] The image processing method according to the embodiment of the present application is described by using a scenario of real-time shooting as an example. In another implementation, the shake correction and perspective distortion correction described above may be performed on the videos in the gallery on the terminal device or in the cloud-side server.

[0455] Please refer to FIG. 24a. FIG. 24a is a schematic flowchart of an image processing method according to an embodiment of the present application. As shown in FIG. 24a, the method includes the following steps.

[0456] 2401: Obtain the video stored in the album, where the video includes a first image and a second image.

[0457] The user may click on the application icon "Gallery" on the home screen of the mobile phone (such as shown in Fig. 24b) to instruct the mobile phone to start the gallery application, and the mobile phone may display the application interface as shown in Fig. 25. The user may long-press on the selected video. In response to the user's long-press operation, the terminal device may obtain the video stored in the album and selected by the user, and may display the interface as shown in Fig. 26. Controls for selection (including the "Select" control, the "Delete" control, the "Share" control, and the "Other" control) may be included within the interface.

[0458] 2402: Activate the camera shake correction function and the perspective distortion correction function.

[0459] The user may click on the "Other" control, and an interface such as shown in Fig. 27 may be displayed in response to the user's click operation. Controls for selection (including the "Camera Shake Correction" control and the "Video Distortion Correction" control) may be included within the interface. The user may click on the "Camera Shake Correction" control and the "Video Distortion Correction" control, and may click on the "Complete" control as shown in Fig. 28.

[0460] 2403: Perform camera shake correction on the first image to obtain the first cropped area.

[0461] 2404: Perform perspective distortion correction on the first cropped area to obtain the third cropped area.

[0462] 2405: Obtain the third image based on the first image and the third cropped area.

[0463] 2406: Perform shake correction on the second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image.

[0464] It should be understood that the shake information in this specification may be acquired and stored in real time when the terminal acquires a video.

[0465] 2407: Perform perspective distortion correction on the second cropped area to obtain a fourth cropped area.

[0466] 2408: Obtain a fourth image based on the second image and the fourth cropped area.

[0467] 2409: Generate a target video based on the third image and the fourth image.

[0468] The terminal device may perform shake correction and perspective distortion correction on the selected video. For the detailed description of steps 2403 to 2408, refer to the description in the embodiment corresponding to FIG. 3. The details will not be described again in this specification. After the shake correction and perspective distortion correction are completed, the terminal device may display an interface as shown in FIG. 29 to prompt the user that the shake correction and perspective distortion correction are completed.

[0469] Example 3: Target Tracking

[0470] Please refer to FIG. 30. FIG. 30 is a schematic flowchart of an image processing method provided in an embodiment of the present application. As shown in FIG. 30, the image processing method provided in this embodiment of the present application includes the following steps.

[0471] 3001: Determine the target object in the acquired first image, and acquire a first cropped region including the target object.

[0472] In some scenarios, the camera may enable the object tracking function. This function can ensure that when the posture remains unchanged, an image including the target object that the terminal device is moving towards is output. For example, the object to be tracked is a human object. Specifically, when the camera is in the wide-angle shooting mode or the ultra-wide-angle shooting mode, the coverage area of the acquired original image is large, and the original image includes the target human object. The area where the human object is located can be identified, and the original image is cropped and zoomed in to obtain a cropped region including the human object. The cropped region is a part of the original image. When the human object moves and does not exceed the shooting range of the camera, the human object may be tracked, the area where the human object is located is acquired in real time, and the original image is cropped and zoomed in to obtain a cropped region including the human object. In this way, when the posture remains unchanged, an image including the human object that the terminal device is moving towards may be output.

[0473] Specifically, refer to FIG. 31. The control "Other" may be included in the shooting interface on the camera. The user may click on the control "Other", and in response to the user's click operation, an application interface as shown in FIG. 32 may be displayed. A control for indicating to the user to enable target tracking may be included in the interface, and the user may click on the control to enable the target tracking function of the camera.

[0474] Similar to the hand shake correction described above, in existing implementations, in order to simultaneously perform hand shake correction and perspective distortion correction on an image, a cropped region including a human object needs to be identified in the original image, and perspective distortion correction needs to be performed on all regions of the original image. In this case, this is equivalent to performing perspective distortion correction on the cropped region including the human object, and the cropped region on which perspective distortion correction is performed is output. However, the above-described method has the following problems: When the target object moves, the positions of the cropped regions including the human object in the former image frame and the latter image frame are different. As a result, the shifts for perspective distortion correction with respect to the position points or pixels in the two image frames that will be output are different. The outputs of hand shake correction for two adjacent frames are images containing the same or basically the same content, but the degrees of transformation of the same object between the images are different, that is, frame-to-frame consistency is lost, so the Jello phenomenon occurs and the quality of the video output deteriorates.

[0475] 3002: Perform perspective distortion correction on the first cropped region to obtain a third cropped region.

[0476] In this embodiment, different from performing perspective distortion correction on all regions of the original image according to the existing implementation, in this embodiment of the present application, the perspective distortion correction is performed on the cropped regions. Since the target object moves, the positions of the cropped regions including the target object in the former image frame and the latter image frame are different, and the shifts for perspective distortion correction for the position points or pixels in the two output image frames are different. However, the sizes of the cropped regions of the two frames determined after the target identification is performed are the same. Therefore, when the perspective distortion correction is performed on each of the cropped regions, the degree of perspective distortion correction performed on adjacent frames is the same (because the distances between each position point or pixel in each cropped region of the two frames and the center point of each cropped region are the same). In addition, the cropped regions of two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object between the cropped regions is the same, that is, frame-to-frame consistency is maintained, so that an obvious Jello phenomenon does not occur, thereby improving the quality of the video output.

[0477] In this embodiment of the present application, the same target object (for example, a human object) may be included within the cropped regions of the former frame and the latter frame. Since the object for perspective distortion correction is the cropped region, for the target object, the difference between the shapes of the transformed target objects in the outputs of the perspective distortion correction in the former frame and the latter frame is very small. Specifically, this difference may be within a preset range. This preset range can be understood as being difficult for the human naked eye to distinguish the shape difference, or being difficult for the human naked eye to discover the Jello phenomenon between the former frame and the latter frame.

[0478] In a possible implementation, a second control is included within a shooting interface on the camera. The second control instructs to enable perspective distortion correction, receives a second operation of the user on the second control, and is used to perform perspective distortion correction on the cropped area in response to the second operation.

[0479] For a more detailed description of step 3002, refer to the description of step 302 in the foregoing embodiments. The same description will not be repeated herein.

[0480] 3003: Obtain a third image based on the first image and the third cropped area.

[0481] For a detailed description of step 3003, refer to the description of step 303 in the foregoing embodiments. The details will not be repeated herein.

[0482] 3004: Determine a target object in the obtained second image, and obtain a second cropped area including the target object, where the second cropped area is related to target movement information, and the target movement information indicates the movement of the target object in the process of the terminal obtaining the first image and the second image.

[0483] For a detailed description of step 3004, refer to the description of step 301 in the foregoing embodiments. The same description will not be repeated herein.

[0484] 3005: Perform perspective distortion correction on the second cropped area to obtain a fourth cropped area.

[0485] For a detailed description of step 3005, refer to the description of step 305 in the foregoing embodiments. The same description will not be repeated herein.

[0486] 3006: Obtain a fourth image based on the second image and the fourth cropped region.

[0487] 3007: Generate a target video based on the third image and the fourth image.

[0488] For a detailed description of step 3006, please refer to the description of step 306 in the foregoing embodiments. Similar descriptions will not be repeated herein.

[0489] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.

[0490] In a possible implementation, before perspective distortion correction is performed on the first cropped region, the method: further includes the steps of detecting that the terminal meets the distortion correction condition and enabling the distortion correction function.

[0491] In a possible implementation, detecting that the terminal meets the distortion correction condition includes one or more of the following cases, but is not limited thereto: Case 1: It is detected that the zoom ratio of the terminal is lower than a first preset threshold.

[0492] In one implementation, whether perspective distortion correction is enabled may be determined based on the zoom ratio of the terminal. When the zoom ratio of the terminal is large (for example, within the full or partial zoom ratio range of the telephoto shooting mode, or within the full or partial zoom ratio range of the medium focal length shooting mode), the degree of perspective distortion in the photo of the image acquired by the terminal is low. The fact that the degree of perspective distortion is low can be understood as meaning that almost no perspective distortion in the photo of the image acquired by the terminal can be identified by the human eye. Therefore, when the zoom ratio of the terminal is large, perspective distortion correction need not be enabled.

[0493] ​In a possible implementation, when the zoom ratio ranges from a to 1, the terminal acquires a video stream in real time by using the front wide-angle camera or the rear wide-angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or when the zoom ratio ranges from 1 to b, the terminal acquires a video stream in real time by using the rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0494] Case 2: A first activation operation of the user to enable the perspective distortion correction function is detected.

[0495] In a possible implementation, the first activation operation includes a first operation on a first control in the shooting interface on the terminal, the first control is used to instruct to enable or disable perspective distortion correction, and the first operation is used to instruct to enable perspective distortion correction.

[0496] In one implementation, the control (referred to as the first control in this embodiment of the present application) used to instruct to enable or disable perspective distortion correction may be included in the shooting interface. The user may enable perspective distortion correction through the first operation on the first control, that is, trigger to enable perspective distortion correction through the first operation on the first control. In this way, the terminal may detect a first activation operation of the user to enable perspective distortion correction, and the first activation operation includes a first operation on a first control in the shooting interface on the terminal.

[0497] In one implementation, a control (referred to as the first control in this embodiment of the present application) used to instruct enabling or disabling perspective distortion correction is displayed in the shooting interface only when it is detected that the zoom ratio of the terminal is lower than a first preset threshold value.

[0498] Case 3: A human face is identified in the shooting scenario; A human face is identified in the shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or A human face is identified in the shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio.

[0499] When a human face exists in the shooting scenario, the degree of transformation of the human face due to perspective distortion correction is more visually obvious, and when the distance between the human face and the terminal is smaller or the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher, the size of the human face in the image is larger, so it should be understood that the degree of transformation of the human face due to perspective distortion correction is even more visually obvious. Therefore, perspective distortion correction needs to be enabled in the aforementioned scenarios. In this embodiment, whether to enable perspective distortion correction is determined by determining the aforementioned conditions related to the human face in the shooting scenario, and as a result, the shooting scenarios in which perspective distortion occurs can be accurately determined, and perspective distortion correction is executed for the shooting scenarios in which perspective distortion occurs. In a shooting scenario where no perspective distortion occurs, perspective distortion correction is not executed. In this way, the image signal is accurately processed and power consumption is reduced.

[0500] This embodiment of the present application is an image processing method, comprising: determining a target object in the acquired first image, and obtaining a first cropped region including the target object; performing perspective distortion correction on the first cropped region to obtain a third cropped region; obtaining a third image based on the first image and the third cropped region; determining a target object in the acquired second image, and obtaining a second cropped region including the target object, where the second cropped region is related to target movement information, and the target movement information indicates the movement of the target object in the process of the terminal acquiring the first image and the second image; performing perspective distortion correction on the second cropped region to obtain a fourth cropped region; and obtaining a fourth image based on the second image and the fourth cropped region, where the third image and the fourth image are used to generate a video, the first image and the second image are acquired by the terminal, and are a pair of original images adjacent to each other in the time domain, and provides an image display method. Different from performing perspective distortion correction on all regions of the original image according to the existing implementation, in this embodiment of the present application, the perspective distortion correction is performed on the cropped region. Since the target object moves, the positions of the cropped regions including the human object in the former image frame and the latter image frame are different, and the shifts for perspective distortion correction for the position points or pixels in the two output image frames are different. However, the sizes of the cropped regions of the two frames determined after target identification is performed are the same. Therefore, when the perspective distortion correction is performed on each of the cropped regions, the degree of perspective distortion correction performed on adjacent frames is the same (because the distances between each position point or pixel and the center point of the cropped region in each of the cropped regions of the two frames are the same).In addition, the cropped regions of two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object in the images is the same. That is, since frame-to-frame consistency is maintained, no obvious Jello phenomenon occurs, thereby improving the quality of the video output.

[0501] The present application further provides an image display device. The image display device may be a terminal device. Please refer to FIG. 33. FIG. 33 is a schematic diagram of the structure of an image processing device according to an embodiment of the present application. As shown in FIG. 33, the image processing device 3300 is: equipped with a shake correction module 3301 configured to perform shake correction on the acquired first image to obtain a first cropped region, and further configured to perform shake correction on the acquired second image to obtain a second cropped region, where the second cropped region is related to the first cropped region and shake information, the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image, and the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain.

[0502] For a detailed description of the shake correction module 3301, please refer to the description of step 301 and step 304. Details will not be described again in this specification.

[0503] The image processing device is further equipped with a perspective distortion correction module 3302 configured to perform perspective distortion correction on the first cropped region to obtain a third cropped region, and further configured to perform perspective distortion correction on the second cropped region to obtain a fourth cropped region.

[0504] For a detailed description of the perspective distortion correction module 3302, please refer to the description of step 302 and step 305. Details will not be described again in this specification.

[0505] The image processing apparatus further includes an image generation module 3303 configured to obtain a third image based on a first image and a third cropped region, obtain a fourth image based on a second image and a fourth cropped region, and generate a target video based on the third image and the fourth image.

[0506] For a detailed description of the image generation module 3303, refer to the descriptions of step 303, step 306, and step 307. Details will not be described again in this specification.

[0507] In a possible implementation, the direction of the offset of the position of the second cropped region in the second image with respect to the position of the first cropped region in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal obtaining the first image and the second image.

[0508] In a possible implementation, the first cropped region represents a first sub-region in the first image, and the second cropped region represents a second sub-region in the second image.

[0509] The similarity between the first image content corresponding to the first sub-region in the first image and the second image content corresponding to the second sub-region in the second image is higher than the similarity between the first image and the second image.

[0510] In a possible implementation, the apparatus: A first detection module 3304 configured to detect that the terminal satisfies a distortion correction condition and activate a distortion correction function before the hand shake correction module performs hand shake correction on the obtained first image. further includes.

[0511] In a possible implementation, detecting that the terminal satisfies a distortion correction condition is: detecting that the zoom ratio of the terminal is lower than a first preset threshold. has, but is not limited to,

[0512] In a possible implementation, when the zoom ratio ranges from a to 1, the terminal acquires a video stream in real time by using a front - wide - angle camera or a rear - wide - angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front - wide - angle camera or the rear - wide - angle camera, and the value of a ranges from 0.5 to 0.9; or When the zoom ratio ranges from 1 to b, the terminal acquires a video stream in real time by using a rear - main camera, where b is the fixed zoom ratio of the rear - telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0513] In a possible implementation, detecting that the terminal meets the distortion correction condition is: detecting a first activation operation of the user to enable the perspective distortion correction function, where the first activation operation includes a first operation on a first control on the shooting interface of the terminal, the first control is used to indicate enabling or disabling perspective distortion correction, and the first operation is used to indicate enabling perspective distortion correction has, but is not limited to,

[0514] In a possible implementation, detecting that the terminal meets the distortion correction condition is: identifying a human face in a shooting scenario; identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or identifying a human face in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio has, but is not limited to,

[0515] In a possible implementation, the device is: Before the hand shake correction module performs hand shake correction on the acquired first image, a second detection module 3305 configured to detect that the terminal satisfies the hand shake correction condition and activate the hand shake correction function is further provided.

[0516] In a possible implementation, detecting that the terminal satisfies the hand shake correction condition includes: detecting that the zoom ratio of the terminal is higher than the fixed zoom ratio of the camera having the minimum zoom ratio on the terminal but is not limited thereto.

[0517] In a possible implementation, detecting that the terminal satisfies the hand shake correction condition includes: detecting a second activation operation of the user to activate the hand shake correction function, where the second activation operation includes a second operation on a second control in the shooting interface on the terminal, the second control is used to instruct to activate or deactivate the hand shake correction function, and the second operation is used to instruct to activate the hand shake correction function but is not limited thereto.

[0518] In a possible implementation, the image generation module is specifically configured to: obtain a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates the mapping relationship between each position point in the second cropped area and the corresponding position point in the second image, and the second mapping relationship indicates the mapping relationship between each position point in the fourth cropped area and the corresponding position point in the second cropped area; and obtain a fourth image based on the second image, the first mapping relationship, and the second mapping relationship is specifically configured to perform.

[0519] In a possible implementation, the image generation module is: Combining the first mapping relationship with the second mapping relationship to determine a target mapping relationship, where the target mapping relationship includes the mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second image; and Determining a fourth image based on the second image and the target mapping relationship is specifically configured to perform.

[0520] In a possible implementation, the first mapping relationship is different from the second mapping relationship.

[0521] In a possible implementation, the perspective distortion correction module is specifically configured to: perform optical distortion correction on the first cropped region to obtain a corrected first cropped region; and perform perspective distortion correction on the corrected first cropped region; The image generation module is specifically configured to: perform optical distortion correction on the third cropped region to obtain a corrected third cropped region; and obtain a third image based on the first image and the corrected third cropped region; The perspective distortion correction module is specifically configured to: perform optical distortion correction on the second cropped region to obtain a corrected second cropped region; and perform perspective distortion correction on the corrected second cropped region; or The image generation module is specifically configured to: perform optical distortion correction on the fourth cropped region to obtain a corrected fourth cropped region; and obtain a fourth image based on the second image and the corrected fourth cropped region.

[0522] This embodiment of the present application provides an image processing apparatus. The apparatus is used by a terminal to acquire a video stream in real time. The apparatus includes: a shake correction module configured to perform shake correction on the acquired first image to obtain a first cropped area, and further configured to perform shake correction on the acquired second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates a shake that occurs in the process of the terminal acquiring the first image and the second image, and the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain; a perspective distortion correction module configured to perform perspective distortion correction on the first cropped area to obtain a third cropped area, and further configured to perform perspective distortion correction on the second cropped area to obtain a fourth cropped area; and an image generation module configured to obtain a third image based on the first image and the third cropped area, obtain a fourth image based on the second image and the fourth cropped area, and further configured to generate a target video based on the third image and the fourth image. In addition, the outputs of the shake correction for two adjacent frames are images including the same or substantially the same content, and the degree of transformation of the same object in the images is the same, that is, the inter-frame consistency is maintained, so that an obvious Jello phenomenon does not occur, thereby improving the quality of the video display.

[0523] The present application further provides an image processing apparatus. The image display apparatus may be a terminal device. Please refer to FIG. 34. FIG. 34 is a schematic diagram of the structure of an image processing apparatus according to an embodiment of the present application. As shown in FIG. 34, the image processing apparatus 3400 includes: configured to determine a target object in the acquired first image and obtain a first cropped region including the target object, and further configured to determine a target object in the acquired second image and obtain a second cropped region including the target object, where the second cropped region is related to target movement information, and the target movement information indicates the movement of the target object in the process of the terminal acquiring the first image and the second image, and the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain.

[0524] For a detailed description of the object determination module 3401, refer to the descriptions of steps 3001 and 3004. Details will not be described again in this specification.

[0525] The apparatus is further provided with a perspective distortion correction module 3402 configured to perform perspective distortion correction on the first cropped region to obtain a third cropped region, and further configured to perform perspective distortion correction on the second cropped region to obtain a fourth cropped region.

[0526] For a detailed description of the perspective distortion correction module 3402, refer to the descriptions of steps 3002 and 3005. Details will not be described again in this specification.

[0527] The apparatus is further provided with an image generation module 3403 configured to obtain a third image based on the first image and the third cropped region, obtain a fourth image based on the second image and the fourth cropped region, and generate a target video based on the third image and the fourth image.

[0528] For a detailed description of the image generation module 3403, refer to the descriptions of steps 3003, 3006, and 3007. Details will not be described again in this specification.

[0529] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.

[0530] In a possible implementation, the apparatus is: A first detection module 3404 configured to detect that the terminal satisfies a distortion correction condition and enable a distortion correction function before the perspective distortion correction module performs perspective distortion correction on the first cropped region. further comprises.

[0531] In a possible implementation, detecting that the terminal satisfies a distortion correction condition is: detecting that the zoom ratio of the terminal is lower than a first preset threshold but not limited thereto.

[0532] In a possible implementation, when the zoom ratio ranges from a to 1, the terminal acquires a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, where the first preset threshold is higher than 1, a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or when the zoom ratio ranges from 1 to b, the terminal acquires a video stream in real time by using a rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15.

[0533] In a possible implementation, detecting that the terminal satisfies a distortion correction condition is: Detecting a first activation operation of a user to activate a perspective distortion correction function, where the first activation operation includes a first operation on a first control in a shooting interface on the terminal, the first control is used to instruct to activate or deactivate perspective distortion correction, and the first operation is used to instruct to activate perspective distortion correction but not limited thereto.

[0534] In a possible implementation, detecting that the terminal meets the distortion correction condition may include: Identifying a human face in a shooting scenario; Identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or Identifying a human face in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio but not limited thereto.

[0535] This embodiment of the present application provides an image processing apparatus. The apparatus is used by a terminal to acquire a video stream in real time. The apparatus includes: an object determination module configured to determine a target object in the acquired first image and acquire a first cropped region including the target object, and further configured to determine a target object in the acquired second image and acquire a second cropped region including the target object, where the second cropped region is related to target movement information, and the target movement information indicates the movement of the target object in the process of the terminal acquiring the first image and the second image, and the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain; a perspective distortion correction module configured to perform perspective distortion correction on the first cropped region to acquire a third cropped region, and further configured to perform perspective distortion correction on the second cropped region to acquire a fourth cropped region; and an image generation module configured to acquire a third image based on the first image and the third cropped region, acquire a fourth image based on the second image and the fourth cropped region, and further configured to generate a target video based on the third image and the fourth image.

[0536] Unlike performing perspective distortion correction for all regions of the original image according to existing implementations, in this embodiment of the present application, perspective distortion correction is performed on the cropped regions. Since the target object moves, the positions of the cropped regions including human objects in the former image frame and the latter image frame are different, and the shifts for perspective distortion correction for the position points or pixels in the two output image frames are different. However, the sizes of the cropped regions of the two frames determined after target identification is performed are the same. Therefore, when perspective distortion correction is performed for each of the cropped regions, the degree of perspective distortion correction performed for adjacent frames is the same (because the distance between each position point or pixel and the center point of the cropped region in each of the cropped regions of the two frames is the same). In addition, the cropped regions of two adjacent frames are images containing the same or basically the same content, and the degree of transformation of the same object between the images is the same, that is, frame-to-frame consistency is maintained, so no obvious Jello phenomenon occurs, thereby improving the quality of the video output.

[0537] A terminal device according to an embodiment of the present application will be described below. The terminal device may be an image processing device in FIG. 33 or FIG. 34. Please refer to FIG. 35. FIG. 35 is a schematic diagram of the structure of a terminal device according to an embodiment of the present application. The terminal device 3500 may be specifically represented as a virtual reality (VR) device, a mobile phone, a tablet, a laptop computer, an intelligent wearable device, etc. This is not limited in this specification. Specifically, the terminal device 3500 includes a receiver 3501, a transmitter 3502, a processor 3503, and a memory 3504 (one or more processors 3503 may exist in the terminal device 3500, and one processor is used as an example in FIG. 35). The processor 3503 may include an application processor 35031 and a communication processor 35032. In some embodiments of the present application, the receiver 3501, the transmitter 3502, the processor 3503, and the memory 3504 may be connected through a bus or in another manner.

[0538] The memory 3504 includes a read-only memory and a random access memory, and may provide instructions and data for the processor 3503. A part of the memory 3504 may further include a non-volatile random access memory (NVRAM). The memory 3504 stores the processor and operating instructions, executable modules or data structures, subsets thereof, or extended sets thereof. The operating instructions may include various operating instructions for implementing various operations.

[0539] The processor 3503 controls the operation of the terminal device. In a specific application, the components of the terminal device are coupled to each other by using a bus system. In addition to the data bus, the bus system may further include a power bus, a control bus, a status signal bus, etc. However, for a clearer description, various types of buses in the figure are referred to as a bus system.

[0540] The method disclosed in the foregoing embodiments of the present application may be applied to the processor 3503 or may be implemented by using the processor 3503. The processor 3503 may be an integrated circuit chip having signal processing capabilities. During implementation, the steps in the foregoing method may be implemented by using hardware integrated logic circuits or instructions in the form of software in the processor 3503. The processor 3503 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or another programmable logic device, discrete gates or transistor logic devices, or discrete hardware assemblies. The processor 3503 may implement or execute the method, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor or the like. The steps in the method disclosed with reference to the embodiments of the present application may be presented directly as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules may be disposed in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is disposed in the memory 3504, and the processor 3503 reads the information in the memory 3504 and combines it with the hardware of the processor to complete the steps in the foregoing method.Specifically, the processor 3503 reads the information in the memory 3504 and, in combination with the hardware of the processor 3503, completes the steps related to the data processing in steps 301 to 304 in the foregoing embodiments and the steps related to the data processing in steps 3001 to 3004 in the foregoing embodiments.

[0541] The receiver 3501 may be configured to receive input digital or character information and generate a signal input related to the settings related to the terminal device and the function control of the terminal device. The transmitter 3502 may be configured to output digital or character information through the first interface. The transmitter 3502 may be further configured to send a command to the disk pack through the first interface to change the data in the disk pack. The transmitter 3502 may further include a display device, for example, a display screen.

[0542] One embodiment of the present application further provides a computer program product. When the computer program product is executed on a computer, the computer is enabled to execute the steps in the image processing method described in the embodiments corresponding to FIGS. 5a, 24a, and 30 in the foregoing embodiments.

[0543] One embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a program for signal processing. When the program is executed on a computer, the computer is enabled to execute the steps in the image processing method in the embodiments of the method described above.

[0544] The image display device provided in this embodiment of the present application may specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, a circuit, etc. The processing unit may execute computer-executable instructions stored in the storage unit. As a result, the chip in the execution device executes the data processing method described in the foregoing embodiment, or the chip in the training device executes the data processing method described in the foregoing embodiment. Optionally, the storage unit is a storage unit in the chip, for example, a register or a cache. Alternatively, the storage unit is in a wireless access device and is a storage unit external to the chip, for example, a read-only memory (ROM) or another type of static storage device capable of storing static information and instructions, or a random access memory (RAM).

[0545] In addition, it should be noted that the device embodiments described above are merely examples. The units described as separate parts may or may not be physically separate, and the parts shown as units may or may not be physical units. They may have one location or may be distributed via a plurality of network units. All or some of the modules may be selected as actually required to achieve the object of the solution means of the embodiment. In addition, in the accompanying drawings corresponding to the device embodiments provided in the present application, the connection relationship between the modules indicates that the modules may have a communication connection that can be specifically implemented through one or a plurality of communication buses or signal wires.

[0546] Based on the description in the foregoing implementation, a person skilled in the art can clearly understand that this application can be implemented by using software in addition to the necessary general-purpose hardware, or by using dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated elements or components, etc. Usually, any function implemented by using a computer program can be easily implemented by using the corresponding hardware. In addition, there may be various types of specific hardware structures used to implement the same function, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, the implementation of the software program is a better implementation in most cases. Based on such an understanding, the technical solution of this application, or the part contributing to the prior art, can be essentially embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk on a computer, a USB flash drive, a removable hard disk, a ROM, a RAM, a magnetic disk, or an optical disk, and includes several instructions that enable a computer device (which can be a personal computer, a server, a network device, etc.) to execute the method described in the embodiments of this application.

[0547] All or some of the foregoing embodiments may be implemented by using software, hardware, firmware, or any combination thereof. When software is used for implementation, all or some of the embodiments may be implemented in the form of a computer program product.

[0548] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the procedures or functions according to the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or another programmable device. The computer instructions may be stored in a computer-readable storage medium or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (e.g., by using coaxial cable, optical fiber, or digital subscriber line (DSL)) or a wireless manner (e.g., via infrared rays, radio waves, or microwaves). The computer-readable storage medium may be any usable medium accessible by a computer or a data storage device integrating one or more usable media, such as a server or a data center. The usable medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a DVD), a semiconductor medium (e.g., a solid-state drive, Solid-State Drive (SSD)), etc. (Other possible items) (Item 1) An image processing method, which is used by a terminal to acquire a video stream in real time, and the method includes: Performing shake correction on the acquired first image to obtain a first cropped area; Performing perspective distortion correction on the first cropped area to obtain a third cropped area; Obtaining a third image based on the first image and the third cropped area; Performing shake correction on the acquired second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image; Performing perspective distortion correction on the second cropped area to obtain a fourth cropped area; Obtaining a fourth image based on the second image and the fourth cropped area; and Generating a target video based on the third image and the fourth image The method is characterized in that the first image and the second image are original image pairs acquired by the terminal and adjacent to each other in the time domain. (Item 2) The direction of the offset of the position of the second cropped area in the second image with respect to the position of the first cropped area in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal acquiring the first image and the second image, according to the method described in Item 1. (Item 3) The first cropped area indicates a first sub-region in the first image, and the second cropped area indicates a second sub-region in the second image; The method according to item 1 or 2, wherein a similarity between first image content corresponding to the first sub-region in the first image and second image content corresponding to the second sub-region in the second image is higher than a similarity between the first image and the second image. (Item 4) Before the step of performing perspective distortion correction on the first cropped area, the method includes: A step of detecting that the terminal satisfies a distortion correction condition, and a step of enabling a distortion correction function The method according to any one of items 1 to 3, further comprising: (Item 5) The step of detecting that the terminal satisfies a distortion correction condition includes: A step of detecting that a zoom ratio of the terminal is lower than a first preset threshold The method according to item 4, comprising: (Item 6) When the zoom ratio ranges from a to 1, a step of acquiring a video stream in real time by using a front wide-angle camera or a rear wide-angle camera by the terminal, where the first preset threshold is higher than 1, and a is a fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or When the zoom ratio ranges from 1 to b, a step of acquiring a video stream in real time by using a rear main camera by the terminal, where b is a fixed zoom ratio of a rear telephoto camera on the terminal or a maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15. The method according to item 5. (Item 7) The step of detecting that the terminal satisfies a distortion correction condition includes: A step of detecting a first activation operation of a user for enabling the perspective distortion correction function, where the first activation operation includes a first operation on a first control in a shooting interface on the terminal, the first control is used to indicate enabling or disabling perspective distortion correction, and the first operation is used to indicate enabling perspective distortion correction The method according to item 4, comprising: (Item 8) The step of detecting that the terminal satisfies a distortion correction condition includes: The stage of identifying a human face in a shooting scenario; The stage of identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or The stage of identifying a human face in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio The method according to item 4, having the above. (Item 9) Before the stage of performing shake correction on the acquired first image, the method includes: The stage of detecting that the terminal meets the shake correction condition and the stage of enabling the shake correction function The method according to any one of items 1 to 8, further comprising the above. (Item 10) The stage of detecting that the terminal meets the shake correction condition includes: The stage of detecting that the zoom ratio of the terminal is higher than the fixed zoom ratio of the camera with the minimum zoom ratio on the terminal The method according to item 9, having the above. (Item 11) The stage of detecting that the terminal meets the shake correction condition includes: The stage of detecting a second activation operation of the user for enabling the shake correction function, where the second activation operation includes a second operation on a second control in the shooting interface on the terminal, the second control is used to indicate enabling or disabling the shake correction function, and the second operation is used to indicate enabling the shake correction function The method according to any one of items 1 to 10, having the above. (Item 12) The stage of acquiring a fourth image based on the second image and the fourth cropped area includes: The stage of acquiring a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates the mapping relationship between each position point in the second cropped area and the corresponding position point in the second image, and the second mapping relationship indicates the mapping relationship between each position point in the fourth cropped area and the corresponding position point in the second cropped area; and The stage of acquiring the fourth image based on the second image, the first mapping relationship, and the second mapping relationship The method according to any one of items 1 to 11, having the above. (Item 13) The step of obtaining the fourth image based on the second image, the first mapping relationship, and the second mapping relationship is as follows: Combining the first mapping relationship with the second mapping relationship to determine a target mapping relationship, where the target mapping relationship includes the mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second image; and Determining the fourth image based on the second image and the target mapping relationship The method according to item 12, comprising the above steps. (Item 14) The method according to item 12 or 13, wherein the first mapping relationship is different from the second mapping relationship. (Item 15) The step of performing perspective distortion correction on the first cropped region is as follows: performing optical distortion correction on the first cropped region to obtain a corrected first cropped region; and performing perspective distortion correction on the corrected first cropped region; The step of obtaining a third image based on the first image and the third cropped region is as follows: performing optical distortion correction on the third cropped region to obtain a corrected third cropped region; and obtaining the third image based on the first image and the corrected third cropped region; The step of performing perspective distortion correction on the second cropped region is as follows: performing optical distortion correction on the second cropped region to obtain a corrected second cropped region; and performing perspective distortion correction on the corrected second cropped region; or The step of obtaining a fourth image based on the second image and the fourth cropped region is as follows: performing optical distortion correction on the fourth cropped region to obtain a corrected fourth cropped region; and obtaining the fourth image based on the second image and the corrected fourth cropped region. The method according to any one of items 1 to 14. (Item 16) An image processing apparatus, which is used by a terminal to obtain a video stream in real time, and the apparatus is as follows: configured to perform camera shake correction on the acquired first image to obtain a first cropped area, and further configured to perform camera shake correction on the acquired second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates a shake that occurs in the process of the terminal acquiring the first image and the second image, and the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain; a perspective distortion correction module configured to perform perspective distortion correction on the first cropped area to obtain a third cropped area, and further configured to perform perspective distortion correction on the second cropped area to obtain a fourth cropped area; and an image generation module configured to obtain a third image based on the first image and the third cropped area, obtain a fourth image based on the second image and the fourth cropped area, and further configured to generate a target video based on the third image and the fourth image The device comprises. (Item 17) The direction of the offset between the position of the second cropped area in the second image and the position of the first cropped area in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal acquiring the first image and the second image. The device according to Item 16. (Item 18) The first cropped area indicates a first sub-area in the first image, and the second cropped area indicates a second sub-area in the second image; The similarity between the first image content corresponding to the first sub-area in the first image and the second image content corresponding to the second sub-area in the second image is higher than the similarity between the first image and the second image. The device according to Item 16 or 17. (Item 19) The device is: A first detection module configured to detect that the terminal meets the distortion correction condition and activate the distortion correction function before the camera shake correction module performs camera shake correction on the acquired first image The device according to any one of items 16 to 18, further comprising (Item 20) Detecting that the terminal satisfies the distortion correction condition includes: Detecting that the zoom ratio of the terminal is lower than a first preset threshold The device according to item 19, which has (Item 21) When the zoom ratio ranges from a to 1, the terminal obtains a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, where the first preset threshold is higher than 1, and a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or When the zoom ratio ranges from 1 to b, the terminal obtains a video stream in real time by using a rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15. The device according to item 20 (Item 22) Detecting that the terminal satisfies the distortion correction condition includes: Detecting a first activation operation of the user to activate the perspective distortion correction function, where the first activation operation includes a first operation on a first control in the shooting interface on the terminal, and the first control is used to instruct to activate or deactivate perspective distortion correction, and the first operation is used to instruct to activate perspective distortion correction The device according to item 19, which has (Item 23) Detecting that the terminal satisfies the distortion correction condition includes: Identifying a human face in a shooting scenario; Identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or Identifying a human face in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio The device according to item 19, which has (Item 24) The device is: A second detection module configured to detect that the terminal satisfies a hand shake correction condition and activate a hand shake correction function before the hand shake correction module performs hand shake correction on the acquired first image The apparatus according to any one of items 16 to 23, further comprising (Item 25) Detecting that the terminal satisfies the hand shake correction condition includes: Detecting that the zoom ratio of the terminal is higher than the fixed zoom ratio of a camera having a minimum zoom ratio on the terminal The apparatus according to item 24, comprising (Item 26) Detecting that the terminal satisfies the hand shake correction condition includes: Detecting a second activation operation of the user for activating the hand shake correction function, where the second activation operation includes a second operation on a second control in the shooting interface on the terminal, the second control is used to instruct activation or deactivation of the hand shake correction function, and the second operation is used to instruct activation of the hand shake correction function The apparatus according to any one of items 16 to 25, comprising (Item 27) The image generation module obtains a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates a mapping relationship between each position point in the second cropped region and the corresponding position point in the second image, and the second mapping relationship indicates a mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second cropped region; and Obtaining the fourth image based on the second image, the first mapping relationship, and the second mapping relationship The apparatus according to any one of items 16 to 26, specifically configured to perform (Item 28) The image generation module: Combines the first mapping relationship with the second mapping relationship to determine a target mapping relationship, where the target mapping relationship includes a mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second image; and Determines the fourth image based on the second image and the target mapping relationship The apparatus according to item 27, specifically configured to perform (Item 29) The first mapping relationship is different from the second mapping relationship, the apparatus according to item 27 or 28 (Item 30) The perspective distortion correction module is specifically configured to: perform optical distortion correction on the first cropped region to obtain a corrected first cropped region; and perform perspective distortion correction on the corrected first cropped region; The image generation module is specifically configured to: perform optical distortion correction on the third cropped region to obtain a corrected third cropped region; and obtain the third image based on the first image and the corrected third cropped region; The perspective distortion correction module is specifically configured to: perform optical distortion correction on the second cropped region to obtain a corrected second cropped region; and perform perspective distortion correction on the corrected second cropped region; or The image generation module is specifically configured to: perform optical distortion correction on the fourth cropped region to obtain a corrected fourth cropped region; and obtain the fourth image based on the second image and the corrected fourth cropped region. The device according to any one of items 16 to 29. (Item 31) An image processing device, the device comprising a processor, a memory, a camera, and a bus, The processor, the memory, and the camera are connected through the bus; The camera is configured to acquire video in real time; The memory is configured to store a computer program or instructions; The processor is configured to call or execute the program or the instructions stored in the memory, and is further configured to call the camera to implement the steps in the method according to any one of items 1 to 15. A device. (Item 32) A computer-readable storage medium comprising a program, when the program is executed on a computer, the computer is enabled to execute the method according to any one of items 1 to 15. A computer-readable storage medium. (Item 33) A computer program product comprising instructions, wherein when the computer program product is executed on a terminal, the terminal is enabled to execute the method according to any one of Items 1 to 15.

Claims

Claim 1 An image processing method, which is used by a terminal to acquire a video stream in real time, and the image processing method includes: Performing shake correction on the acquired first image to obtain a first cropped area; Performing perspective distortion correction on the first cropped area to obtain a third cropped area; Obtaining a third image based on the first image and the third cropped area; Performing shake correction on the acquired second image to obtain a second cropped area, where the second cropped area is related to the first cropped area and shake information, and the shake information indicates the shake that occurs in the process of the terminal acquiring the first image and the second image; Performing perspective distortion correction on the second cropped area to obtain a fourth cropped area; Obtaining a fourth image based on the second image and the fourth cropped area; and Generating a target video based on the third image and the fourth image wherein the first image and the second image are a pair of original images acquired by the terminal and adjacent to each other in the time domain, Before the step of performing perspective distortion correction on the first cropped area, the image processing method includes: Detecting that the terminal meets the distortion correction condition, and enabling the distortion correction function further comprising, The step of detecting that the terminal meets the distortion correction condition includes: Detecting that the zoom ratio of the terminal is lower than a first preset threshold having, When the zoom ratio ranges from a to 1, the step of the terminal acquiring a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, where the first preset threshold is higher than 1, and a is the fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or When the zoom ratio ranges from 1 to b, the terminal acquires a video stream in real time by using the rear main camera, where b is the fixed zoom ratio of the rear telephoto camera on the terminal or the maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15. Further comprising Method.

2. The direction of the offset of the position of the second cropped area in the second image with respect to the position of the first cropped area in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal acquiring the first image and the second image. The method according to claim 1.

3. The first cropped area indicates a first sub-area in the first image, and the second cropped area indicates a second sub-area in the second image; The similarity between the first image content corresponding to the first sub-area in the first image and the second image content corresponding to the second sub-area in the second image is higher than the similarity between the first image and the second image. The method according to claim 1 or 2.

4. The step of the terminal detecting that it meets the distortion correction condition is: Detecting a first activation operation of the user to activate the perspective distortion correction function, where the first activation operation includes a first operation on a first control in the shooting interface on the terminal, and the first control is used to indicate to activate or deactivate the perspective distortion correction, and the first operation is used to indicate to activate the perspective distortion correction. The method according to claim 1, having.

5. The step of the terminal detecting that it meets the distortion correction condition is: Identifying a human face in a shooting scenario; Identifying a human face in a shooting scenario, where the distance between the human face and the terminal is lower than a preset value; or Identifying a human face in a shooting scenario, where the ratio of the pixels occupied by the human face in the image corresponding to the shooting scenario is higher than a preset ratio. The method according to claim 1, having.

6. Before the step of performing shake correction on the acquired first image, the method comprises: detecting that the terminal satisfies a shake correction condition, and enabling a shake correction function The method according to any one of claims 1 to 5, further comprising.

7. The step of detecting that the terminal satisfies a shake correction condition comprises: detecting that the zoom ratio of the terminal is higher than the fixed zoom ratio of a camera having a minimum zoom ratio on the terminal The method according to claim 6, comprising.

8. The step of detecting that the terminal satisfies a shake correction condition comprises: detecting a second enabling operation of a user for enabling a shake correction function, where the second enabling operation includes a second operation on a second control in a shooting interface on the terminal, the second control is used to indicate enabling or disabling the shake correction function, and the second operation is used to indicate enabling the shake correction function The method according to any one of claims 1 to 7, comprising.

9. The step of acquiring a fourth image based on the second image and the fourth cropped region comprises: acquiring a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates a mapping relationship between each position point in the second cropped region and the corresponding position point in the second image, and the second mapping relationship indicates a mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second cropped region; and acquiring the fourth image based on the second image, the first mapping relationship, and the second mapping relationship The method according to any one of claims 1 to 8, comprising.

10. The step of acquiring the fourth image based on the second image, the first mapping relationship, and the second mapping relationship comprises: combining the first mapping relationship with the second mapping relationship to determine a target mapping relationship, where the target mapping relationship includes a mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second image; and determining the fourth image based on the second image and the target mapping relationship The method according to claim 9, comprising

11. The method according to claim 9 or 10, wherein the first mapping relationship is different from the second mapping relationship.

12. The step of performing perspective distortion correction on the first cropped region comprises: performing optical distortion correction on the first cropped region to obtain a corrected first cropped region; and performing perspective distortion correction on the corrected first cropped region. The step of obtaining a third image based on the first image and the third cropped region comprises: performing optical distortion correction on the third cropped region to obtain a corrected third cropped region; and obtaining the third image based on the first image and the corrected third cropped region. The step of performing perspective distortion correction on the second cropped region comprises: performing optical distortion correction on the second cropped region to obtain a corrected second cropped region; and performing perspective distortion correction on the corrected second cropped region; or The step of obtaining a fourth image based on the second image and the fourth cropped region comprises: performing optical distortion correction on the fourth cropped region to obtain a corrected fourth cropped region; and obtaining the fourth image based on the second image and the corrected fourth cropped region. The method according to any one of claims 1 to 11.

13. An image processing apparatus, which is used by a terminal to acquire a video stream in real time, the image processing apparatus comprising: A hand shake correction module configured to perform hand shake correction on the acquired first image to obtain a first cropped region, and further configured to perform hand shake correction on the acquired second image to obtain a second cropped region, wherein the second cropped region is related to the first cropped region and shake information, the shake information indicating a shake that occurs in the process of the terminal acquiring the first image and the second image, and the first image and the second image being a pair of original images acquired by the terminal and adjacent to each other in the time domain. configured to perform perspective distortion correction on the first cropped area to obtain a third cropped area, and further configured to perform perspective distortion correction on the second cropped area to obtain a fourth cropped area; and an image generation module configured to obtain a third image based on the first image and the third cropped area, obtain a fourth image based on the second image and the fourth cropped area, and further configured to generate a target video based on the third image and the fourth image comprising the image processing apparatus:[[]]END]] a first detection module configured to detect that the terminal meets a distortion correction condition and activate a distortion correction function before the shake correction module performs shake correction on the acquired first image further comprising detecting that the terminal meets a distortion correction condition includes:[[]]END]] detecting that a zoom ratio of the terminal is lower than a first preset threshold having when the zoom ratio ranges from a to 1, the terminal obtaining a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, where the first preset threshold is higher than 1, and a is a fixed zoom ratio of the front wide-angle camera or the rear wide-angle camera, and the value of a ranges from 0.5 to 0.9; or when the zoom ratio ranges from 1 to b, the terminal obtaining a video stream in real time by using a rear main camera, where b is a fixed zoom ratio of a rear telephoto camera on the terminal or a maximum zoom of the terminal, the first preset threshold ranges from 1 to 15, and the value of b ranges from 3 to 15 further comprising device.[[]]END]] [

14. ] The direction of the offset of the position of the second cropped area in the second image with respect to the position of the first cropped area in the first image is opposite to the shake direction on the image plane of the shake that occurs in the process of the terminal acquiring the first image and the second image. The device according to claim 13.[[]]END]] [

15. ] The first cropped area indicates a first sub-area in the first image, and the second cropped area indicates a second sub-area in the second image; The similarity between the first image content corresponding to the first sub-region in the first image and the second image content corresponding to the second sub-region in the second image is higher than the similarity between the first image and the second image. The apparatus according to claim 13 or 14.

16. The apparatus is: A second detection module configured to detect that the terminal satisfies a hand shake correction condition and activate a hand shake correction function before the hand shake correction module performs hand shake correction on the acquired first image. The apparatus according to any one of claims 13 to 15, further comprising.

17. Detecting that the terminal satisfies a hand shake correction condition is: Detecting a second activation operation of the user for activating the hand shake correction function, where the second activation operation includes a second operation on a second control in a shooting interface on the terminal, and the second control is used to indicate enabling or disabling the hand shake correction function, and the second operation is used to indicate enabling the hand shake correction function. The apparatus according to any one of claims 13 to 16, having.

18. The image generation module obtains a first mapping relationship and a second mapping relationship, where the first mapping relationship indicates a mapping relationship between each position point in the second cropped region and the corresponding position point in the second image, and the second mapping relationship indicates a mapping relationship between each position point in the fourth cropped region and the corresponding position point in the second cropped region; and Obtaining the fourth image based on the second image, the first mapping relationship, and the second mapping relationship. The apparatus according to any one of claims 13 to 17, specifically configured to perform.

19. The perspective distortion correction module is specifically configured to perform optical distortion correction on the first cropped region to obtain a corrected first cropped region; and perform perspective distortion correction on the corrected first cropped region. The image generation module is specifically configured to: perform optical distortion correction on the third cropped region to obtain a corrected third cropped region; and obtain the third image based on the first image and the corrected third cropped region; The perspective distortion correction module is specifically configured to: perform optical distortion correction on the second cropped region to obtain a corrected second cropped region; and perform perspective distortion correction on the corrected second cropped region; or The image generation module is specifically configured to: perform optical distortion correction on the fourth cropped region to obtain a corrected fourth cropped region; and obtain the fourth image based on the second image and the corrected fourth cropped region, the apparatus according to any one of claims 13 to 18.

20. An image processing device, the image processing device comprising a processor, a memory, a camera, and a bus, The processor, the memory, and the camera are connected through the bus; The camera is configured to acquire video in real time; The memory is configured to store a computer program or instructions; The processor is further configured to call the camera to implement the steps in the method according to any one of claims 1 to 12 by calling or executing the computer program or the instructions stored in the memory, a device.

21. A computer-readable storage medium comprising a program, wherein when the program is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 12, a computer-readable storage medium.

22. A computer program comprising instructions, wherein when the computer program is executed on a terminal, the terminal is enabled to execute the method according to any one of claims 1 to 12, a computer program.

Citation Information

Patent Citations

  • Imaging apparatus, its control method, program and storage medium

    JP2005195656A

  • Image pickup device and image processing method

    JP2005252626A

  • Image processing device, electronic apparatus, image processing method, and program

    JP2018097870A

  • Video Stabilization

    JP2020520603A

  • Imaging device and operation control method thereof

    WO2013190946A1