An image processing method and related device
By correcting perspective distortion in the cropped area after image stabilization, the problems of image shake and perspective distortion when shooting with a handheld device are solved, improving video quality and avoiding the rolling shutter effect, thus achieving more stable video output.
Patent Information
- Application Number
- CN202110333994.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-29
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-03-29
AI Technical Summary
When users shoot with their handheld devices, image jitter and perspective distortion are caused by unstable posture. Existing technologies have poor perspective distortion correction effects, resulting in low video quality and a tendency for the "jelly effect" to occur.
By correcting perspective distortion in the cropped area after image stabilization, the cropped area size of adjacent frames is made consistent, and the degree of deformation of the same object is kept consistent, avoiding inter-frame differences. The cropped area is used for perspective distortion correction to generate the target video.
It improves video output quality, avoids noticeable jello effect, maintains frame consistency, and enhances video clarity and stability.
Smart Images

Figure CN115131222B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to an image processing method and related device. BACKGROUND
[0002] When the camera is in a wide-angle camera image acquisition mode or a medium-focus camera image acquisition mode, the original image acquired will have obvious perspective distortion. Specifically, perspective distortion (or referred to as 3D distortion) is a 3D object imaging deformation phenomenon (such as horizontal stretching, radial stretching, and a combination of the two) caused by the difference in depth of the photographed scene, which causes different magnification ratios when the spatial three-dimensional shape is mapped to the camera surface, and thus the image needs to be corrected for perspective distortion.
[0003] When a user holds a terminal to take a picture, the user cannot keep the posture of the terminal stable, which causes the terminal to appear to shake in the image plane. That is, the user wants to hold the terminal to take a picture of a target region, but due to the unstable posture of the user holding the terminal, the target region appears to shift greatly between adjacent image frames captured by the camera, which causes the video captured by the camera to appear blurred.
[0004] However, the existing perspective distortion correction technology is not mature, and the quality of the corrected video is low. SUMMARY
[0005] The embodiments of the present application provide an image display method, which does not have obvious Jello phenomenon and improves the quality of video output.
[0006] In a first aspect, the present application provides an image processing method, which is applied to real-time acquisition of a video stream by a terminal; the method comprises:
[0007] performing anti-shake processing on the acquired first image to obtain a first cropped region;
[0008] performing perspective distortion correction processing on the first cropped region to obtain a third cropped region;
[0009] obtaining a third image according to the first image and the third cropped region;
[0010] The second image is subjected to anti-shake processing to obtain a second cropped region; the second cropped region is related to the first cropped region and the shake information; the shake information is used to indicate the shake that occurs in the process of the terminal collecting the first image and the second image; the anti-shake processing is to determine the displacement (including the displacement direction and the displacement distance) of the picture in the framing frame in the previous frame of the original image to the next frame of the original image. In one implementation, the displacement of the above-mentioned offset can be determined based on the shake that occurs in the process of the terminal collecting two adjacent frames of the original image, specifically, the attitude change that occurs in the process of the terminal collecting two adjacent frames of the original image, which can be obtained based on the shooting attitude corresponding to the framing frame in the previous frame of the original image and the attitude of the terminal shooting the current frame of the original image;
[0011] The second cropped region is subjected to perspective distortion correction processing to obtain a fourth cropped region;
[0012] A fourth image is obtained according to the second image and the fourth cropped region; the third image and the fourth image are used to generate a target video;
[0013] The first image and the second image are a pair of time-domain adjacent original images collected by the terminal.
[0014] Unlike the perspective distortion correction of the entire region of the original image in the prior art, in the embodiment of the present application, the perspective distortion correction is performed on the first cropped region. Since the content of two adjacent frames of original images collected by the camera will be dislocated when the terminal device is shaken, the position of the cropped region to be output after the anti-shake processing of the two frames of images in the original image is different, that is, the distance of the cropped region to be output after the anti-shake processing of the two frames of images from the center point of the original image is different, but the size of the cropped region to be output after the anti-shake processing of the two frames is consistent. Therefore, if the perspective distortion correction is performed on the cropped region obtained after the anti-shake processing, the degree of perspective distortion correction performed on adjacent frames is consistent (because the distance of each sub-region in the cropped region to be output after the anti-shake processing of the two frames from the center point of the cropped region to be output after the anti-shake processing is consistent). And since the cropped regions to be output after the anti-shake processing of adjacent frames are images including the same or substantially the same content, the deformation degree of the same object between images is the same, that is, the inter-frame consistency is maintained, and thus the Jello phenomenon is not obvious, and the quality of the video output is improved. It should be understood that the inter-frame consistency can also be referred to as time-domain consistency, which is used to indicate that the regions with the same content in adjacent frames have the same processing result.
[0015] Therefore, in the embodiments of the present application, perspective distortion correction is not performed on the first image, but is performed on the first clipping region obtained through the anti-shake processing, to obtain a third clipping region. For the adjacent frame (the second image) of the first image, the same processing procedure is performed, that is, perspective distortion correction is performed on the second clipping region obtained through the anti-shake processing, to obtain a third clipping region. The deformation correction degree of the same object between the third clipping region and the third clipping region is the same, the inter-frame consistency is maintained, and thus the Jello phenomenon is not obvious, and the quality of the video output is improved.
[0016] In the embodiments of the present application, the clipping regions of the front and back frames can include the same target object (for example, a person object). Since the object of the perspective distortion correction is the clipping region obtained through the anti-shake processing, for the target object, the difference in the deformation degree of the target object caused by the perspective distortion correction of the front and back frames is very small. The very small difference can be understood as being difficult to distinguish from the human eye, or the Jello phenomenon between the front and back frames is difficult to distinguish from the human eye.
[0017] The image processing method of the first aspect can also be understood as follows: for the processing of a single image (for example, the second image) in the collected video stream, the second image can be subjected to anti-shake processing to obtain a second clipping region. The second clipping region is related to the first clipping region and the shake information. The first clipping region is obtained by performing anti-shake processing on an image frame adjacent to the second image and located before the second image in the time domain. The shake information is used to indicate the shake that occurs in the process of collecting the second image and the image frame adjacent to the second image and located before the second image in the time domain. The second clipping region is subjected to perspective distortion correction processing to obtain a fourth clipping region. A fourth image is obtained according to the second image and the fourth clipping region. According to the fourth image, a target video is generated. It should be understood that the above-mentioned second image can refer to any non-first image in the collected video stream. The so-called non-first image refers to an image frame that is not the first in the time domain in the video stream. The above-mentioned image processing method for the second image can be performed on multiple images in the video stream to obtain a target video.
[0018] In a possible implementation, the offset direction of the position of the second clipping region on the second image relative to the position of the first clipping region on the first image is opposite to the shake direction of the shake that occurs in the process of collecting the first image and the second image in the image plane.
[0019] The anti-shake processing is to eliminate the shaking of the collected video caused by the posture change of the user when holding the terminal device. Specifically, the anti-shake processing is to crop the original image collected by the camera to obtain the output in the framing frame, and the output in the framing frame is the anti-shake output. When the terminal has a large posture change in a very short time, the picture in the framing frame of a frame image will have a large offset in the next frame image. For example, when the terminal collects the n+1th frame, compared with the nth frame, the terminal camera has a shaking in the A direction (it should be understood that the shaking in the A direction can mean that the main optical axis of the terminal camera has a shaking in the A direction, or the optical center of the terminal camera has a shaking in the A direction in the image plane), and the picture in the framing frame of the n+1th frame is offset to the opposite direction of the A direction in the nth frame.
[0020] In a possible implementation, the first cropping region is used to indicate a first sub-region in the first image; and the second cropping region is used to indicate a second sub-region in the second image.
[0021] The similarity of the first sub-region corresponding first image content in the first image and the second sub-region corresponding second image content in the second image is greater than the similarity of the first image and the second image. The image content similarity can be understood as the similarity of the scene in the image, the similarity of the background in the image, or the similarity of the target subject in the image.
[0022] In a possible implementation, the similarity between the images can be understood as the overall similarity of the image content in the image region at the same position. Due to the shaking of the terminal, the image content between the first image and the second image is dislocated, so that the same image content is in different positions in the image, that is, the image content similarity of the first image and the second image in the image region at the same position is lower than the image content similarity of the first sub-region and the second sub-region in the image region at the same position.
[0023] In a possible implementation, the similarity between the images can also be understood as the similarity between the image content included in the images. Due to the shaking of the terminal, there is a part of image content in the edge region of the first image that is not in the second image, and similarly, there is a part of image content in the edge region of the second image that is not in the first image, that is, there is image content that is not in each other between the first image and the second image, and the similarity of the image content included in the first image and the second image is lower than the similarity of the image content included in the first sub-region and the second sub-region.
[0024] In addition, if the granularity of similarity determination is a pixel, the similarity can be represented by the overall pixel value similarity of the same position pixels between images.
[0025] In a possible implementation, before the perspective distortion correction processing on the first cropped region, the method further includes:
[0026] The distortion correction function is enabled when it is detected that the terminal meets the distortion correction condition. It can also be understood that the current shooting environment, shooting parameter, or shooting content of the terminal meets the condition that needs to be anti-shake.
[0027] It should be understood that the timing of the action of enabling the distortion correction function when it is detected that the terminal meets the distortion correction condition can be before the step of performing anti-shake processing on the collected first image, or after the anti-shake processing on the collected first image, and before the perspective distortion correction processing on the first cropped region. The present application does not make any limitation.
[0028] In an implementation, the perspective distortion correction needs to be enabled, that is, the perspective distortion correction is performed on the first cropped region only when it is detected that the current perspective distortion correction is in an enabled state.
[0029] In a possible implementation, detecting that the terminal meets the distortion correction condition includes one or more of the following situations, but is not limited to the following situations:
[0030] Situation 1: It is detected that the shooting magnification of the terminal is less than a first preset threshold.
[0031] In an implementation, whether to enable the perspective distortion correction can be determined based on the shooting magnification of the terminal. In the case that the shooting magnification of the terminal is large (for example, in the full or partial magnification range of the telephoto shooting mode, or in the full or partial magnification range of the medium focal length shooting mode), the perspective distortion degree of the image collected by the terminal is low. The so-called low perspective distortion degree can be understood as the case that the human eye almost cannot distinguish the perspective distortion of the image collected by the terminal. Therefore, in the case that the shooting magnification of the terminal is large, the perspective distortion correction does not need to be enabled.
[0032] In a possible implementation, when the shooting magnification is a~1, the terminal collects a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, and the first preset threshold is greater than 1, where a is the inherent magnification of the front wide-angle camera or the rear wide-angle camera, and the value range of a is 0.5~0.9; or,
[0033] When the photographing magnification is 1-b, the terminal adopts the rear main camera to collect a video stream in real time, b is the inherent magnification of a rear long-focus camera included in the terminal or is the maximum zoom value of the terminal, the first preset threshold is 1-15, and the value range of b is 3-15.
[0034] Case 2: A first enabling operation of starting the perspective distortion correction function is detected.
[0035] In a possible implementation, the first enabling operation includes a first operation on a first control on a photographing interface of the terminal, and the first control is used to indicate starting or stopping the perspective distortion correction, and the first operation is used to indicate starting the perspective distortion correction.
[0036] In an implementation, the photographing interface can include a control (referred to as a first control in embodiments of the present application) used to indicate starting or stopping the perspective distortion correction. A user can trigger starting the perspective distortion correction by a first operation on the first control, that is, enable the perspective distortion correction by the first operation on the first control. Then, the terminal can detect a first enabling operation of starting the perspective distortion correction by the user, and the first enabling operation includes the first operation on the first control on the photographing interface of the terminal.
[0037] In an implementation, the control (referred to as a first control in embodiments of the present application) used to indicate starting or stopping the perspective distortion correction is displayed on the photographing interface only when it is detected that the photographing magnification of the terminal is less than a first preset threshold.
[0038] Case 3: It is identified that a face exists in a photographing scene; or,
[0039] It is identified that a face exists in a photographing scene, and a distance between the face and the terminal is less than a preset value; or,
[0040] It is identified that a face exists in a photographing scene, and a pixel proportion of the face in an image corresponding to the photographing scene is greater than a preset proportion.
[0041] It is understood that when a human face exists in the shooting scene, the deformation degree of the human face caused by the perspective distortion correction will be visually more obvious, and the smaller the distance between the human face and the terminal or the larger the pixel proportion of the human face in the image corresponding to the shooting scene, i.e., the larger the area of the human face in the image, the more obvious the deformation degree of the human face caused by the perspective distortion correction will be visually. Therefore, in the above scenario, it is necessary to enable the perspective distortion correction. In this embodiment, whether to enable the perspective distortion correction is determined through the determination of the above conditions related to the human face in the shooting scene, so that it can be accurately judged whether the shooting scene will appear perspective distortion, and the shooting scene that will appear perspective distortion is processed for perspective distortion correction, while the shooting scene that will not appear perspective distortion is not processed for perspective distortion correction, thereby realizing accurate processing of the image signal and saving power consumption.
[0042] It should be understood that the above shooting scene can be understood as an image collected by the camera before the first image is collected, or when the first image is the first frame image in the video collected by the camera, the shooting scene can be the first image or an image close to the first image in the time domain. Whether a human face exists in the shooting scene, the distance between the human face and the terminal, and the pixel proportion of the human face in the image corresponding to the shooting scene can be determined by a neural network or other means. The preset value can be 0-10 m, for example, the preset value can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, which is not limited here. The pixel proportion can be 30%-100%, for example, the pixel proportion can be 30%, 30%, 30%, 30%, 30%, 30%, 30%, 30%, which is not limited here.
[0043] In a possible implementation, before the first image collected is processed for anti-shake, the method further includes:
[0044] The anti-shake processing function is enabled when it is detected that the terminal meets the anti-shake condition. It can also be understood that the current shooting environment, shooting parameter, or shooting content of the terminal meets the condition that needs to be anti-shaken.
[0045] In a possible implementation, detecting that the terminal meets the anti-shake condition includes but is not limited to one or more of the following situations:
[0046] Situation 1: It is detected that the shooting magnification of the terminal is greater than a second preset threshold.
[0047] In an implementation, the difference between the second preset threshold and the inherent magnification of the target camera on the terminal device is within a preset range, where the target camera is the camera with the smallest inherent magnification on the terminal.
[0048] In an implementation, whether to enable the anti-shake processing function can be determined based on the shooting magnification of the terminal. In a case where the shooting magnification of the terminal is small (for example, in a full or partial magnification range of a wide-angle shooting mode), the physical area size corresponding to the image displayed in the shooting interface viewfinder is already small. If anti-shake processing is performed, that is, further image cropping is performed, the physical area size corresponding to the image displayed in the shooting interface viewfinder will be smaller. Therefore, the anti-shake processing function can be enabled only when the shooting magnification of the terminal is greater than a certain preset threshold.
[0049] The target camera can be a front wide-angle camera, and the preset range can be 0-0.3. For example, when the intrinsic magnification of the front wide-angle camera is 0.5, the second preset threshold can be 0.5, 0.6, 0.7, or 0.8.
[0050] Case 2: The second enabling operation of the user to start the anti-shake processing function is detected.
[0051] In an implementation, the second enabling operation includes a second operation on a second control on the shooting interface of the terminal, and the second control is used to indicate starting or stopping the anti-shake processing function, and the second operation is used to indicate starting the anti-shake processing function.
[0052] In an implementation, the shooting interface can include a control (referred to as a second control in embodiments of the present application) used to indicate starting or stopping the anti-shake processing function. The user can trigger starting the anti-shake processing function by a second operation on the second control, that is, enable the anti-shake processing function by the second operation on the second control. Then, the terminal can detect the second enabling operation of the user to start the anti-shake processing function, and the second enabling operation includes the second operation on the second control on the shooting interface of the terminal.
[0053] In an implementation, the control (referred to as a second control in embodiments of the present application) used to indicate starting or stopping the anti-shake processing function is displayed on the shooting interface only when it is detected that the shooting magnification of the terminal is greater than or equal to a second preset threshold.
[0054] In a possible implementation, the fourth image is obtained according to the second image and the fourth cropping region, including: obtaining a first mapping relationship and a second mapping relationship, the first mapping relationship is used to represent the mapping relationship between each position point of the second cropping region and the corresponding position point in the second image, and the second mapping relationship is used to represent the mapping relationship between each position point of the fourth cropping region and the corresponding position point in the second cropping region; the fourth image is obtained according to the second image, the first mapping relationship and the second mapping relationship.
[0055] In a possible implementation, the fourth image is obtained according to the second image, the first mapping relationship, and the second mapping relationship, including:
[0056] The first mapping relationship and the second mapping relationship are coupled to determine a target mapping relationship, the target mapping relationship including a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second image;
[0057] The fourth image is determined according to the second image and the target mapping relationship.
[0058] The first mapping relationship and the second mapping relationship are coupled as index tables. Through the coupling of the index tables, a Warp operation is only needed once after the output is obtained based on the coupled index tables after the perspective distortion correction, instead of once after the cropped region is obtained in the anti-shake processing and once after the perspective distortion correction, thereby reducing the overhead of the Warp operation. The Warp operation can refer to an affine transformation operation on an image, and a specific implementation can refer to an existing Warp technology.
[0059] In a possible implementation, the first mapping relationship is different from the second mapping relationship.
[0060] In a possible implementation, the perspective distortion correction processing on the first cropped region includes: optical distortion correction processing on the first cropped region to obtain a corrected first cropped region; perspective distortion correction processing on the corrected first cropped region; or the like.
[0061] The third image is obtained according to the first image and the third cropped region, including: optical distortion correction processing on the third cropped region to obtain a corrected third cropped region; the third image is obtained according to the first image and the corrected third cropped region; or the like.
[0062] The perspective distortion correction processing on the second cropped region includes: optical distortion correction processing on the second cropped region to obtain a corrected second cropped region; perspective distortion correction processing on the corrected second cropped region; or the like.
[0063] The fourth image is obtained according to the second image and the fourth cropped region, including: optical distortion correction processing on the fourth cropped region to obtain a corrected fourth cropped region; the fourth image is obtained according to the second image and the corrected fourth cropped region.
[0064] Compared with coupling processing of placing the optical distortion correction index table and the perspective distortion index table in the distortion correction module, on one hand, since the optical distortion is an important cause of the jello phenomenon, controlling the effectiveness of both in the video anti-shake backend, the embodiment can solve the problem that the perspective distortion correction module is uncontrollable to the jello phenomenon when the video anti-shake module is closed.
[0065] In a second aspect, the application provides an image processing device, which is applied to a terminal for collecting a video stream in real time; the device comprises:
[0066] an anti-shake processing module, configured to perform anti-shake processing on a collected first image to obtain a first cropping region, and perform anti-shake processing on a collected second image to obtain a second cropping region; the second cropping region is related to the first cropping region and shake information; the shake information is used to indicate shake occurring in a process in which the terminal collects the first image and the second image; the first image and the second image are a pair of original images collected by the terminal in time domain and adjacent in time sequence;
[0067] a perspective distortion correction module, configured to perform perspective distortion correction processing on the first cropping region to obtain a third cropping region, and perform perspective distortion correction processing on the second cropping region to obtain a fourth cropping region;
[0068] an image generation module, configured to obtain a third image according to the first image and the third cropping region, and obtain a fourth image according to the second image and the fourth cropping region, and generate a target video according to the third image and the fourth image.
[0069] In a possible implementation, a position of the second cropping region on the second image is opposite to a shake direction of shake occurring in a process in which the terminal collects the first image and the second image on an image plane.
[0070] In a possible implementation, the first cropping region is used to indicate a first sub-region in the first image; the second cropping region is used to indicate a second sub-region in the second image.
[0071] A similarity between first image content corresponding to the first sub-region in the first image and second image content corresponding to the second sub-region in the second image is greater than a similarity between the first image and the second image.
[0072] In a possible implementation, the device further comprises:
[0073] The first detection module is configured to detect that the terminal satisfies a distortion correction condition before the anti-shake processing module performs anti-shake processing on the collected first image, and enable a distortion correction function.
[0074] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0075] The detection that a shooting magnification of the terminal is less than a first preset threshold.
[0076] In a possible implementation, when the shooting magnification is a~1, the terminal collects a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, the first preset threshold is greater than 1, where a is an intrinsic magnification of the front wide-angle camera or the rear wide-angle camera, and the value range of a is 0.5 to 0.9; or,
[0077] When the shooting magnification is 1 to b, the terminal collects a video stream in real time by using a rear main camera, b is an intrinsic magnification of a rear telephoto camera included in the terminal or is a zoom maximum value of the terminal, the first preset threshold is 1 to 15, and the value range of b is 3 to 15.
[0078] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0079] The detection of a first enabling operation of a user to start a perspective distortion correction function, where the first enabling operation includes a first operation on a first control on a shooting interface of the terminal, the first control is used to indicate starting or stopping perspective distortion correction, and the first operation is used to indicate starting perspective distortion correction.
[0080] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0081] The identification that a face exists in a shooting scene; or,
[0082] The identification that a face exists in a shooting scene and a distance between the face and the terminal is less than a preset value; or,
[0083] The identification that a face exists in a shooting scene and a pixel proportion of the face in an image corresponding to the shooting scene is greater than a preset proportion.
[0084] In a possible implementation, the apparatus further includes:
[0085] The second detection module is configured to detect that the terminal satisfies an anti-shake condition before the anti-shake processing module performs anti-shake processing on the collected first image, and enable an anti-shake processing function.
[0086] In a possible implementation, the detection that the terminal meets the anti-shake condition comprises:
[0087] The detection that the shooting magnification of the terminal is greater than the intrinsic magnification of the camera with the smallest magnification in the terminal.
[0088] In a possible implementation, the detection that the terminal meets the anti-shake condition comprises:
[0089] The detection of a second enabling operation of the user to start the anti-shake processing function, the second enabling operation comprising a second operation on a second control on a shooting interface of the terminal, the second control being used to indicate starting or closing the anti-shake processing function, and the second operation being used to indicate starting the anti-shake processing function.
[0090] In a possible implementation, the image generation module is specifically configured to obtain a first mapping relationship and a second mapping relationship, the first mapping relationship being used to represent a mapping relationship between each position point of the second cropped area and a corresponding position point in the second image, and the second mapping relationship being used to represent a mapping relationship between each position point of the fourth cropped area and a corresponding position point in the second cropped area.
[0091] The fourth image is obtained according to the second image, the first mapping relationship, and the second mapping relationship.
[0092] In a possible implementation, the perspective distortion correction module is specifically configured to determine a target mapping relationship according to the first mapping relationship and the second mapping relationship, the target mapping relationship comprising a mapping relationship between each position point of the fourth cropped area and a corresponding position point in the second image.
[0093] The fourth image is determined according to the target mapping relationship and the second image.
[0094] In a possible implementation, the first mapping relationship is different from the second mapping relationship.
[0095] In a possible implementation, the perspective distortion correction module is specifically configured to perform optical distortion correction processing on the first cropped area to obtain a corrected first cropped area, perform perspective distortion correction processing on the corrected first cropped area, or the like.
[0096] The image generation module is specifically configured to perform optical distortion correction processing on the third cropped area to obtain a corrected third cropped area, obtain a third image according to the first image and the corrected third cropped area, or the like.
[0097] The perspective distortion correction module is specifically configured to perform optical distortion correction processing on the second cropped region to obtain a corrected second cropped region, and perform perspective distortion correction processing on the corrected second cropped region.
[0098] The image generation module is specifically configured to perform optical distortion correction processing on the fourth cropped region to obtain a corrected fourth cropped region, and obtain a fourth image according to the second image and the corrected fourth cropped region.
[0099] In a third aspect, the present application provides an image processing method, which is applied to a terminal for collecting a video stream in real time. The method comprises the following steps:
[0100] determining a target object in a first image collected by the terminal to obtain a first cropped region comprising the target object;
[0101] performing perspective distortion correction processing on the first cropped region to obtain a third cropped region;
[0102] obtaining a third image according to the first image and the third cropped region;
[0103] determining the target object in a second image collected by the terminal to obtain a second cropped region comprising the target object; wherein the second cropped region is related to target movement information, and the target movement information is used to indicate the movement of the target object during the process of collecting the first image and the second image by the terminal;
[0104] performing perspective distortion correction processing on the second cropped region to obtain a third cropped region;
[0105] obtaining a fourth image according to the second image and the third cropped region;
[0106] generating a target video according to the third image and the fourth image;
[0107] wherein the first image and the second image are a pair of original images collected by the terminal in time domain and adjacent in time.
[0108] The image processing method of the third aspect can also be understood as: for processing a single image (e.g., the second image) in the collected video stream, the target object in the collected second image can be determined to obtain a second cropping region including the target object; the second cropping region is related to target movement information, and the target movement information indicates movement of the target object in the process of collecting the second image and an image frame adjacent to the second image and located before the second image in the time domain; the second cropping region is subjected to perspective distortion correction processing to obtain a third cropping region; a fourth image is obtained according to the second image and the third cropping region; and a target video is generated according to the fourth image. It should be understood that the second image can refer to any non-first image in the collected video stream, and the non-first image refers to an image frame that is not the first in the time domain in the video stream. The image processing process of the above-mentioned processing of the second image can obtain a target video.
[0109] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.
[0110] Unlike the perspective distortion correction of the entire region of the original image in the prior implementation, in the embodiment of the present application, the perspective distortion correction is performed on the cropping region. Due to the movement of the target object, the positions of the cropping region including the human object in the front and back image frames are different, which further causes the displacement of the perspective distortion correction of the position points or pixel points in the output two image frames to be different. However, the sizes of the cropping regions determined after the target recognition between the two frames are consistent. Therefore, if the perspective distortion correction is performed on the cropping region, the degree of the perspective distortion correction of the adjacent frames is consistent (because the distances of the position points or pixel points in the two frames relative to the center point of the cropping region are consistent). Moreover, the adjacent two cropping regions are images including the same or substantially the same content, and the deformation degree of the same object between the images is the same, that is, the inter-frame consistency is maintained, and thus the obvious Jello phenomenon does not occur, and the quality of the video output is improved.
[0111] In a possible implementation, before the perspective distortion correction processing on the first cropping region, the method further includes:
[0112] The distortion correction function is enabled when it is detected that the terminal meets the distortion correction condition.
[0113] In a possible implementation, detecting that the terminal meets the distortion correction condition includes, but is not limited to, one or more of the following situations: situation 1: it is detected that the shooting magnification of the terminal is less than a first preset threshold.
[0114] In an implementation, whether to enable the perspective distortion correction can be determined based on a shooting magnification of the terminal. In a case where the shooting magnification of the terminal is large (for example, in a full or partial magnification range of a telephoto shooting mode, or in a full or partial magnification range of a medium focal length shooting mode), the degree of perspective distortion of a picture in an image captured by the terminal is low. The degree of perspective distortion being low can be understood as a case where the human eye cannot almost distinguish that there is perspective distortion in the picture in the image captured by the terminal. Therefore, in the case where the shooting magnification of the terminal is large, the perspective distortion correction does not need to be enabled.
[0115] In a possible implementation, when the shooting magnification is a~1, the terminal uses a front wide-angle camera or a rear wide-angle camera to capture a video stream in real time, the first preset threshold is greater than 1, where the a is an inherent magnification of the front wide-angle camera or the rear wide-angle camera, and the a is in a range of 0.5 to 0.9; or,
[0116] When the shooting magnification is 1~b, the terminal uses a rear main camera to capture a video stream in real time, the b is an inherent magnification of a rear telephoto camera included in the terminal or is a zoom maximum value of the terminal, the first preset threshold is 1 to 15, and the b is in a range of 3 to 15.
[0117] Case 2: A first enabling operation of the user to start the perspective distortion correction function is detected.
[0118] In a possible implementation, the first enabling operation includes a first operation on a first control on a shooting interface of the terminal, the first control being used to indicate starting or stopping the perspective distortion correction, and the first operation being used to indicate starting the perspective distortion correction.
[0119] In an implementation, the shooting interface can include a control (referred to as a first control in embodiments of the present application) used to indicate starting or stopping the perspective distortion correction. The user can trigger starting the perspective distortion correction by means of a first operation on the first control, that is, enable the perspective distortion correction by means of the first operation on the first control. Then, the terminal can detect a first enabling operation of the user to start the perspective distortion correction, the first enabling operation including the first operation on the first control on the shooting interface of the terminal.
[0120] In an implementation, the control (referred to as a first control in embodiments of the present application) used to indicate starting or stopping the perspective distortion correction is displayed on the shooting interface only when it is detected that the shooting magnification of the terminal is less than a first preset threshold.
[0121] Case 3: It is identified that a face exists in a shooting scene; or,
[0122] identifying that a face exists in the shooting scene, and a distance between the face and the terminal is less than a preset value; or
[0123] identifying that a face exists in the shooting scene, and a pixel proportion of the face in an image corresponding to the shooting scene is greater than a preset proportion.
[0124] It is understood that when a face exists in the shooting scene, the deformation degree of the face caused by the perspective distortion correction will be visually more obvious, and the smaller the distance between the face and the terminal or the greater the pixel proportion of the face in the image corresponding to the shooting scene, the greater the area of the face in the image, and the deformation degree of the face caused by the perspective distortion correction will be visually more obvious. Therefore, in the above scenarios, the perspective distortion correction needs to be enabled. In this embodiment, whether to enable the perspective distortion correction is determined through the determination of the above conditions related to the face in the shooting scene, so that the shooting scene in which the perspective distortion will occur can be accurately judged, and the shooting scene in which the perspective distortion will occur is processed by the perspective distortion correction, while the shooting scene in which the perspective distortion will not occur is not processed by the perspective distortion correction, thereby realizing accurate processing of the image signal and saving power consumption.
[0125] In a fourth aspect, the present application provides an image processing device, which is applied to a terminal for collecting a video stream in real time; the device comprises:
[0126] an object determination module, configured to determine a target object in a first image collected by the terminal to obtain a first cropping region comprising the target object, and determine the target object in a second image collected by the terminal to obtain a second cropping region comprising the target object; the second cropping region is related to target movement information, and the target movement information is used to indicate movement of the target object in a process in which the terminal collects the first image and the second image; the first image and the second image are a pair of original images collected by the terminal in time domain and adjacent in time.
[0127] a perspective distortion correction module, configured to perform perspective distortion correction processing on the first cropping region to obtain a third cropping region, and perform perspective distortion correction processing on the second cropping region to obtain a third cropping region.
[0128] an image generation module, configured to obtain a third image according to the first image and the third cropping region, obtain a fourth image according to the second image and the third cropping region, and generate a target video according to the third image and the fourth image.
[0129] In an optional implementation, a position of the target object in the second image is different from a position of the target object in the first image.
[0130] In a possible implementation, the apparatus further includes:
[0131] The first detection module is configured to detect that the terminal satisfies a distortion correction condition before the perspective distortion correction module performs the perspective distortion correction processing on the first clipping area, and enable the distortion correction function.
[0132] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0133] The detection that a shooting magnification of the terminal is less than a first preset threshold.
[0134] In a possible implementation, when the shooting magnification is a~1, the terminal uses a front wide-angle camera or a rear wide-angle camera to collect a video stream in real time, the first preset threshold is greater than 1, where a is an intrinsic magnification of the front wide-angle camera or the rear wide-angle camera, and the value range of a is 0.5-0.9; or,
[0135] When the shooting magnification is 1-b, the terminal uses a rear main camera to collect a video stream in real time, b is an intrinsic magnification of a rear telephoto camera included in the terminal or is a zoom maximum value of the terminal, the first preset threshold is 1-15, and the value range of b is 3-15.
[0136] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0137] The detection of a first enabling operation of a user to start the perspective distortion correction function, where the first enabling operation includes a first operation on a first control on a shooting interface of the terminal, the first control is used to indicate starting or stopping the perspective distortion correction, and the first operation is used to indicate starting the perspective distortion correction.
[0138] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0139] The identification that a face exists in a shooting scene; or,
[0140] The identification that a face exists in a shooting scene, and a distance between the face and the terminal is less than a preset value; or,
[0141] The identification that a face exists in a shooting scene, and a pixel proportion of the face in an image corresponding to the shooting scene is greater than a preset proportion.
[0142] In a fifth aspect, the present application provides an image processing device, including: a processor, a memory, a camera, and a bus, wherein: the processor, the memory, and the camera are connected through the bus.
[0143] the camera, configured to collect video in real time;
[0144] the memory, configured to store computer programs or instructions;
[0145] the processor, configured to call or execute the programs or instructions stored on the memory, and to call the camera to implement the steps of the first aspect and any possible implementation manner of the first aspect, and the steps of the third aspect and any possible implementation manner of the third aspect.
[0146] In a sixth aspect, a computer storage medium is provided, which includes computer instructions, when the computer instructions are run on an electronic device or a server, the steps of the first aspect and any possible implementation manner of the first aspect, and the steps of the third aspect and any possible implementation manner of the third aspect are executed.
[0147] In a seventh aspect, a computer program product is provided, when the computer program product is run on an electronic device or a server, the steps of the first aspect and any possible implementation manner of the first aspect, and the steps of the third aspect and any possible implementation manner of the third aspect are executed.
[0148] In an eighth aspect, a chip system is provided, which includes a processor, configured to support the execution device or the training device to implement the functions involved in the above aspects, for example, to send or process the data involved in the above method; or, information. In a possible design, the chip system further includes a memory, the memory is configured to save the necessary program instructions and data of the execution device or the training device. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0149] The application provides an image display method, which is applied to terminal real-time collection of a video stream; the method comprises the following steps: performing anti-shake processing on a first image collected to obtain a first clipping area; performing perspective distortion correction processing on the first clipping area to obtain a third clipping area; obtaining a third image according to the first image and the third clipping area; performing anti-shake processing on a second image collected to obtain a second clipping area; the second clipping area is related to the first clipping area and shaking information; the shaking information is used for representing shaking occurring in the process that the terminal collects the first image and the second image; performing perspective distortion correction processing on the second clipping area to obtain a fourth clipping area; obtaining a fourth image according to the second image and the fourth clipping area; the third image and the fourth image are used for generating a target video; wherein the first image and the second image are a pair of original images collected by the terminal, which are adjacent in time domain and in sequence; and because two adjacent frames of anti-shake output are images including the same or basically the same content, the deformation degree of the same object between the images is the same, that is, the inter-frame consistency is kept, and therefore, the Jello phenomenon does not occur obviously, and the quality of video display is improved. BRIEF DESCRIPTION OF DRAWINGS
[0150] Figure 1 A structural schematic diagram of a terminal device provided by the embodiment of the application is shown in the figure;
[0151] Figure 2 A software structural block diagram of the terminal device of the embodiment of the application is shown in the figure;
[0152] Figure 3 An embodiment schematic diagram of an image processing method provided by the embodiment of the application is shown in the figure;
[0153] Figure 4a A perspective distortion in the embodiment of the application is shown in the figure;
[0154] Figure 4b A terminal shaking in the embodiment of the application is shown in the figure;
[0155] Figure 4c An anti-shake processing in the embodiment of the application is shown in the figure;
[0156] Figure 4d An anti-shake processing in the embodiment of the application is shown in the figure;
[0157] Figure 4e A Jello phenomenon is shown in the figure;
[0158] Figure 5a An embodiment schematic diagram of an image processing method provided by the embodiment of the application is shown in the figure;
[0159] Figure 5bA schematic of a terminal interface in an embodiment of the application;
[0160] Figure 6 A schematic of a terminal interface in an embodiment of the application;
[0161] Figure 7 A schematic of a terminal interface in an embodiment of the application;
[0162] Figure 8 A schematic of a terminal interface in an embodiment of the application;
[0163] Figure 9 A schematic of a terminal interface in an embodiment of the application;
[0164] Figure 10 A schematic of a terminal interface in an embodiment of the application;
[0165] Figure 11 A schematic of a terminal interface in an embodiment of the application;
[0166] Figure 12 A schematic of a terminal interface in an embodiment of the application;
[0167] Figure 13 A schematic of a terminal interface in an embodiment of the application;
[0168] Figure 14 A schematic of a terminal interface in an embodiment of the application;
[0169] Figure 15 A schematic of a terminal interface in an embodiment of the application;
[0170] Figure 16 A schematic of a terminal interface in an embodiment of the application;
[0171] Figure 17 A schematic of a terminal interface in an embodiment of the application;
[0172] Figure 18 A schematic of a terminal interface in an embodiment of the application;
[0173] Figure 19 A schematic of a terminal interface in an embodiment of the application;
[0174] Figure 20 A schematic of a terminal interface in an embodiment of the application;
[0175] Figure 21 A schematic of a terminal interface in an embodiment of the application;
[0176] Figure 22aA diagram of a Jello phenomenon in an embodiment of the present application;
[0177] Figure 22b A diagram of index table establishment in an embodiment of the present application;
[0178] Figure 22c A diagram of index table establishment in an embodiment of the present application;
[0179] Figure 22d A diagram of index table establishment in an embodiment of the present application;
[0180] Figure 22e A diagram of index table establishment in an embodiment of the present application;
[0181] Figure 22f A diagram of index table establishment in an embodiment of the present application;
[0182] Figure 23 A diagram of index table establishment in an embodiment of the present application;
[0183] Figure 24a An embodiment diagram of an image processing method provided in an embodiment of the present application;
[0184] Figure 24b A diagram of a terminal interface in an embodiment of the present application;
[0185] Figure 25 A diagram of a terminal interface in an embodiment of the present application;
[0186] Figure 26 A diagram of a terminal interface in an embodiment of the present application;
[0187] Figure 27 A diagram of a terminal interface in an embodiment of the present application;
[0188] Figure 28 A diagram of a terminal interface in an embodiment of the present application;
[0189] Figure 29 A diagram of a terminal interface in an embodiment of the present application;
[0190] Figure 30 An embodiment diagram of an image display method provided in an embodiment of the present application;
[0191] Figure 31 A diagram of a terminal interface in an embodiment of the present application;
[0192] Figure 32 A diagram of a terminal interface in an embodiment of the present application;
[0193] Figure 33 A structure diagram of an image display device provided by an embodiment of the present application is shown in FIG. 1.
[0194] Figure 34 A structure diagram of an image display device provided by an embodiment of the present application is shown in FIG. 1.
[0195] Figure 35 A structure diagram of a terminal device provided by an embodiment of the present application is shown in FIG. 1. DETAILED DESCRIPTION
[0196] The embodiments of the present application will be described below in conjunction with the accompanying drawings. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0197] The embodiments of the present application will be described below in conjunction with the accompanying drawings. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0198] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or devices containing a series of units do not have to be limited to those units, but can include other units not clearly listed or inherent to these processes, methods, products or devices.
[0199] For ease of understanding, the structure of the terminal 100 provided by the embodiments of the present application will be described by way of example. Referring to Figure 1 , Figure 1 A structure diagram of a terminal device provided by an embodiment of the present application is shown in FIG. 1.
[0200] As Figure 1As shown, the terminal 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0201] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the terminal 100. In other embodiments of the present application, the terminal 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0202] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.
[0203] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0204] The processor 110 can also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has recently used or is likely to use again. If the processor 110 needs to use the instructions or data again, it can be retrieved directly from the memory. This avoids repeated accesses and reduces the latency of the processor 110, thus improving the efficiency of the system.
[0205] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0206] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example, the processor 110 can be coupled to the touch sensor 180K through an I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface to realize the touch function of the terminal 100.
[0207] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple sets of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to realize communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the I2S interface to realize the function of answering a phone through a Bluetooth headset.
[0208] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 can be coupled with the wireless communication module 160 through a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface, realizing the function of answering a phone call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0209] The UART interface is a universal serial bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is usually used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface, realizing the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface, realizing the function of playing music through a Bluetooth headset.
[0210] The MIPI interface can be used to connect the processor 110 and peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface, realizing the shooting function of the terminal 100. The processor 110 and the display screen 194 communicate through the DSI interface, realizing the display function of the terminal 100.
[0211] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 and the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0212] Specifically, the video (including a sequence of image frames, for example, including the first image and the second image in the present application) collected by the camera 193 can be but not limited to transmitted to the processor 110 through the interface (such as the CSI interface or the GPIO interface) described above for connecting the camera 193 and the processor 110.
[0213] The processor 110 can acquire instructions from the memory, and perform video processing (such as the anti-shake processing, the perspective distortion correction processing, etc. in the present application) on the video collected by the camera 193 based on the acquired instructions, to obtain a processed image (such as the third image and the fourth image in the present application).
[0214] The processor 110 can, but is not limited to, deliver the processed image to the display screen 194 through the interface (such as the DSI interface or the GPIO interface) described above for connecting the display screen 194 and the processor 110, and then the display screen 194 can perform video display.
[0215] The USB interface 130 is an interface conforming to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal 100, and can also be used to transmit data between the terminal 100 and a peripheral device. It can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as AR devices, etc.
[0216] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative, and does not constitute a structural limitation on the terminal 100. In other embodiments of the present application, the terminal 100 can also use different interface connection methods or combinations of multiple interface connection methods in the above embodiments.
[0217] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through the wireless charging coil of the terminal 100. The charging management module 140 can charge the battery 142 while also supplying power to electronic devices through the power management module 141.
[0218] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), etc. In other embodiments, the power management module 141 can also be arranged in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be arranged in the same device.
[0219] The wireless communication function of the terminal 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0220] The antenna 1 and the antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the terminal 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.
[0221] The mobile communication module 150 can provide a solution including 2G / 3G / 4G / 5G wireless communication applied to the terminal 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed electromagnetic waves to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated by the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be arranged in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 and at least part of the modules of the processor 110 can be arranged in the same device.
[0222] The modem processor can include a modulator and a demodulator. The modulator is used to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the loudspeaker 170A, the microphone 170B, etc.), or displays an image or a video through the display screen 194. In some embodiments, the modem processor can be an independent device. In some other embodiments, the modem processor can be independent of the processor 110, and arranged in the same device as the mobile communication module 150 or other functional modules.
[0223] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the terminal 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, frequency-modulate them, amplify them, and radiate them as electromagnetic waves via the antenna 2.
[0224] In some embodiments, antenna 1 and mobile communication module 150 of terminal 100 are coupled, and antenna 2 and wireless communication module 160 are coupled, so that terminal 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0225] Terminal 100 implements display functions through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 can include one or more GPUs that execute program instructions to generate or change display information. Specifically, one or more GPUs in processor 110 can implement rendering tasks of images (such as rendering tasks related to images that need to be displayed in the present application) and deliver the rendering results to the application processor or other display drivers, which trigger the display screen 194 to display videos.
[0226] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), or the like. In some embodiments, the terminal 100 can include one or N display screens 194, where N is a positive integer greater than 1. The display screen 194 can display a target video in the embodiments of the present application. In one implementation, the terminal 100 can run an application related to photographing. When the terminal opens the application related to photographing, the display screen 194 can display a shooting interface. The shooting interface can include a viewfinder, and the target video can be displayed in the viewfinder.
[0227] The terminal 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor.
[0228] The ISP is configured to process the data fed back by the camera 193. For example, when photographing, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert the electrical signal into an image visible to the naked eye. The ISP can also optimize the algorithm for the noise, brightness, and skin color of the image. The ISP can also optimize the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be arranged in the camera 193.
[0229] The camera 193 is configured to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to an ISP to be converted into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into an image signal in a standard format, such as RGB, YUV, or the like. In some embodiments, the terminal 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0230] After the DSP converts the digital image signal into an image signal in a standard format, such as RGB, YUV, or the like, the processor 110 can perform further image processing on the original image (e.g., the first image and the second image in the embodiments of the present application). The image processing includes, but is not limited to, anti-shake processing, perspective distortion correction processing, optical distortion correction processing, and cropping processing to adapt to the size of the display screen 194. The processed image (e.g., the third image and the fourth image in the embodiments of the present application) can be displayed in the viewfinder frame of the shooting interface displayed on the display screen 194.
[0231] In the embodiments of the present application, the number of cameras 193 in the terminal 100 can be at least two. For example, there can be two cameras, one of which is a front camera and the other of which is a rear camera; for example, there can be three cameras, one of which is a front camera and the other two of which are rear cameras; for example, there can be four cameras, one of which is a front camera and the other three of which are rear cameras. It should be noted that the camera 193 can be one or more of a wide-angle camera, a main camera, or a telephoto camera.
[0232] For example, in the case of two cameras, the front camera can be a wide-angle camera, and the rear camera can be a main camera. In this case, the field of view of the image captured by the rear camera is large, and the image information is rich.
[0233] For example, in the case of three cameras, the front camera can be a wide-angle camera, and the rear cameras can be a wide-angle camera and a main camera.
[0234] For example, in the case of four cameras, the front camera can be a wide-angle camera, and the rear cameras can be a wide-angle camera, a main camera, and a telephoto camera.
[0235] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the terminal 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0236] The video codec is used to compress or decompress digital video. The terminal 100 can support one or more video codecs. In this way, the terminal 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0237] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by drawing on the structure of a biological neural network, such as drawing on the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, intelligent cognitive applications of the terminal 100 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.
[0238] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions. For example, music, video, etc. Files are saved in the external memory card.
[0239] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during the use of the terminal 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various function applications and data processing of the terminal 100 by running instructions stored in the internal memory 121 and / or instructions stored in the memory arranged in the processor.
[0240] The terminal 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.
[0241] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some of the functions of the audio module 170 can be disposed in the processor 110.
[0242] The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The terminal 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0243] The receiver 170B, also referred to as a "earpiece", is configured to convert an audio electrical signal into a sound signal. When the terminal 100 answers a call or a voice message, the user can listen to the voice by holding the receiver 170B close to the ear.
[0244] The microphone 170C, also referred to as a "microphone", "sound collector", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can make a sound by holding the mouth close to the microphone 170C, and input the sound signal into the microphone 170C. The terminal 100 can be provided with at least one microphone 170C. In other embodiments, the terminal 100 can be provided with two microphones 170C, in addition to collecting sound signals, noise reduction functions can also be achieved. In other embodiments, the terminal 100 can also be provided with three, four or more microphones 170C, to achieve the functions of collecting sound signals, noise reduction, and identifying the source of the sound, and to achieve the functions of directional recording, etc.
[0245] The earphone interface 170D is configured to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0246] The pressure sensor 180A is configured to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. The pressure sensor 180A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The terminal 100 determines the intensity of the force according to the change in capacitance. When a touch operation is applied to the display screen 194, the terminal 100 detects the intensity of the touch operation according to the pressure sensor 180A. The terminal 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view short messages is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed.
[0247] The shooting interface displayed by the display screen 194 can include a first control and a second control. The first control is configured to turn on or turn off the anti-shake processing function, and the second control is configured to turn on or turn off the perspective distortion correction function. For example, the user can perform an enabling operation on the display screen to turn on the anti-shake processing function. The enabling operation can be a click operation on the first control. The terminal 100 can determine, according to the detection signal of the pressure sensor 180A, that the click position on the display screen is the position of the first control, and then generate an operation instruction to turn on the anti-shake processing function, and then enable the anti-shake function according to the operation instruction to turn on the anti-shake processing function. For example, the user can perform an enabling operation on the display screen to turn on the perspective distortion correction function. The enabling operation can be a click operation on the second control. The terminal 100 can determine, according to the detection signal of the pressure sensor 180A, that the click position on the display screen is the position of the second control, and then generate an operation instruction to turn on the perspective distortion correction function, and then enable the perspective distortion correction function according to the operation instruction to turn on the perspective distortion correction function.
[0248] The gyro sensor 180B can be used to determine the motion posture of the terminal 100. In some embodiments, the angular velocity of the terminal 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can be used for anti-shake photography. For example, when the shutter is pressed, the gyro sensor 180B detects the angle of shaking of the terminal 100, calculates the distance that the lens module needs to compensate according to the angle, and lets the lens offset the shaking of the terminal 100 by reverse movement to achieve anti-shake. The gyro sensor 180B can also be used for navigation and motion sensing game scenarios.
[0249] The barometric sensor 180C is used to measure air pressure. In some embodiments, the terminal 100 calculates the altitude, assists in positioning and navigation by using the air pressure value measured by the barometric sensor 180C.
[0250] The magnetic sensor 180D includes a Hall sensor. The terminal 100 can detect the opening and closing of a flip cover with the magnetic sensor 180D. In some embodiments, when the terminal 100 is a flip phone, the terminal 100 can detect the opening and closing of the flip cover according to the magnetic sensor 180D. In turn, according to the detected opening and closing state of the cover or the opening and closing state of the flip cover, the terminal 100 can set features such as automatic unlocking of the flip cover.
[0251] The acceleration sensor 180E can detect the magnitude of acceleration of the terminal 100 in various directions (typically three axes). When the terminal 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device and applied to landscape / portrait screen switching, pedometers, and other applications.
[0252] The distance sensor 180F is used to measure distance. The terminal 100 can measure distance by infrared or laser. In some embodiments, in a shooting scenario, the terminal 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0253] The proximity light sensor 180G can include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The terminal 100 emits infrared light outwardly through the light-emitting diode. The terminal 100 detects infrared reflected light from nearby objects using the photodiode. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 100. When insufficient reflected light is detected, the terminal 100 can determine that there is no object near the terminal 100. The terminal 100 can use the proximity light sensor 180G to detect that the user is holding the terminal 100 close to the ear for a call, so as to automatically turn off the screen to achieve power saving. The proximity light sensor 180G can also be used for automatic unlocking and locking of the cover in cover mode and pocket mode.
[0254] Ambient light sensor 180L is used to sense ambient light brightness. Terminal 100 can adaptively adjust display screen 194 brightness according to sensed ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also cooperate with proximity light sensor 180G to detect whether terminal 100 is in a pocket to prevent accidental touch.
[0255] Fingerprint sensor 180H is used to collect a fingerprint. Terminal 100 can use collected fingerprint characteristics to implement fingerprint unlocking, access application lock, take photos with fingerprint, answer incoming calls with fingerprint, and the like.
[0256] Temperature sensor 180J is used to detect temperature. In some embodiments, terminal 100 uses temperature detected by temperature sensor 180J to implement temperature processing strategies. For example, when temperature reported by temperature sensor 180J exceeds a threshold, terminal 100 implements performance reduction of a processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In another embodiment, when temperature is lower than another threshold, terminal 100 heats battery 142 to avoid abnormal shutdown of terminal 100 caused by low temperature. In yet another embodiment, when temperature is lower than yet another threshold, terminal 100 implements voltage boost of output voltage of battery 142 to avoid abnormal shutdown caused by low temperature.
[0257] Touch sensor 180K, also referred to as a "touch device". Touch sensor 180K can be disposed on display screen 194, and touch sensor 180K and display screen 194 form a touch screen, also referred to as a "touch panel". Touch sensor 180K is used to detect a touch operation acting on or near it. Touch sensor 180K can transmit the detected touch operation to an application processor to determine a touch event type. Visual output related to the touch operation can be provided through display screen 194. In another embodiment, touch sensor 180K can also be disposed on a surface of terminal 100, which is different from the position of display screen 194.
[0258] Bone conduction sensor 180M can obtain a vibration signal. In some embodiments, bone conduction sensor 180M can obtain a vibration signal of a human body sound part vibration bone block. Bone conduction sensor 180M can also contact a human body pulse to receive a blood pressure pulsation signal. In some embodiments, bone conduction sensor 180M can also be disposed in a headset to form a bone conduction headset. Audio module 170 can analyze a voice signal based on the vibration signal of the sound part vibration bone block obtained by bone conduction sensor 180M to implement a voice function. Application processor can analyze heart rate information based on the blood pressure pulsation signal obtained by bone conduction sensor 180M to implement a heart rate detection function.
[0259] The keys 190 include a power key, a volume key, and the like. The keys 190 can be mechanical keys. Alternatively, the keys 190 can be touch keys. The terminal 100 can receive key input and generate key signal input related to user settings and function control of the terminal 100.
[0260] The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts and touch vibration feedback. For example, touch operations for different applications (e.g., taking pictures, playing audio, and the like) can correspond to different vibration feedback effects. Touch operations on different regions of the display screen 194 can also correspond to different vibration feedback effects. Different application scenarios (e.g., time reminders, received messages, alarms, games, and the like) can also correspond to different vibration feedback effects. The touch vibration feedback effects can also be customizable.
[0261] The indicator 192 can be an indicator light that can be used to indicate a charging state, a power change, and can also be used to indicate messages, missed calls, notifications, and the like.
[0262] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the terminal 100. The terminal 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. The same SIM card interface 195 can simultaneously insert multiple cards. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external storage cards. The terminal 100 interacts with a network through a SIM card to achieve functions such as calls and data communication. In some embodiments, the terminal 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 100 and cannot be separated from the terminal 100.
[0263] The software system of the terminal 100 can use a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. Embodiments of the present disclosure exemplarily illustrate the software structure of the terminal 100 using an Android system with a layered architecture as an example.
[0264] Figure 2 is a software structure block diagram of the terminal 100 of the present disclosure.
[0265] A layered architecture divides software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.
[0266] The application layer can include a series of application packages.
[0267] As shown in Figure 2 , the application packages can include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0268] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.
[0269] As shown in Figure 2 , the application framework layer can include window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0270] The window manager is used to manage window programs. The window manager can get the size of the display screen, determine whether there is a status bar, lock the screen, and take screenshots, etc.
[0271] The content provider is used to store and obtain data, and make the data accessible to the application. The data can include video, image, audio, dialed and received calls, browsing history and bookmarks, phonebook, etc.
[0272] The view system includes visual controls, such as controls that display text, controls that display pictures, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface that includes a short message notification icon can include a view that displays text and a view that displays a picture.
[0273] The phone manager is used to provide the communication function of the terminal 100. For example, the management of the call state (including call connection, call hang-up, etc.).
[0274] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, etc.
[0275] The notification manager enables applications to display notification information in the status bar, which can be used to convey a message of the notification type, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify the completion of the download, message reminders, etc. The notification manager can also be a notification in the form of a chart or a scroll bar text appearing in the system top status bar, such as a notification of the background running application, and can also be a notification in the form of a dialog window appearing on the screen. For example, the text information in the status bar prompts, the prompt sound, the vibration of the electronic device, the flashing of the indicator light, etc.
[0276] The Android Runtime includes a core library and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0277] The core library contains two parts: one part is the function function that the java language needs to call, and the other part is the core library of Android.
[0278] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform the management of the object life cycle, the stack management, the thread management, the security and the exception management, and the garbage collection function.
[0279] The system library can include multiple functional modules. For example: surface manager, media library, three-dimensional graphics processing library (for example: OpenGL ES), 2D graphics engine (for example: SGL) and the like.
[0280] The surface manager is used to manage the display subsystem, and provides a fusion of 2D and 3D layers for multiple applications.
[0281] The media library supports multiple commonly used audio, video format playback and recording, and static image files and the like. The media library can support multiple audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG and the like.
[0282] The three-dimensional graphics processing library is used to realize three-dimensional graphics drawing, image rendering, synthesis, and layer processing and the like.
[0283] The 2D graphics engine is a drawing engine for 2D drawing.
[0284] The kernel layer is the layer between hardware and software. The kernel layer at least contains display driver, camera driver, audio driver, sensor driver.
[0285] The working flow of the terminal 100 software and hardware is exemplarily explained below in combination with the photographing scene.
[0286] When the touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, a timestamp of the touch operation, and the like). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer, and identifies a control corresponding to the input event. Taking an example in which the touch operation is a touch single-click operation, and the control corresponding to the single-click operation is a control of the camera application icon, the camera application invokes an interface of the application framework layer, starts the camera application, and then starts the camera driver by invoking the kernel layer, captures a still image or a video by the camera 193, and the captured video can be the first image and the second image in the embodiment of the present application.
[0287] For ease of understanding, the image processing method provided in the embodiment of the present application is described in detail in combination with the accompanying drawings and application scenarios.
[0288] The embodiment of the present application can be applied in instant photographing, video post-processing, target tracking, and the like, which are described as follows.
[0289] Example 1: Instant photographing
[0290] In the embodiment of the present application, in the scene of instant photographing or shooting of the terminal, the camera of the terminal can collect a video stream in real time, and display a preview picture generated based on the video stream collected by the camera on a shooting interface.
[0291] The video stream collected by the camera can include a first image and a second image, the first image and the second image being two frame raw images adjacent in time domain and collected by the terminal, specifically, the video collected by the terminal in real time is a raw image sequence arranged in sequence in time domain, the second image being a raw image arranged adjacent to and after the first image in time domain, for example, the video collected by the terminal in real time includes a 0th image, a 1st image, a 2nd image, …, an Xth image arranged in time domain, the first image being an nth image, the second image being an n+1th image, n being an integer greater than or equal to 0 and less than X.
[0292] The first image and the second image are raw images collected by the camera of the terminal, and the concept of the raw image is introduced as follows.
[0293] When the terminal is taking a photo, the shutter can be opened, and then light can be transmitted to the photosensitive element of the camera through the lens. The photosensitive element of the camera can convert the light signal into an electrical signal, and transmit the electrical signal to a processor such as an image signal processor (ISP) and a digital signal processor (DSP) for conversion into an image. The image can be referred to as a raw image captured by the camera. The first image and the second image described in the embodiments of the present application can be the raw image. The image processing method provided in the embodiments of the present application can be used to perform image processing on the raw image to obtain a target image (the third image and the fourth image) after processing, and then the preview screen including the third image and the fourth image can be displayed in the framing box 402. It should be understood that the first image and the second image can also be images obtained after size cropping of the images obtained after ISP and DSP processing. The size cropping can be performed to adapt to the size of the display screen of the terminal or the size of the framing box of the shooting interface.
[0294] In the embodiments of the present application, the video stream captured by the camera in real time can be subjected to anti-shake processing and perspective distortion correction. Next, taking the first image and the second image in the embodiments of the present application as an example, how to perform anti-shake processing and perspective distortion correction on the first image and the second image is described.
[0295] In some scenarios, the user can take a target area by holding the terminal device, where the target area can include an object in a moving state, or the target area can be a region in a stationary state, i.e., does not include an object in a moving state.
[0296] When the video stream is subjected to anti-shake processing and perspective distortion correction, the preview screen displayed in the framing box can exhibit a jelly effect (or referred to as Jello phenomenon). The Jello phenomenon refers to the deformation and change of the screen like jelly. Next, the causes of the jelly jello effect are described.
[0297] First, the perspective distortion of an image and the perspective distortion correction are introduced.
[0298] When the camera is in a wide-angle shooting mode or a main camera shooting mode, the raw image captured will have obvious perspective distortion. Specifically, perspective distortion (or referred to as 3D distortion) is caused by the difference in depth of the shooting scene, which causes the 3D object imaging deformation phenomenon (such as horizontal stretching, radial stretching, and a combination of the two) when the spatial three-dimensional shape is mapped to the surface of the camera. Therefore, perspective distortion correction of the image is needed.
[0299] It should be understood that the so-called 3D object imaging deformation can be understood as that the object on the image has image distortion caused by deformation compared with the real object observed by human eyes visually.
[0300] In the image with perspective distortion, the object farther from the image center position has more obvious perspective distortion than the object closer to the image center position, that is, the length of imaging deformation is higher (the stretching in the horizontal direction, the stretching in the radial direction and the combination of the two is greater). Therefore, in order to obtain an image without perspective distortion problem, in the process of correcting the perspective distortion of the image, the displacement to be corrected is different for the position points or pixel points at different distances from the image center position, and the displacement to be corrected is greater for the position points or pixel points farther from the image center position.
[0301] Referring to Figure 4a , Figure 4a Fig. 1 is a schematic diagram of a perspective distortion image, as Figure 4a shown, the person in the edge region of the image has a certain 3D deformation.
[0302] In some scenarios, the user may appear to shake during the process of photographing by holding the terminal device.
[0303] When the user holds the terminal to photograph, the user cannot keep the posture of the terminal always in a stable state, which causes the terminal to appear to shake in the image plane direction (that is, the terminal has a large posture change in a very short time). The image plane can be the plane where the imaging surface of the camera is located, specifically, the plane where the light-sensing surface of the light-sensing element of the camera is located. From the perspective of the user, the image plane can be a plane substantially identical or completely identical to the plane where the display screen is located. The shaking in the image plane direction can include shaking in various directions on the image plane, such as left, right, up, down, left-up, left-down, right-up or right-down, and the like.
[0304] When the user holds the terminal to photograph a target region, if the user holds the terminal to appear to shake in the A direction, the image content including the target region in the original image captured by the terminal camera will shift to the opposite direction of the A direction. For example, when the user holds the terminal to appear to shake in the left-up direction, the image content including the target region in the original image captured by the terminal camera will shift to the opposite direction of the A direction.
[0305] It should be understood that the shaking of the terminal in the A direction as described above can be understood as that the terminal appears to shake in the A direction from the perspective of the user holding the terminal.
[0306] Referring to Figure 4b , Figure 4b Fig. 2 is a schematic diagram of a terminal shaking scenario, as Figure 4bAs shown, at time A1, the user holds the terminal device to take a picture of object 1. At this time, the shooting interface can display object 1. Between times A1 and A2, the user holds the terminal device and shakes to the lower left. At this time, object 1 in the original image captured by the camera shifts to the upper right.
[0307] When a user holds a handheld device to photograph a target area, the unstable posture of the handheld device causes a significant shift in the target area between adjacent image frames captured by the camera (in other words, a significant content misalignment between two adjacent image frames). When the device shakes, this content misalignment between adjacent image frames typically results in blurry video. It should be understood that the misalignment is more severe in images captured by the main camera when the device shakes compared to a wide-angle camera, and even more severe in images captured by a telephoto camera when the device shakes compared to the main camera.
[0308] When performing image stabilization, two adjacent frames can be registered and the edge portion can be cropped to obtain a clear image (which can be called image stabilization output or image stabilization viewfinder, such as the first cropping area and the third cropping area in the embodiment of this application).
[0309] Image stabilization is used to eliminate video shake caused by changes in posture when a user holds a handheld device. Specifically, image stabilization involves cropping the original image captured by the camera to obtain the output within the viewfinder. This viewfinder output is the stabilized output. When the device experiences a significant posture change within a short period, the image within the viewfinder of one frame will shift considerably in the next frame. For example, if the device experiences a shake in direction A in the image plane direction when capturing frame n+1 compared to frame n (this could mean the camera's main optical axis shifts in the direction A, or the camera's optical center shifts in the direction A in the image plane direction), then the image within the viewfinder of frame n+1 will be shifted in the opposite direction of A in frame n.
[0310] by Figure 4b For example, between times A1 and A2, the user's handheld terminal device shakes to the lower left. At this time, object 1 in the original image captured by the camera shifts to the upper right. Therefore, the cropping area required for image stabilization is (…). Figure 4b The dashed box shown also shifts to the upper right accordingly.
[0311] The anti-shake processing is to determine the displacement (including the displacement direction and the displacement distance) of the picture in the framing frame in the previous frame of the original image offset in the next frame of the original image. In one implementation, the displacement of the above-mentioned offset can be determined based on the shaking of the terminal in the process of collecting the adjacent two frames of the original image. Specifically, it can be the attitude change of the terminal in the process of collecting the adjacent two frames of the original image. The attitude change can be obtained based on the shooting attitude corresponding to the framing frame in the previous frame of the original image and the attitude of the terminal shooting the current frame of the original image.
[0312] After the anti-shake processing, when the user holds the terminal device and a large degree of shaking occurs, although there is a large difference between the pictures of the adjacent frames of the original image, there will be no large difference or no difference between the pictures in the framing frame obtained by cropping the adjacent frames of the original image, and thus the display picture in the framing frame of the shooting interface will not appear a large shaking and blur.
[0313] Next, a specific process of anti-shake processing is given in combination with Figure 4c , and referring to Figure 4c , for the n-th frame of the original image collected by the camera, the attitude parameter of the handheld terminal when shooting the n-th frame of the image can be obtained. The handheld terminal attitude parameter can indicate the attitude of the terminal when collecting the n-th frame, and the handheld terminal attitude parameter can be obtained based on the information collected by the sensor carried by the terminal. The sensor carried by the terminal can be but is not limited to a gyroscope, and the collected information can be but is not limited to the rotation angular velocity of the terminal on the x-axis, the rotation angular velocity on the y-axis and the rotation angular velocity on the z-axis.
[0314] In addition, the stable attitude parameter of the previous frame of the image, that is, the attitude parameter of the n-1-th frame of the image, can also be obtained. The stable attitude parameter of the previous frame of the image can refer to the shooting attitude corresponding to the framing frame in the n-1-th frame of the image. Based on the stable attitude parameter of the previous frame of the image and the attitude parameter of the handheld terminal when shooting the n-th frame of the image, the shaking direction and amplitude (the shaking direction and amplitude can also be referred to as the shaking displacement) of the framing frame in the original image on the image plane can be calculated. The size of the framing frame is pre-set, and when n is 0, that is, when the first frame of the image is collected, the position of the framing frame can be located at the center position of the original image. After obtaining the shaking direction and amplitude of the framing frame in the original image on the image plane, the position of the framing frame in the n-th frame of the original image can be determined, and it is determined whether the framing frame boundary exceeds the range of the input image (that is, the n-th frame of the original image). If the framing frame boundary exceeds the range of the input image due to the large shaking amplitude, the stable attitude parameter of the n-th frame of the image is adjusted so that the boundary of the framing frame is located within the boundary of the input original image. For example Figure 4bAs shown, if the viewfinder boundary extends beyond the input image in the upper right direction, the stabilization pose parameters of the nth frame image are adjusted to move the viewfinder position to the lower left, thus ensuring the viewfinder boundary is within the boundary of the original input image. A stabilization index table is then output, specifically indicating the pose correction displacement of each pixel or position point in the image after shake removal. The viewfinder output can be determined based on the output stabilization index table and the nth frame original image. If the shake is small or nonexistent, ensuring the viewfinder boundary does not extend beyond the input image, the stabilization index table can be directly output.
[0315] Similarly, when performing image stabilization on the (n+1)th frame, the pose parameters of the nth frame can be obtained. These pose parameters refer to the shooting pose corresponding to the viewfinder in the nth frame. Based on the pose parameters of the nth frame and the pose parameters of the handheld terminal in the (n+1)th frame, the shake direction and amplitude (also known as shake displacement) of the viewfinder in the image plane can be calculated. The size of the viewfinder is pre-set. Then, the position of the viewfinder in the (n+1)th frame can be determined, and it can be checked whether the viewfinder boundary exceeds the range of the input image (i.e., the (n+1)th frame). If the viewfinder boundary exceeds the range of the input image due to excessive shake amplitude, the stabilization pose parameters of the (n+1)th frame are adjusted so that the viewfinder boundary is within the boundary of the input image, and the image stabilization index table is output. If the shake amplitude is small or nonexistent, and the viewfinder boundary does not exceed the range of the input image, the image stabilization index table can be directly output.
[0316] To more clearly describe the above image stabilization process, we will use n=0, that is, take the acquisition of the first image frame as an example, to describe the image stabilization process.
[0317] Reference Figure 4d , Figure 4d This is a schematic diagram of a stabilization process in an embodiment of this application, such as... Figure 4d As shown, the center portion of the image in frame 0 is cropped to serve as the stabilized output for frame 0 (i.e., the viewfinder output for frame 0). Later, due to hand-held shaking, the user's hand position changes, causing the video footage to shake. The viewfinder boundary of frame 1 can be calculated based on the handheld hand position parameters in frame 1 and the stable position parameters of the image in frame 0. Figure 4dAs shown, the position of the framing box of the first frame of the original image is determined to be offset to the right, and the framing box of the first frame is output. Assuming that there is a very large jitter when the second frame of the image is captured, the framing box boundary of the second frame is calculated according to the pose parameters of the second frame of the handheld terminal and the pose parameters after the pose correction of the first frame of the image, and the framing box boundary of the second frame exceeds the boundary of the input original image of the second frame, the stable pose parameters are adjusted so that the framing box boundary in the original image of the second frame is located within the boundary of the original image of the second frame, and then the framing box of the second frame of the original image is output.
[0318] In some shooting scenes, the anti-shake processing and perspective distortion correction of the collected original image can cause the jello effect in the picture, and the reasons are as follows:
[0319] In some existing implementations, in order to perform perspective distortion correction and anti-shake processing on the image at the same time, the anti-shake framing box needs to be confirmed on the original image, and the perspective distortion correction needs to be performed on the entire region of the original image. Therefore, the image region in the anti-shake framing box is equivalent to being subjected to perspective distortion correction, and the image region in the anti-shake framing box is output. However, the above-mentioned method has the following problems: Because the adjacent two frames of images captured by the camera will usually be misaligned in content when the terminal device is jittering, the positions of the image regions to be output after the anti-shake processing of the two frames of images in the original image are different, that is, the distances of the image regions to be output after the anti-shake processing of the two frames of images from the center point of the original image are different, and then the perspective distortion correction displacements of the position points or pixel points in the output two frames of images are different. Since the adjacent two frames of anti-shake output images include the same or substantially the same content, the deformation degrees of the same objects between the images are different, that is, the inter-frame consistency is lost, and then the jello phenomenon occurs, which reduces the quality of the video output.
[0320] More specifically, the causes of the jello phenomenon can be as Figure 4e As shown, there is obvious jitter between the first frame and the second frame, and the anti-shake framing box is moved from the center of the picture to the edge of the picture. The content in the small box is the anti-shake output, and the large box is the original image. When the position of the checkerboard content in the original image changes, the correction displacement of the perspective distortion correction is different, so that the grid spacing of the front and rear frames changes differently, and then the anti-shake output of the two frames of images will observe the obvious grid stretching phenomenon (that is, the jello phenomenon).
[0321] The embodiments of the present application provide an image processing method, which can perform anti-shake processing and perspective distortion correction on the image while ensuring inter-frame consistency, avoiding the jello phenomenon between images, or reducing the deformation degree of the jello phenomenon between images.
[0322] Next, it is described in detail how to perform the image stabilization and perspective distortion correction while ensuring the inter-frame consistency. Embodiments of the present application take the first image and the second image as an example to illustrate the image stabilization and perspective distortion correction, wherein the first image and the second image are two original images adjacent in time domain and collected by the terminal, specifically, the video collected by the terminal in real time is a sequence of original images arranged in sequence in time domain, the second image is an original image arranged adjacent to and after the first image in time domain, for example, the video collected by the terminal in real time includes the 0th image, the 1st image, the 2nd image, …, the Xth image arranged in time domain, the first image is the nth image, the second image is the n+1th image, and n is an integer greater than or equal to 0 and less than X.
[0323] In embodiments of the present application, the first image collected can be subjected to the image stabilization to obtain the first cropping region, wherein the first image is an original image in a video collected by the terminal camera in real time, and the process of how the camera collects the original image can refer to the description in the above embodiments, which will not be repeated here.
[0324] In one implementation, the first image is the first original image in the video collected by the terminal camera in real time, and the first cropping region is located in the central region of the first image, which can refer to the description in the above embodiments about how to perform the image stabilization on the nth image, which will not be repeated here.
[0325] In one implementation, the first image is an original image after the first original image in the video collected by the terminal camera in real time, and the first cropping region is determined based on the shaking occurred by the terminal in the process of collecting the original image adjacent to and before the first image, which can refer to the description in the above embodiments about how to perform the image stabilization on the nth (n≠0) image, which will not be repeated here.
[0326] In one implementation, the first image can be subjected to the image stabilization by default, which means that the image stabilization is directly performed on the collected original image without enabling the image stabilization function, for example, the image stabilization is performed on the collected video when the shooting interface is opened. The way to enable the image stabilization function can be based on the user's enabling operation, or by analyzing the shooting parameters of the current shooting scene and the content of the collected video to determine whether to enable the image stabilization function.
[0327] In one implementation, the terminal can be detected to satisfy the image stabilization condition to enable the image stabilization function, and then the first image can be subjected to the image stabilization.
[0328] In the specific implementation process, the terminal satisfying the image stabilization condition includes but is not limited to one or more of the following situations:
[0329] Case 1: it is detected that the photographing magnification of the terminal is greater than a second preset threshold.
[0330] In an implementation, whether to enable the anti-shake processing function can be determined based on the photographing magnification of the terminal. In a case where the photographing magnification of the terminal is small (for example, in a full or partial magnification range of a wide-angle photographing mode), the physical area size corresponding to the picture displayed in the viewfinder frame of the photographing interface is already small. If the anti-shake processing is still performed, that is, further image cropping is performed, the physical area size corresponding to the picture displayed in the viewfinder frame will be smaller. Therefore, the anti-shake processing function can be enabled only when the photographing magnification of the terminal is greater than a certain preset threshold.
[0331] Specifically, before the first image is subjected to the anti-shake processing, it needs to be detected that the photographing magnification of the terminal is greater than or equal to a second preset threshold, which is the minimum magnification value at which the anti-shake processing function can be enabled. When the photographing magnification of the terminal is a~1, the terminal uses a front wide-angle camera or a rear wide-angle camera to collect a video stream in real time, where a is the inherent magnification of the front wide-angle camera or the rear wide-angle camera, and the value range of a is 0.5~0.9. Therefore, the second preset threshold can be a magnification value greater than or equal to a and less than or equal to 1, for example, the second preset threshold can be a. In this case, the anti-shake processing function is not enabled only when the photographing magnification of the terminal is equal to a. For example, the second preset threshold can be 0.6, 0.7 or 0.8, which is not limited here.
[0332] Case 2: it is detected that a second enabling operation of the user to enable the anti-shake processing function.
[0333] In an implementation, the photographing interface can include a control (referred to as a second control in embodiments of the present application) for indicating whether to enable or disable the anti-shake processing function. The user can trigger the enabling of the anti-shake processing function by means of a second operation on the second control, that is, the anti-shake processing function is enabled by means of the second operation on the second control. Then, the terminal can detect the second enabling operation of the user to enable the anti-shake processing function, which includes the second operation on the second control on the photographing interface of the terminal.
[0334] In an implementation, the control (referred to as a second control in embodiments of the present application) for indicating whether to enable or disable the anti-shake processing function is displayed on the photographing interface only when it is detected that the photographing magnification of the terminal is greater than or equal to the second preset threshold. In an implementation, the first cropped area can also be subjected to optical distortion correction processing to obtain a corrected first cropped area.
[0335] In an implementation, the first cropped region can be subjected to the perspective distortion correction by default, by which it is meant that the perspective distortion correction is directly performed on the first cropped region without the need to enable the perspective distortion correction, for example, the perspective distortion correction is performed on the first cropped region after the anti-shake processing when the shooting interface is opened. The manner of enabling the perspective distortion correction can be based on a user's enabling operation or by analyzing the shooting parameters of the current shooting scene and the content of the collected video to determine whether to enable the perspective distortion correction.
[0336] In an implementation, it can be detected that the terminal satisfies the distortion correction condition, the distortion correction function is enabled, and then the perspective distortion correction is performed on the first cropped region.
[0337] It should be understood that the timing of performing the action of detecting that the terminal satisfies the distortion correction condition and enabling the distortion correction function can be before the step of performing the anti-shake processing on the collected first image, or after the anti-shake processing on the collected first image and before the perspective distortion correction processing on the first cropped region. The present application does not make any limitation.
[0338] In the specific implementation process, detecting that the terminal satisfies the distortion correction condition includes but is not limited to one or more of the following situations:
[0339] Situation 1: It is detected that the shooting magnification of the terminal is less than a first preset threshold.
[0340] In an implementation, whether to enable the perspective distortion correction can be determined based on the shooting magnification of the terminal. In the case that the shooting magnification of the terminal is large (for example, in the full or partial magnification range of the telephoto shooting mode, or in the full or partial magnification range of the medium focal length shooting mode), the degree of perspective distortion of the image collected by the terminal is low, by which it is meant that the human eye can hardly distinguish that the image collected by the terminal has perspective distortion. Therefore, in the case that the shooting magnification of the terminal is large, the perspective distortion correction does not need to be enabled.
[0341] Specifically, before the first cropped region is subjected to the perspective distortion correction by default, it is necessary to detect that the shooting magnification of the terminal is less than a first preset threshold, and the first preset threshold is the maximum magnification value at which the perspective distortion correction can be enabled. That is, when the shooting magnification of the terminal is less than the first preset threshold, the perspective distortion correction can be enabled.
[0342] When the shooting magnification of the terminal is a1, the terminal uses a front wide-angle camera or a rear wide-angle camera to collect a video stream in real time, wherein the shooting magnification range corresponding to the image captured by the wide-angle camera can be a1, the minimum value a can be the inherent magnification of the wide-angle camera, the value range of a can be 0.50.9, for example, a can be but not limited to 0.5, 0.6, 0.7, 0.8 or 0.9. It should be understood that the wide-angle camera can include a front wide-angle camera and a rear wide-angle camera. The shooting magnification range corresponding to the front wide-angle camera and the rear wide-angle camera can be the same or different, for example, the shooting magnification range corresponding to the front wide-angle camera can be a11, the minimum value a1 can be the inherent magnification of the front wide-angle camera, the value range of a1 can be 0.50.9, for example, a1 can be but not limited to 0.5, 0.6, 0.7, 0.8 or 0.9, for example, the shooting magnification range corresponding to the rear wide-angle camera can be a21, the minimum value a2 can be the inherent magnification of the rear wide-angle camera, the value range of a2 can be 0.50.9, for example, a2 can be but not limited to 0.5, 0.6, 0.7, 0.8 or 0.9. In the case that the terminal uses a front wide-angle camera or a rear wide-angle camera to collect a video stream in real time, the first preset threshold is greater than 1, for example, the first preset threshold can be 2, 3, 4, 5, 6, 7 or 8, which is not limited here.
[0343] When the shooting magnification of the terminal is 1b, the terminal uses a rear main camera to collect a video stream in real time, in the case that the terminal is also integrated with a rear long-focus camera, b can be the inherent magnification of the rear long-focus camera, the value range of b can be 315, for example, b can be but not limited to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15. In the case that the terminal is not integrated with a rear long-focus camera, the value b can be the maximum zoom value of the terminal, the value range of b can be 315, for example, b can be but not limited to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15. In the case that the terminal uses a rear main camera to collect a video stream in real time, the first preset threshold is 115, for example, the first preset threshold can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15, which is not limited here.
[0344] It should be understood that the range of the above-mentioned first preset threshold (115) can include two end point magnifications (1 and 15), or can not include the two end point magnifications (1 and 15), or include one of the two end point magnifications.
[0345] For example, the range of the first preset threshold can be the interval [1, 15].
[0346] Exemplarily, the first preset threshold value can range from 1 to 15.
[0347] Exemplarily, the first preset threshold value can range from 1 to 15.
[0348] Exemplarily, the first preset threshold value can range from 1 to 15.
[0349] Case 2: detecting a first enabling operation of the user to start the perspective distortion correction function.
[0350] In an implementation, the shooting interface can include a control (referred to as a first control in the embodiments of the present application) for indicating the start or stop of the perspective distortion correction. The user can trigger the start of the perspective distortion correction by a first operation on the first control, i.e., enable the perspective distortion correction by the first operation on the first control. Then the terminal can detect a first enabling operation of the user to start the perspective distortion correction, and the first enabling operation includes the first operation on the first control on the shooting interface of the terminal.
[0351] In an implementation, the control (referred to as a first control in the embodiments of the present application) for indicating the start or stop of the perspective distortion correction is displayed on the shooting interface only when it is detected that the shooting magnification of the terminal is less than a first preset threshold value.
[0352] Case 3: determining to enable the perspective distortion correction based on recognizing that there is a face in the shooting scene; or, recognizing that there is a face in the shooting scene and the distance between the face and the terminal is less than a preset value; or, recognizing that there is a face in the shooting scene and the pixel proportion of the face in the image corresponding to the shooting scene is greater than a preset proportion.
[0353] It should be understood that when there is a face in the shooting scene, the deformation degree of the face caused by the perspective distortion correction will be more visually obvious, and the smaller the distance between the face and the terminal or the greater the pixel proportion of the face in the image corresponding to the shooting scene, i.e., the larger the area of the face in the image, the more visually obvious the deformation degree of the face caused by the perspective distortion correction. Therefore, in the above scenarios, the perspective distortion correction needs to be enabled. In the embodiments, whether to enable the perspective distortion correction is determined by the above conditions related to the face in the shooting scene, which can accurately determine the shooting scene in which the perspective distortion will occur and perform the perspective distortion correction on the shooting scene in which the perspective distortion will occur, and does not perform the perspective distortion correction on the shooting scene in which the perspective distortion will not occur, thereby achieving accurate processing of the image signal and saving power consumption.
[0354] In the above manner, when the anti-shake processing function and the perspective distortion correction function are enabled, the following steps 301 to 306 can be performed, and details can be referred to the above embodiments. Figure 3 , Figure 3 An embodiment of the image processing method provided in the embodiments of the present application is shown in the following embodiment schematic diagram. Figure 3 As shown in the above embodiment schematic diagram, the image processing method provided in the present application includes the following steps.
[0355] 301. Perform anti-shake processing on the collected first image to obtain a first cropping region.
[0356] In the embodiments of the present application, the first image can be collected and anti-shake processing can be performed on the first image to obtain a first cropping region, wherein the first image is an original image in a video collected by a terminal camera in real time, and the process of how the camera collects the original image can be referred to the description in the above embodiments, which will not be repeated here.
[0357] In one implementation, the first image is the first frame of the original image in the video collected by the terminal camera in real time, and the first cropping region is located in the center region of the first image, and details can be referred to the description of how to perform anti-shake processing on the 0th frame in the above embodiments, which will not be repeated here.
[0358] In one implementation, the first image is an original image after the first frame of the original image in the video collected by the terminal camera in real time, and the first cropping region is determined based on the shaking occurred in the process of collecting the first image and the original image adjacent to and before the first image, and details can be referred to the description of how to perform anti-shake processing on the nth (n is not equal to 0) frame in the above embodiments, which will not be repeated here.
[0359] In one implementation, the first cropping region can also be subjected to optical distortion correction processing to obtain a corrected first cropping region.
[0360] 302. Perform perspective distortion correction processing on the first cropping region to obtain a third cropping region.
[0361] In the embodiments of the present application, after the first cropping region is obtained, perspective distortion correction processing can be performed on the first cropping region to obtain a third cropping region.
[0362] It should be understood that the above shooting scene can be understood as an image collected by the camera before the first image is collected, or when the first image is the first frame image in the video collected by the camera, the shooting scene can be the first image or an image close to the first image in the time domain. Whether the shooting scene contains a face, the distance between the face and the terminal, and the pixel ratio of the face in the image corresponding to the shooting scene can be determined by a neural network or other means. The preset value can be 0-10m, for example, the preset value can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, which is not limited here. The pixel ratio can be 30%-100%, for example, the pixel ratio can be 30%, 30%, 30%, 30%, 30%, 30%, 30%, 30%, which is not limited here.
[0363] The above describes how to enable perspective distortion correction, and the following describes how to perform perspective distortion correction:
[0364] Unlike the existing implementation of performing perspective distortion correction on the entire region of the original image, in the embodiment of the present application, perspective distortion correction is performed on the first cropped region. Since the content of the adjacent two frames of original images collected by the camera will usually be misaligned when the terminal device is shaking, the position of the cropped region to be output after the two frames of images are processed for anti-shake is different in the original image, that is, the distance of the cropped region to be output after the two frames of images are processed for anti-shake from the center point of the original image is different. However, the size of the cropped region to be output after the two frames are processed for anti-shake is consistent, so if perspective distortion correction is performed on the cropped region obtained after anti-shake processing, the degree of perspective distortion correction performed on adjacent frames is consistent (because the distance of each sub-region in the cropped region output after anti-shake of the two frames from the center point of the cropped region output after anti-shake is consistent). And since the adjacent two frames of anti-shake output cropped regions are images containing the same or substantially the same content, the deformation degree of the same object between images is the same, that is, the inter-frame consistency is maintained, thereby avoiding the Jello phenomenon and improving the quality of the video output. It should be understood that the inter-frame consistency can also be referred to as time domain consistency, which is used to represent that the regions with the same content in adjacent frames have the same processing result.
[0365] Therefore, in the embodiments of the present application, the perspective distortion correction is not performed on the first image, but is performed on the first clipping region obtained through the anti-shake processing, to obtain a third clipping region. For the adjacent frame (the second image) of the first image, the same processing procedure is performed, that is, the perspective distortion correction is performed on the second clipping region obtained through the anti-shake processing, to obtain a third clipping region. The deformation correction degree of the same object between the third clipping region and the third clipping region is the same, the inter-frame consistency is maintained, and thus the obvious Jello phenomenon is avoided, and the quality of the video output is improved.
[0366] In the embodiments of the present application, the clipping regions of the front and back frames can include the same target object (for example, a person object). Since the object of the perspective distortion correction is the clipping region obtained through the anti-shake processing, for the target object, the front and back frames make the difference in the deformation degree of the target object very small through the perspective distortion correction. The very small difference can be understood as being difficult to be distinguished by the naked eye, or the Jello phenomenon is difficult to be distinguished by the naked eye between the front and back frames.
[0367] Specifically, refer to Figure 22a The output obtained after the perspective distortion correction is performed on the whole region of the original image has the obvious Jello phenomenon, and the output obtained after the perspective distortion correction is performed on the clipping region obtained through the anti-shake processing does not have the obvious Jello phenomenon.
[0368] In one implementation, the third clipping region can be further subjected to optical distortion correction processing to obtain a corrected third clipping region.
[0369] 303. Obtain a third image according to the first image and the third clipping region.
[0370] In the embodiments of the present application, the third clipping region can indicate the output of the perspective distortion correction. Specifically, the third clipping region can be used to indicate how to clip from the first image and the displacement of the pixel points or position points, to obtain the image region that needs to be output.
[0371] In one implementation, the third clipping region can be used to indicate which pixel points or position points in the first image need to be mapped or selected after the anti-shake processing and the perspective distortion correction, and how the pixel points or position points that need to be mapped or selected in the first image need to be displaced.
[0372] Specifically, the third clipping region can include each pixel point or position point in the output of the perspective distortion correction, and a mapping relationship between the pixel points or position points in the first image. Further, the third image can be obtained according to the first image and the third clipping region, and the third image can be displayed in the viewfinder of the shooting interface. It should be understood that the third image can be obtained according to the first image and the third clipping region by a warp operation, wherein the warp operation refers to affine transformation of the image.
[0373] 304, performing anti-shake processing on the collected second image to obtain a second clipping region; and the second clipping region is related to the first clipping region and the shake information; the shake information is used to indicate the shake of the terminal in the process of collecting the first image and the second image.
[0374] The shake information indicated by the shake information can be the shake of the terminal in a time period from a certain time point before the terminal collects the first image to a time point when the terminal collects the second image, a time point after the terminal collects the first image (also before the time point when the terminal collects the second image) to the time point when the terminal collects the second image, a time period from the time point after the terminal collects the first image to a certain time point after the time point when the terminal collects the second image, a time period from the time point after the terminal collects the first image to a certain time point before the time point when the terminal collects the second image, a time period from the time point when the terminal collects the first image to a certain time point before the time point when the terminal collects the second image, a time period from the time point when the terminal collects the first image to a certain time point after the time point when the terminal collects the second image, or a time period from the time point when the terminal collects the first image to a certain time point after the terminal collects a certain image before collecting the first image, a time point after the terminal collects the second image or a certain time point after the terminal collects a certain image after collecting the second image.
[0375] Specifically, the shake information indicated by the shake information can be the shake of the terminal in a time period from a certain time point before the terminal collects the first image to a time point when the terminal collects the second image, a time point after the terminal collects the first image (also before the time point when the terminal collects the second image) to the time point when the terminal collects the second image, a time period from the time point after the terminal collects the first image to a certain time point after the time point when the terminal collects the second image, a time period from the time point after the terminal collects the first image to a certain time point before the time point when the terminal collects the second image, a time period from the time point after the terminal collects the first image to a certain time point before the time point when the terminal collects the second image, a time period from the time point after the terminal collects the first image to a certain time point after the time point when the terminal collects the second image, or a time period from the time point when the terminal collects the first image to a certain time point after the terminal collects a certain image before collecting the first image, a time point after the terminal collects the second image or a certain time point after the terminal collects a certain image after collecting the second image.
[0376] In one implementation, a direction of shift of the position of the second cropped region on the second image relative to the position of the first cropped region on the first image is opposite to a direction of shake in an image plane of shake occurred in the process of the terminal capturing the first image and the second image.
[0377] More description about step 304 can refer to the description about how to perform the anti-shake processing on the captured first image in the above-mentioned embodiments, which will not be repeated here.
[0378] It should be understood that there is no strict time sequence limitation between step 304 and step 302 and step 303, step 304 can be performed before step 302, after step 301, step 304 can also be performed before step 303, after step 302, step 304 can also be performed after step 303, and the present application does not limit.
[0379] 305. performing perspective distortion correction processing on the second cropped region to obtain a fourth cropped region.
[0380] In the embodiments of the present application, the second cropped region can be subjected to distortion correction processing to obtain a second mapping relationship, the second mapping relationship including the mapping relationship between each position point of the fourth cropped region and the corresponding position point in the second cropped region; the fourth cropped region is determined according to the second cropped region and the second mapping relationship.
[0381] In the embodiments of the present application, the first mapping relationship including the mapping relationship between each position point of the second cropped region and the corresponding position point in the second image can be determined according to the shake information and the second image, then the second cropped region can be determined according to the first mapping relationship and the second image; the second mapping relationship including the mapping relationship between each position point of the fourth cropped region and the corresponding position point in the second cropped region can be obtained by performing distortion correction processing on the second cropped region; and the fourth cropped region is determined according to the second cropped region and the second mapping relationship.
[0382] In an implementation, the first mapping relationship and the second mapping relationship can also be coupled to determine a target mapping relationship, the target mapping relationship including a mapping relationship between each position point of the fourth clipping region and a corresponding position point in the second image; and the fourth image is determined according to the target mapping relationship and the second image, which is equivalent to coupling the first mapping relationship and the second mapping relationship by an index table. Through the coupling of the index table, it is not necessary to perform a warp operation after obtaining the clipping region in the anti-shake processing, and then perform a warp operation after the perspective distortion correction, but only need to perform a warp operation after obtaining the output based on the coupled index table after the perspective distortion correction, thereby reducing the overhead of the warp operation.
[0383] wherein the first mapping relationship can be referred to as an anti-shake index table, and the second mapping relationship can be referred to as a perspective distortion correction index table. Next, how to determine the correction index table (or referred to as the second mapping relationship) used for perspective distortion correction is described in detail: first, how to determine the mapping relationship from the second clipping region output by the anti-shake to the fourth clipping region is described as follows:
[0384] x = x0(1 + K 01 *r x 2 + K 02 *r x 4 )
[0385] y = y0(1 + K 11 *r y 2 + K 12 *r y 4 );
[0386] wherein (x0, y0) is the normalized coordinate of a pixel point or a position point in the second clipping region output by the anti-shake, and the relationship between the pixel coordinate (xi, yi), the image width W, and the image height H can be as follows:
[0387] x0 = (xi - 0.5 * W) / W;
[0388] y0 = (yi - 0.5 * H) / H;
[0389] wherein (x, y) is the normalized coordinate of a corresponding position point or a pixel point in the fourth clipping region, which needs to be converted into a target pixel coordinate (xo, yo) according to the following formula:
[0390] xo = x * W + 0.5 * W;
[0391] yo = y * H + 0.5 * H;
[0392] In the formula, K01, K02, K11, K12 are relational parameters, and the specific values are determined by the field angle of the distorted image; r_x is the distortion distance in the x direction of (x0, y0), and r_y is the distortion distance in the y direction of (x0, y0), which can be determined by the following relational expression:
[0393]
[0394] In the formula, α_x = 1, α_y, β_x, β_y are relational coefficients, and the value range is 0.0-1.0. When the values are 1.0, 1.0, 1.0, 1.0, r_x = r_y is the distance from the pixel point (x0, y0) to the center of the image; when the values are 1.0, 0.0, 0.0, 1.0, r_x is the distance from the pixel point (x0, y0) to the x-direction center axis of the image, and r_y is the distance from the pixel point (x0, y0) to the y-direction center axis of the image. The coefficients can be adjusted according to the perspective distortion correction effect and the bending degree of the background straight line to achieve the correction effect after balancing the two.
[0395] Next, how to perform the anti-shake processing and perspective distortion correction on the first image based on the index table (the first mapping relationship and the second mapping relationship) is described. Figure 22b
[0396] In one implementation, the optical distortion correction index table T1 can be determined according to the imaging module related parameters of the camera, wherein the optical distortion correction index table T1 includes the mapping points of the optical distortion correction image position to the first image. The granularity of the optical distortion correction index table T1 can be a pixel or a position point. Then, the anti-shake index table T2 can be generated, which includes the mapping points of the anti-shake compensated image position to the first image. Then, the optical + anti-shake index table T3 can be generated, which can include the mapping points of the optical distortion correction + anti-shake compensated image position to the first image (it should be understood that the above optical distortion correction processing is optional, that is, the optical distortion correction processing can not be performed, and the first image can be directly subjected to anti-shake processing, and then the anti-shake index table T2 can be obtained); then, based on the anti-shake output image information (including the optical distortion correction + anti-shake compensated image position), the perspective distortion correction index table T4 is determined, which can include the mapping points of the fourth clipping region to the second clipping region of the anti-shake output. Based on T3 and T4, the coupling index table T5 can be generated by the coordinate system switching lookup table method, referring to Figure 22c, specifically, the mapping relationship between the fourth clipping region and the first image can be established by reverse twice lookup from the fourth clipping region: taking a position point a in the fourth clipping region as an example, for the currently looked-up position point a in the fourth clipping region, the mapping position b of the point a in the anti-shake output coordinate system can be indexed according to the perspective distortion correction table T4; then in the optical distortion + anti-shake output index table T3, the optical distortion correction + anti-shake correction displacement delta1, delta2, …, deltan of the neighboring grid points c0, c1, …, cn of the point b and the coordinates d0, d1, …, dn in the first image can be known according to the mapping relationship; finally, the optical + anti-shake correction displacement of the point b is calculated based on delta1, delta2, …, deltan by using an interpolation algorithm, and then the mapping position d of the point b in the first image is known; thus the mapping relationship between the position point a in the fourth clipping region and the mapping position d of the first image is established. For each position point in the fourth clipping region, the above process is performed to obtain the coupling index table T5.
[0397] Then the image deformation interpolation algorithm can be performed according to the coupling index table T5 and the second image to output the fourth image.
[0398] Taking the processing of the second image as an example, the specific flowchart can refer to Figure 22e , Figure 22f and Figure 22c , wherein Figure 22e describes how to construct the index table and the process of warp based on the anti-shake, perspective distortion and optical distortion coupling index table T5 and the second image, Figure 22f , Figure 22c and Figure 22c are different from Figure 22e , in which the anti-shake index table and the optical distortion correction index table are coupled, Figure 22f , in which the perspective distortion correction index table and the optical distortion correction index table are coupled, Figure 22d , in which all distortion correction processes (including optical distortion correction and perspective distortion correction) are coupled in the distortion correction module of the video anti-shake backend. Figure 23 The process of coordinate system switching lookup table is shown. It should be understood that in one implementation, the perspective distortion index table and the anti-shake index table can be coupled first, and then the optical distortion index table is coupled, and the specific process is as follows: based on the anti-shake index table T2 and the perspective distortion index table T3, the anti-shake and perspective distortion coupling index table T4 is generated by the coordinate system switching lookup table; then starting from the perspective distortion correction output coordinate, the mapping relationship between the perspective distortion correction output and the anti-shake input coordinate is established by reverse twice lookup, specifically, refer to Figure 5aFor the current search point a in the perspective distortion correction table T3, its mapping position b in the anti-shake output (perspective correction input) coordinate system can be indexed; then in the anti-shake output index table T2, the mapping relationship between the neighborhood grid points c0, c1, …, cn of point b and the anti-shake input coordinates d0, d1, …, dn can be used to know their anti-shake correction displacements delta1', delta2', …, deltan'; after the anti-shake correction displacement of point b is calculated by using the interpolation algorithm, its mapping position d in the anti-shake input coordinate system is known; thus the mapping relationship between the perspective distortion correction output point a and the anti-shake input coordinate point d is established. For each point of the perspective distortion correction table T3, the above process is performed to obtain the coupling index table T4 of perspective distortion correction and anti-shake.
[0399] Then the three-coupling index table T5 of anti-shake, perspective distortion and optical distortion can be generated again by the coordinate system switching lookup table method based on T1 and T4. Starting from the fourth clipping region, reverse twice lookup table is used to establish the coordinate mapping relationship between the fourth clipping region and the second image: specifically, for the current search position point a in the fourth clipping region, the mapping position d of the position point a in the anti-shake input (optical distortion correction output) coordinate system can be indexed according to the coupling index table T4 of perspective distortion correction and anti-shake; then in the optical distortion index table T1, the mapping relationship between the neighborhood grid points d0, d1, …, dn of point d and the original image coordinates e0, e1, …, en can be used to know their optical distortion correction displacements delta1'', delta2'', …, deltan''; after the optical distortion correction displacement of point d is calculated by using the interpolation algorithm, its mapping position e in the first image coordinate system is known; thus the mapping relationship between the position point a and the mapping position e of the second image is established. For each position point of the fourth clipping region, the above process is performed to obtain the three-coupling index table T5.
[0400] Compared with the coupling processing of the optical distortion correction index table and the perspective distortion index table in the distortion correction module, on the one hand, since the optical distortion is an important cause of the jello phenomenon, the effectiveness of the two is controlled in the video anti-shake backend, and the embodiment can solve the problem that the perspective distortion correction module cannot control the jello phenomenon when the video anti-shake module closes the optical distortion correction.
[0401] In an implementation, in order to make the front and rear frames keep strict inter-frame consistency, the same correction index table (or referred to as the second mapping relationship) can be used for the perspective distortion correction of the front and rear frames, so that the front and rear frames keep strict inter-frame consistency.
[0402] More descriptions about step 305 can refer to the above descriptions about how to perform the perspective distortion correction processing on the first clipping region in the embodiments, and similar parts will not be described here.
[0403] 306、obtaining a fourth image according to the second image and the fourth cropped region.
[0404] In an implementation, a first mapping relationship and a second mapping relationship can be obtained, the first mapping relationship is used to represent a mapping relationship between each position point of the second cropped region and a corresponding position point in the second image, and the second mapping relationship is used to represent a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second cropped region; the fourth image is obtained according to the second image, the first mapping relationship and the second mapping relationship.
[0405] More description about step 306 can refer to the description about obtaining a third image according to the first image and the third cropped region in the above-mentioned embodiments, which will not be repeated here.
[0406] 307、generating a target video according to the third image and the fourth image.
[0407] In an implementation, a target video can be generated according to the third image and the fourth image, and the target video can be displayed. Specifically, for any two images in a video stream collected by a terminal device in real time, an output image processed by anti-shake and perspective distortion correction can be obtained based on steps 301 to 306, and then a video processed by anti-shake and perspective distortion correction (i.e., the above-mentioned target video) can be obtained, for example, the first image can be the nth frame in the video stream, and the second image can be the n+1th frame, the n-2th frame and the n-1th frame can also be used as the first image and the second image respectively for processing in steps 301 to 306, the n-1th frame and the nth frame can also be used as the first image and the second image respectively for processing in steps 301 to 306, and the n+1th frame and the n+2th frame can also be used as the first image and the second image respectively for processing in steps 301 to 306. This is not repeated and exhausted here.
[0408] In an implementation, after obtaining the target video, the target video can be displayed in a viewfinder frame of a shooting interface of the camera.
[0409] It should be understood that the above-mentioned shooting interface can be a preview interface in a shooting mode, and the target video is a preview picture in the shooting mode, and the user can trigger a camera shutter by clicking a shooting control of the shooting interface or by other triggering manners, and then a shooting image can be obtained.
[0410] It should be understood that the above shooting interface can be a preview interface before starting video recording in a video recording mode, and the target video is a preview picture before starting video recording in the video recording mode. The user can click a shooting control of the shooting interface or trigger the camera to start recording by other triggering manners.
[0411] It should be understood that the above shooting interface can be a preview interface after starting video recording in a video recording mode, and the target video is a preview picture after starting video recording in the video recording mode. The user can click a shooting control of the shooting interface or trigger the camera to end recording by other triggering manners, and save the video recorded during recording.
[0412] The present application provides an image display method, applied to a terminal for real-time acquisition of a video stream; comprising: performing anti-shake processing on a first image acquired to obtain a first cropping region; performing perspective distortion correction processing on the first cropping region to obtain a third cropping region; obtaining a third image according to the first image and the third cropping region; performing anti-shake processing on a second image acquired to obtain a second cropping region; and the second cropping region is related to the first cropping region and shake information; the shake information is used to represent the shake occurring in the process of the terminal acquiring the first image and the second image; performing perspective distortion correction processing on the second cropping region to obtain a fourth cropping region; obtaining a fourth image according to the second image and the fourth cropping region; the third image and the fourth image are used to generate a target video; wherein the first image and the second image are a pair of original images acquired by the terminal in time domain and adjacent in time. And since the adjacent two frames of anti-shake output are images including the same or basically the same content, the deformation degree of the same object between the images is the same, that is, the inter-frame consistency is maintained, and thus the obvious Jello phenomenon does not occur, and the quality of video display is improved.
[0413] Next, the image processing method provided by the embodiment of the present application will be described in combination with the interaction on the terminal side.
[0414] Reference Figure 5a , Figure 5a The flow of the image processing method in the embodiment of the present application is shown in FIG. 1, and the image processing method provided by the embodiment of the present application comprises: Next, the process that the terminal starts the camera application and opens the shooting interface is described first in combination with the interaction between the terminal and the user.
[0415] 501, start the camera and display the shooting interface.
[0416] The embodiment of the present application can be applied in the photographing scene of the terminal, wherein the photographing can include image shooting and video recording. Specifically, the user can open the application related to photographing (which can also be referred to as the camera application in the present application) on the terminal, and the terminal can display the shooting interface for photographing.
[0417] In this embodiment of the application, the shooting interface can be a preview interface in photo mode, a preview interface in video recording mode before starting video recording, or a preview interface in video recording mode after starting video recording.
[0418] Figure 5a Figure 6
[0419] For example, users can instruct the device to open the camera application by touching specific controls on the phone screen, pressing specific physical buttons or button combinations, inputting voice commands, or using air gestures. Upon receiving the user's instruction to open the camera, the device can launch the camera and display the shooting interface.
[0420] For example: Figure 6 As shown, users can open the camera app by clicking the "Camera" app icon 401 on the phone's home screen. The phone will then display the following: Figure 6 The shooting interface shown.
[0421] For example, even when the phone is locked, users can use a swipe-right gesture on the screen to instruct the phone to open the camera app, and the phone can display something like this: Figure 6 The shooting interface shown.
[0422] Alternatively, when the phone is locked, users can tap the "Camera" app shortcut icon on the lock screen to open the camera app. The phone can also display something like this. Next, the shooting magnification of the terminal is described in detail: The shooting interface shown.
[0423] For example, when other applications are running on the phone, users can also click the corresponding control to open the camera application to take pictures. For instance, when a user is using an instant messaging application (such as WeChat), they can also select the camera function control to instruct the phone to open the camera application to take pictures and record videos.
[0424] like Next, the different camera types are described respectively As shown, a camera's shooting interface typically includes a viewfinder 402, shooting controls, and other function controls (such as "large aperture," "portrait," "photo," and "video"). The viewfinder displays a preview of the video captured by the camera, allowing the user to determine when to instruct the terminal to perform a shooting operation based on this preview. This instruction could be, for example, by clicking the shooting control or pressing a volume button.
[0425] In an implementation, the user can trigger the camera to enter the image shooting mode by a click operation on the "shooting" control. When the camera is in the image shooting mode, the viewfinder of the shooting interface can be used to display a preview picture generated based on the video collected by the camera. The user can determine the timing of instructing the terminal to perform the shooting operation based on the preview picture in the viewfinder. The user can click the shooting button to shoot an image.
[0426] In an implementation, the user can trigger the camera to enter the video recording mode by a click operation on the "recording" control. When the camera is in the video recording mode, the viewfinder of the shooting interface can be used to display a preview picture generated based on the video collected by the camera. The user can determine the timing of instructing the terminal to start the video recording based on the preview picture in the viewfinder. The user can click the shooting button to start the video recording. At this time, the preview picture in the video recording process can be displayed in the viewfinder. Then, the user can click the shooting button to end the video recording.
[0427] The third image and the fourth image in the embodiments of the present application can be displayed in the preview picture in the image shooting mode, or in the preview picture before starting the video recording in the video recording mode, or in the preview picture after starting the video recording in the video recording mode.
[0428] In some embodiments, the shooting interface can further include a shooting magnification indication 403. The user can adjust the shooting magnification of the terminal based on adjusting the shooting magnification indication 403. The shooting magnification of the terminal can be described as the magnification of the preview picture in the viewfinder of the shooting interface.
[0429] It should be understood that the user can also adjust the magnification of the preview picture in the viewfinder of the shooting interface through other interactions (for example, based on the volume key of the terminal). The present application is not limited.
[0430] Figure 6
[0431] If multiple cameras are integrated on the terminal, each camera has its own inherent magnification. The inherent magnification of the camera can be understood as the zooming / zooming multiple of the focal length of the camera corresponding to the reference focal length. The reference focal length is usually the focal length of the main camera on the terminal.
[0432] In some embodiments, the user adjusts the magnification of the preview picture in the framing frame in the shooting interface, and then the terminal can perform digital zoom on the image captured by the camera, that is, the terminal ISP or other processor enlarges the area of each pixel of the image captured by the camera at the intrinsic magnification, and correspondingly reduces the framing range of the image, so that the processed image presents an image equivalent to the image captured by the camera at other shooting magnification. The shooting magnification of the terminal in this application can be understood as the shooting magnification of the image presented by the processed image.
[0433] The image captured by each camera can correspond to a shooting magnification range, I. The shooting interface includes the second control for indicating the opening or closing of the anti-shake processing and the first control for indicating the opening of the perspective distortion correction. Figure 7 :
[0434] In one implementation, the terminal can be integrated with one or more of a front short-focus (wide-angle) camera, a rear short-focus (wide-angle) camera, a rear medium-focus camera, and a rear long-focus camera.
[0435] Taking a terminal integrated with a short-focus (wide-angle) camera (including a front wide-angle camera, a rear wide-angle camera), a rear medium-focus camera, and a rear long-focus camera as an example. Among them, in the case where the relative position of the terminal and the photographed object does not change, the focal length of the short-focus (wide-angle) camera is the smallest, the field of view angle is the largest, and the size of the object in the captured image is the smallest. The focal length of the medium-focus camera is greater than that of the short-focus (wide-angle) camera, the field of view angle is smaller than that of the short-focus (wide-angle) camera, and the size of the object in the captured image is larger than that of the short-focus (wide-angle) camera. The focal length of the long-focus camera is the largest, the field of view angle is the smallest, and the size of the object in the captured image is the largest.
[0436] Among them, the field of view angle is used to indicate the maximum angle range that the camera can capture during the process of shooting an image. That is, if the object to be photographed is within this angle range, the object to be photographed will be captured by the mobile phone. If the object to be photographed is outside this angle range, the photographed object will not be captured by the mobile phone. Generally, the larger the field of view angle of the camera, the larger the shooting range using the camera. The smaller the field of view angle of the camera, the smaller the shooting range using the camera. It can be understood that "field of view angle" can be replaced by "field of view range", "field of view range", "field of view area", "imaging range" or "imaging field of view" and the like.
[0437] Generally, the user uses the mid-focus camera most frequently, and thus, the mid-focus camera is usually set as the main camera. The focal length of the main camera is set as a reference focal length, and the intrinsic magnification of the main camera is usually "1x". In some embodiments, digital zoom can be performed on the image captured by the main camera, that is, the area of each pixel of the "1x" image captured by the main camera is enlarged by the ISP or other processors in the mobile phone, and the field of view of the image is correspondingly reduced, so that the processed image presents an image equivalent to that captured by the main camera at other shooting magnifications (for example, "2x"). That is, the image captured by the main camera can correspond to a shooting magnification range, for example, "1x" to "5x". It should be understood that, in the case where the terminal also integrates a rear long-focus camera, the shooting magnification range corresponding to the image captured by the main camera can be 1-b, where the maximum magnification value b in the shooting magnification range can be the intrinsic magnification of the rear long-focus camera, and b can be in the range of 3-15, for example, but not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15.
[0438] It should be understood that, in the case where the terminal does not integrate a rear long-focus camera, the maximum magnification value b of the shooting magnification range corresponding to the image captured by the main camera can be the zoom maximum value of the terminal, and b can be in the range of 3-15, for example, but not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15.
[0439] It should be understood that, in the shooting magnification range (1-b) corresponding to the image captured by the main camera, the two end magnifications (1 and b) can be included, can not be included, or one of the two end magnifications can be included.
[0440] For example, the shooting magnification range corresponding to the image captured by the main camera can be the interval [1, b].
[0441] For example, the shooting magnification range corresponding to the image captured by the main camera can be the interval (1, b].
[0442] For example, the shooting magnification range corresponding to the image captured by the main camera can be the interval [1, b).
[0443] For example, the shooting magnification range corresponding to the image captured by the main camera can be the interval (1, b).
[0444] Similarly, the multiple of the focal length of the long-focus camera and the focal length of the main camera can be the shooting magnification of the long-focus camera, or referred to as the inherent magnification of the long-focus camera. For example, the focal length of the long-focus camera can be 5 times the focal length of the main camera, that is, the inherent magnification of the long-focus camera is "5x". Similarly, the image captured by the long-focus camera can also be digitally zoomed. That is, the image captured by the long-focus camera can correspond to a shooting magnification range, for example, "5x" to "50x". It should be understood that the shooting magnification range corresponding to the image captured by the long-focus camera can be b~c, and the maximum magnification value c of the shooting magnification range corresponding to the image captured by the long-focus camera can be in the range of 3~50.
[0445] Similarly, the multiple of the focal length of the short-focus (wide-angle) camera and the focal length of the main camera can be the shooting magnification of the short-focus (wide-angle) camera. For example, the focal length of the short-focus camera can be 0.5 times the focal length of the main camera, that is, the shooting magnification of the long-focus camera is "0.5x". Similarly, the image captured by the short-focus (wide-angle) camera can also be digitally zoomed. That is, the image captured by the long-focus camera can correspond to a shooting magnification range, for example, "0.5x" to "1x".
[0446] Wherein, the shooting magnification range corresponding to the image captured by the wide-angle camera can be a~1, the minimum value a can be the inherent magnification of the wide-angle camera, and the value range of a can be 0.5~0.9, for example, a can be but not limited to 0.5, 0.6, 0.7, 0.8 or 0.9.
[0447] It should be understood that the wide-angle camera can include a front wide-angle camera and a rear wide-angle camera. The shooting magnification range corresponding to the front wide-angle camera and the rear wide-angle camera can be the same or different, for example, the shooting magnification range corresponding to the front wide-angle camera can be a1~1, the minimum value a1 can be the inherent magnification of the front wide-angle camera, and the value range of a1 can be 0.5~0.9, for example, a1 can be but not limited to 0.5, 0.6, 0.7, 0.8 or 0.9, for example, the shooting magnification range corresponding to the rear wide-angle camera can be a2~1, the minimum value a2 can be the inherent magnification of the rear wide-angle camera, and the value range of a2 can be 0.5~0.9, for example, a2 can be but not limited to 0.5, 0.6, 0.7, 0.8 or 0.9.
[0448] It should be understood that the shooting magnification range (a~1) corresponding to the image captured by the wide-angle camera can include two end point magnifications (a and 1), can not include two end point magnifications (a and 1), or include one of the two end point magnifications.
[0449] For example, the shooting magnification range corresponding to the image captured by the main camera can be the interval [a, 1].
[0450] For example, the magnification range corresponding to the image captured by the main camera can be the interval (a, 1).
[0451] For example, the magnification range corresponding to the image captured by the main camera can be the interval [a, 1).
[0452] For example, the magnification range corresponding to the image captured by the main camera can be the interval (a, 1).
[0453] 502. A first enabling operation is detected when the user enables the perspective distortion correction function. The first enabling operation includes a first operation on a first control on the shooting interface of the terminal. The first control is used to indicate whether perspective distortion correction is enabled or disabled. The first operation is used to indicate whether perspective distortion correction is enabled.
[0454] 503. A second enabling operation is detected that the user has enabled the image stabilization function. The second enabling operation includes a second operation on a second control on the shooting interface of the terminal. The second control is used to indicate whether the image stabilization function is enabled or disabled. The second operation is used to indicate whether the image stabilization function is enabled.
[0455] First, describe how to adjust the shooting magnification when shooting with the terminal.
[0456] In some embodiments, users can manually adjust the shooting magnification when the terminal is shooting.
[0457] For example: Figure 7 As shown, users can adjust the shooting magnification of the terminal by operating the shooting magnification indicator 403 in the shooting interface. For example, when the current shooting magnification of the terminal is "1×", the user can change the shooting magnification of the terminal to "5×" by clicking the shooting magnification indicator 403 once or multiple times.
[0458] For example, users can reduce the shooting magnification of the terminal by pinching with two (or three) fingers in the shooting interface, or increase the shooting magnification of the terminal by sliding two (or three) fingers outward (in the opposite direction of pinching).
[0459] For example, users can also change the shooting magnification of the terminal by dragging the zoom ruler in the shooting interface.
[0460] For example, users can also change the camera magnification of the device by switching the currently used camera in the shooting interface or shooting settings interface. For instance, if the user selects to switch to the telephoto camera, the device will automatically increase the shooting magnification.
[0461] For example, the user can also change the shooting magnification of the mobile phone by selecting a long-focus shooting scene control or a long-distance shooting scene control in the shooting interface or the shooting setting interface.
[0462] In some other embodiments, the terminal can also automatically identify the specific scene of the image captured by the camera and automatically adjust the shooting magnification according to the identified specific scene. For example, if the terminal identifies that the image captured by the camera is a scene with a large field of view, such as the sea, mountains, or forests, the zoom magnification can be automatically reduced. For another example, if the terminal identifies that the image captured by the camera is a distant object, such as a bird flying in the distance or an athlete on a sports field, the zoom magnification can be automatically increased, which is not limited in the present application.
[0463] In one implementation, the shooting magnification of the terminal can be detected in real time, and the anti-shake processing function can be enabled when the shooting magnification of the terminal is greater than a second preset threshold, and a second control or prompt related to enabling the anti-shake processing function can be displayed on the shooting interface.
[0464] In one implementation, the shooting magnification of the terminal can be detected in real time, and the perspective distortion correction can be enabled when the shooting magnification of the terminal is less than a first preset threshold, and a first control or prompt related to the perspective distortion correction can be displayed on the shooting interface.
[0465] In one implementation, the anti-shake processing function and the perspective distortion correction function can be enabled by displaying a control for triggering the anti-shake processing and the perspective distortion correction on the camera interface, through interaction with the user, or automatically triggering the anti-shake processing function and the perspective distortion correction function in the above mode.
[0466] Or by displaying a control for triggering the anti-shake processing and the perspective distortion correction on the camera shooting parameter adjustment interface, enabling the anti-shake processing function and the perspective distortion correction function through interaction with the user, or automatically triggering part or all of the anti-shake processing function and the perspective distortion correction function in the above mode.
[0467] Next, the details are described as follows:
[0468] Figure 7 Figure 8
[0469] Taking the case that the camera is in the main camera shooting mode and the shooting interface of the camera is the preview interface in the shooting mode. Referring to Figure 9 At this time, the second control 405 (video stabilization) for indicating the anti-shake processing can be displayed on the shooting interface, and the first control 404 (perspective distortion correction) for indicating the perspective distortion correction can be displayed on the shooting interface. Figure 10 Figure 11 The display content of the second control DC (distortion correction) can prompt the user that the anti-shake processing is not enabled, and the display content of the first control can prompt the user that the perspective distortion correction is not enabled at this time. Referring to Figure 12 , if the user wants to enable the anti-shake processing function, the second control 405 can be clicked. In response to the user's click operation on the second control, the interface content shown in FIG. 4B can be displayed. Since the anti-shake processing function is enabled at this time, the display content of the second control 405 can prompt the user that the anti-shake processing is enabled at this time. If the user wants to enable the perspective distortion correction function, the first control 404 can be clicked. In response to the user's click operation on the second control, the perspective distortion correction function is enabled at this time, and the display content of the first control 404 can prompt the user that the perspective distortion correction is enabled at this time. Figure 13 If the user wants to disable the anti-shake processing function, the second control 405 can be clicked. In response to the user's click operation on the second control, the anti-shake processing function is disabled at this time, and the display content of the second control 405 can prompt the user that the anti-shake processing is disabled at this time. If the user wants to disable the perspective distortion correction function, the first control 404 can be clicked. The perspective distortion correction function is disabled at this time, and the display content of the first control 404 can prompt the user that the perspective distortion correction is disabled at this time.
[0470] In one implementation, when the camera is in the wide-angle shooting mode, the shooting interface of the camera can include a second control for indicating that the anti-shake processing is enabled and a first control for indicating that the perspective distortion correction is enabled.
[0471] Referring to
[0472] , the user can adjust the shooting magnification used by the mobile phone by operating the shooting magnification indication 403 in the shooting interface to adjust to the wide-angle shooting mode of the camera, and then the shooting interface shown in FIG. 4B can be displayed. If the user wants to enable the anti-shake processing function, the second control 405 can be clicked. In response to the user's click operation on the second control, the anti-shake processing function is enabled, and the display content of the second control 405 can prompt the user that the anti-shake processing is enabled at this time. If the user wants to enable the perspective distortion correction function, the first control 404 can be clicked. In response to the user's click operation on the second control, the perspective distortion correction function is enabled at this time, and the display content of the first control 404 can prompt the user that the perspective distortion correction is enabled at this time. II. The second control for indicating the opening of the anti-shake processing and the first control for triggering the opening of the perspective distortion correction are displayed in the shooting parameter adjustment interface of the camera. Figure 14
[0473] If the user wants to turn off the anti-shake processing function, the second control 405 can be clicked. In response to the user's click operation on the second control, the anti-shake processing function is turned off, and the display content of the second control 405 can prompt the user that the anti-shake processing is turned off at this time. If the user wants to turn off the perspective distortion correction function, the first control 404 can be clicked. At this time, the perspective distortion correction function is turned off, and the display content of the first control 404 can prompt the user that the perspective distortion correction is turned off at this time.
[0474] It should be understood that in the long-focus shooting mode, the perspective distortion is not obvious, and the perspective distortion correction can not be performed on the image. Therefore, in the shooting interface in the long-focus shooting mode, the second control for indicating that the anti-shake processing is turned on can be displayed, and the first control for indicating that the perspective distortion correction is turned on can not be displayed, or a prompt indicating that the user has turned off the perspective distortion correction is displayed to indicate that the perspective distortion correction is turned off by default. The prompt can not be interacted with the user to switch between turning on and turning off the perspective distortion correction. For example, refer to Figure 14 , the user can adjust the shooting magnification of the mobile phone by operating the shooting magnification indication 403 in the shooting interface to adjust to the long-focus shooting mode of the camera, and then the shooting interface shown in Figure 15 can be displayed, wherein Figure 15 the shooting interface does not include the first control for indicating that the perspective distortion correction is turned on.
[0475] Figure 15 Figure 16
[0476] In an implementation, in the wide-angle shooting mode, the super-wide-angle shooting mode, and the medium-focus shooting mode of the camera, the second control for indicating that the anti-shake processing is turned on and the first control for triggering the perspective distortion correction to be turned on can be displayed in the shooting parameter adjustment interface of the camera.
[0477] In the embodiments of the present application, refer to Figure 17 , Figure 17 the shooting interface of the camera shown in III. The second control for indicating the opening of the anti-shake processing is displayed, and the perspective distortion correction is opened by default. may include a first control for indicating that the shooting parameter adjustment interface is opened. The user can click the first control, and the terminal device can receive the third operation of the user on the first control. Then, in response to the third operation, the shooting parameter adjustment interface shown in Figure 18 may be opened. As shown in Figure 19 , the shooting parameter adjustment interface includes the second control for indicating that the anti-shake processing is turned on and the first control for indicating that the perspective distortion correction is turned on. Since the anti-shake processing and the perspective distortion correction are not turned on at this time, the display content of the second control can prompt the user that the anti-shake processing is not turned on at this time, and the display content of the first control can prompt the user that the anti-shake processing is not turned on at this time.If the user wants to enable the image stabilization function, they can click the second control. In response to the user's click on the second control, the following will be displayed: IV. The second control for indicating the opening of the anti-shake processing is displayed in the shooting parameter adjustment interface of the camera, and the perspective distortion correction is opened by default. The interface shown indicates that image stabilization is enabled, so the second control displays a message indicating that image stabilization is now active. If the user wants to enable perspective distortion correction, they can click the first control. In response to this click, perspective distortion correction is activated, and the first control displays a message indicating that it is now enabled.
[0478] If a user wants to disable image stabilization, they can click the second control. In response to this click, image stabilization will be disabled, and the second control will display a message indicating that image stabilization is now off. Similarly, if a user wants to disable perspective distortion correction, they can click the first control. Perspective distortion correction will then be disabled, and the first control will display a message indicating that perspective distortion correction is now off.
[0479] It should be understood that after the user returns to the shooting interface, the shooting interface can also display prompts indicating whether image stabilization and perspective distortion correction are currently enabled.
[0480] In one implementation, perspective distortion is not significant in telephoto shooting mode, so perspective distortion correction may not be applied to the image. Therefore, in telephoto shooting mode, the camera's shooting parameter adjustment interface can display a second control to indicate whether image stabilization is enabled, instead of the first control to indicate whether perspective distortion correction is enabled. Alternatively, a prompt indicating that perspective distortion correction is disabled by default can be displayed, and this prompt cannot be interactively switched between enabling and disabling perspective distortion correction. For example, refer to... V. The first control for indicating the opening of the perspective distortion correction is displayed, and the anti-shake processing is opened by default. , Figure 20 The shooting parameter adjustment interface shown does not include a first control for indicating whether perspective distortion correction is enabled, but includes a prompt indicating that perspective distortion correction is disabled by default.
[0481] Figure 21
[0482] In one implementation, when the camera is in wide-angle shooting mode, ultra-wide-angle shooting mode, or mid-telephoto shooting mode, the shooting interface can display a second control to indicate that image stabilization is enabled, instead of displaying a first control to indicate that perspective distortion correction is enabled, or display a prompt indicating that perspective distortion correction is enabled by default. This prompt cannot be interactively switched between enabling and disabling perspective distortion correction.
[0483] In an implementation, when the camera is in the wide-angle shooting mode, the shooting interface of the camera can include a second control for indicating that the anti-shake processing is turned on, and the perspective distortion correction is turned on by default.
[0484] In an implementation, when the camera is in the ultra-wide-angle shooting mode, the shooting interface of the camera can include a second control for indicating that the anti-shake processing is turned on, and the perspective distortion correction is turned on by default.
[0485] In an implementation, when the camera is in the mid-focus shooting mode, the shooting interface of the camera can include a second control for indicating that the anti-shake processing is turned on, and the perspective distortion correction is turned on by default.
[0486] Referring to VI. The first control for triggering the opening of the perspective distortion correction is displayed in the shooting parameter adjustment interface of the camera, and the anti-shake processing is opened by default. , a second control 405 for indicating that the anti-shake processing is turned on can be displayed in the shooting interface of the camera, and the first control for indicating that the perspective distortion correction is turned on is not displayed, or referring to VII. The perspective distortion correction and the anti-shake processing are opened by default. , a prompt 404 indicating that the user has turned on the perspective distortion correction is displayed to indicate that the perspective distortion correction is turned on by default, and the prompt 404 cannot be interacted with the user to switch between turning on and turning off the perspective distortion correction. Since the anti-shake processing is not turned on at this time, the display content of the second control can prompt the user that the anti-shake processing is not turned on at this time. If the user wants to turn on the anti-shake processing function, the second control can be clicked. In response to the clicking operation of the user on the second control, the anti-shake processing function is turned on, and the display content of the second control can prompt the user that the anti-shake processing is turned on at this time.
[0487] In another embodiment, the mobile phone can also hide some controls on the shooting interface to avoid the image being blocked by the controls as much as possible and improve the user visual experience. Figure 24a
[0488] In an implementation, in the wide-angle shooting mode, the ultra-wide-angle shooting mode, and the mid-focus shooting mode of the camera, a second control for indicating that the anti-shake processing is turned on can be displayed in the shooting parameter adjustment interface of the camera, and the first control for indicating that the perspective distortion correction is turned on is not displayed, or a prompt indicating that the user has turned on the perspective distortion correction is displayed to indicate that the perspective distortion correction is turned on by default, and the prompt cannot be interacted with the user to switch between turning on and turning off the perspective distortion correction.
[0489] In an implementation, a third image collected by the camera can be acquired, the third image is an image frame collected before the first image is collected, the perspective distortion correction is performed on the third image to obtain a fourth image, and the fourth image is displayed in the shooting interface of the camera. In this embodiment, if the camera does not turn on the anti-shake processing and turns on the perspective distortion correction by default, the image displayed in the shooting interface at this time is the image (the fourth image) that has been subjected to the perspective distortion correction but not subjected to the anti-shake processing.
[0490] Figure 24a
[0491] In an implementation, when the camera is in the wide-angle shooting mode, the ultra-wide-angle shooting mode, and the mid-focus shooting mode, the shooting interface can display the first control for indicating to turn on the perspective distortion correction, without displaying the second control for indicating to turn on the anti-shake processing, or display a prompt indicating that the user has turned on the anti-shake processing, to indicate that the anti-shake processing is turned on by default, which cannot be interacted with the user to switch between turning on and turning off the anti-shake processing.
[0492] In an implementation, when the camera is in the wide-angle shooting mode, the shooting interface of the camera can include the first control for indicating to turn on the perspective distortion correction, and the anti-shake processing is turned on by default.
[0493] In an implementation, when the camera is in the ultra-wide-angle shooting mode, the shooting interface of the camera can include the first control for indicating to turn on the perspective distortion correction, and the anti-shake processing is turned on by default.
[0494] In an implementation, when the camera is in the mid-focus shooting mode, the shooting interface of the camera can include the first control for indicating to turn on the perspective distortion correction, and the anti-shake processing is turned on by default.
[0495] Referring to Figure 24a , the first control 404 for indicating to turn on the perspective distortion correction can be displayed in the shooting interface of the camera, without displaying the second control for indicating to turn on the anti-shake processing, or referring to Figure 24b , a prompt 405 indicating that the user has turned on the anti-shake processing is displayed, to indicate that the anti-shake processing is turned on by default, which cannot be interacted with the user to switch between turning on and turning off the anti-shake processing. Since the perspective distortion correction is not turned on at this time, the display content of the first control can prompt the user that the perspective distortion correction is not turned on at this time, and if the user wants to turn on the perspective distortion correction function, the first control can be clicked. In response to the click operation of the user on the first control, the perspective distortion correction function is turned on, and the display content of the first control can prompt the user that the perspective distortion correction is turned on at this time.
[0496] Figure 25 Figure 26
[0497] In an implementation, when the camera is in the wide-angle shooting mode, the ultra-wide-angle shooting mode, and the mid-focus shooting mode, the first control for indicating to turn on the perspective distortion correction can be displayed in the shooting parameter adjustment interface of the camera, without displaying the second control for indicating to turn on the anti-shake processing, or displaying a prompt indicating that the user has turned on the anti-shake processing, to indicate that the anti-shake processing is turned on by default, which cannot be interacted with the user to switch between turning on and turning off the anti-shake processing.
[0498] In an implementation, the fourth image captured by the camera can be acquired, the fourth image being an image frame captured before the first image is captured, and the fourth image is subjected to the anti-shake processing to obtain a second anti-shake output, and the second anti-shake output indicating image is displayed in the shooting interface of the camera. In this embodiment, if the camera does not open the perspective distortion correction and the anti-shake processing is opened by default, the image displayed in the shooting interface is the image subjected to the anti-shake processing and not subjected to the perspective distortion correction (the second anti-shake output indicating image).
[0499] Figure 27
[0500] In an implementation, when the camera is in the wide-angle shooting mode, the ultra-wide-angle shooting mode, and the mid-focus shooting mode, the shooting interface can not display the first control for indicating that the perspective distortion correction is opened, and can not display the second control for indicating that the anti-shake processing is opened, but the perspective distortion correction and the anti-shake processing are opened by default; or a prompt indicating that the user has opened the anti-shake processing is displayed to represent that the anti-shake processing is opened by default, the prompt can not be interacted with the user to switch between the opening and the closing of the anti-shake processing, and a prompt indicating that the user has opened the perspective distortion correction is displayed to represent that the perspective distortion correction is opened by default, the prompt can not be interacted with the user to switch between the opening and the closing of the perspective distortion correction.
[0501] The above description is based on the shooting interface before the camera starts shooting. In a possible implementation, the shooting interface can also be the shooting interface before the camera starts recording, or the shooting interface after the camera starts recording.
[0502] Figure 28 Figure 3
[0503] For example, when the user does not operate the shooting interface for a long time, the second control and the first control in the above embodiments can be stopped from being displayed; after detecting that the user clicks the touch screen, the hidden second control and the first control can be restored to be displayed.
[0504] For how to perform the anti-shake processing on the captured first image to obtain the first cropping region, and perform the perspective distortion correction processing on the first cropping region to obtain the third cropping region, refer to the description of steps 301 and 302 in the above embodiments, which will not be repeated here.
[0505] 504. The first image is subjected to anti-shake processing to obtain a first cropped region; the first cropped region is subjected to perspective distortion correction processing to obtain a third cropped region; a third image is obtained based on the first image and the third cropped region; the second image is subjected to anti-shake processing to obtain a second cropped region; and the second cropped region is related to the first cropped region and the shaking information; the shaking information is used to represent the shaking that occurs during the process of the terminal acquiring the first image and the second image; the second cropped region is subjected to perspective distortion correction processing to obtain a fourth cropped region; a fourth image is obtained based on the second image and the fourth cropped region.
[0506] The description of step 504 can be found in the descriptions of steps 301 to 305, and will not be repeated here.
[0507] 505. Generate a target video based on the third image and the fourth image, and display the target video.
[0508] In one implementation, a target video can be generated and displayed based on the third and fourth images. Specifically, for any two frames in a video stream captured in real time by a terminal device, an output image after image stabilization and perspective distortion correction can be obtained based on steps 301 to 306 above, and then a video after image stabilization and perspective distortion correction (i.e., the target video) can be obtained. For example, the first image can be the nth frame in the video stream, and the second image can be the (n+1)th frame. Then, the (n-2)th and (n-1)th frames can also be processed as the first and second images respectively in steps 301 to 306 above, and the (n-1)th and nth frames can also be processed as the first and second images respectively in steps 301 to 306 above, and the (n+1)th and (n+2)th frames can also be processed as the first and second images respectively in steps 301 to 306 above.
[0509] In one implementation, after obtaining the target video, the target video can be displayed in the viewfinder of the camera's shooting interface.
[0510] It should be understood that the above shooting interface can be a preview interface in photo mode. The target video is the preview screen in photo mode. Users can click the shooting control on the shooting interface or trigger the camera to take a picture through other triggering methods, and then obtain the captured image.
[0511] It should be understood that the above shooting interface can be a preview interface in video recording mode before starting video recording. Thus, the target video is a preview screen in video recording mode before starting video recording. Users can click the shooting control on the shooting interface or trigger the camera to start recording through other triggering methods.
[0512] It should be understood that the above shooting interface can be a preview interface after starting video recording in a video recording mode, and the target video is a preview picture after starting video recording in the video recording mode. The user can click a shooting control of the shooting interface or trigger the camera to end recording by other triggering manners, and save the video shot during recording.
[0513] Example 2: Video post-processing
[0514] The image processing method in the embodiments of the present application is described above taking the instant shooting scene as an example. In another implementation, the above anti-shake processing and perspective distortion correction processing can be performed on a video in a gallery in a terminal device or a cloud server.
[0515] Reference Figure 29 , Figure 30 The flow of an image processing method provided by the embodiments of the present application is shown in FIG. 2, which includes the following steps. Figure 30
[0516] 2401, acquire a video saved in the album, the video including a first image and a second image.
[0517] The user can click the "gallery" application icon on the mobile phone desktop (for example, as shown in FIG. 3) to instruct the mobile phone to start the gallery application. The mobile phone can display an application interface as shown in FIG. 4. The user can long-press a selected video. The terminal device can acquire the video saved in the album and selected by the user in response to the long-press operation of the user, and display an interface as shown in FIG. 5. The interface can include selection controls (including a "selection" control, a "deletion" control, a "share to" control, and a "more" control). Figure 30 Figure 31 Figure 32
[0518] 2402, enable anti-shake function and perspective distortion correction function.
[0519] The user can click the "more" control. In response to the click operation of the user, an interface as shown in FIG. 6 is displayed. The interface can include selection controls (including an "anti-shake processing" control and a "video distortion correction" control). The user can click the "anti-shake processing" control and the "video distortion correction" control, and click a "complete" control as shown in FIG. 7. Figure 33 Figure 33
[0520] 2403, perform anti-shake processing on the first image to obtain a first cropping region;
[0521] 2404, perform perspective distortion correction processing on the first cropping region to obtain a third cropping region;
[0522] 2405、obtaining a third image according to the first image and the third cropped region;
[0523] 2406, performing anti-shake processing on the second image to obtain a second cropped region; the second cropped region is related to the first cropped region and shake information; the shake information is used to indicate shake occurring in a process in which the terminal collects the first image and the second image;
[0524] It should be understood that the shake information herein can be obtained and saved in real time by the terminal when collecting the video.
[0525] 2407, performing perspective distortion correction processing on the second cropped region to obtain a fourth cropped region;
[0526] 2408, obtaining a fourth image according to the second image and the fourth cropped region;
[0527] 2409, generating a target video according to the third image and the fourth image.
[0528] The terminal device can perform anti-shake processing and perspective distortion correction processing on the selected video, and the specific description of steps 2403 to 2408 can refer to the description in the corresponding embodiments of the present application, which will not be repeated here. After the terminal device completes the anti-shake processing and the perspective distortion correction processing, an interface as shown in FIG. 13B can be displayed to prompt the user that the anti-shake processing and the perspective distortion correction processing are completed. Figure 33 Figure 34
[0529] Example 3, target tracking
[0530] Referring to Figure 34 , Figure 34 The flow of an image processing method provided by an embodiment of the present application is shown in FIG. 13A, and the image processing method provided by an embodiment of the present application includes the following steps. Figure 33
[0531] 3001, determining a target object in a collected first image to obtain a first cropped region including the target object;
[0532] In some scenarios, the camera can start an object tracking function, which can ensure that the terminal device outputs images including a target object that is moving while the posture of the terminal device remains unchanged. Taking a person object as an example, specifically, when the camera is in a wide-angle shooting mode or an ultra-wide-angle shooting mode, the coverage range of the obtained original image is large, and the original image includes a target person object. The area where the person object is located can be identified, and the original image can be cropped and enlarged to obtain a cropped area including the person object. The cropped area is part of the original image. When the person object moves and does not exceed the shooting range of the camera, the person object can be tracked, and the area where the person object is located can be obtained in real time. The original image can be cropped and enlarged to obtain a cropped area including the person object. In this way, the terminal device can also output images including a moving person object while the posture of the terminal device remains unchanged.
[0533] Specifically, refer to Figure 34 The shooting interface of the camera can include a "more" control. The user can click the "more" control. In response to the click operation of the user, an application interface as shown in Figure 35 The interface can include a control indicating that the user starts target tracking. The user can click the control to start the target tracking function of the camera.
[0534] Similar to the above anti-shake processing, in existing implementations, in order to perform perspective distortion correction on images while performing anti-shake processing, the cropped area including the person object needs to be confirmed on the original image, and perspective distortion correction needs to be performed on the entire area of the original image. The cropped area including the person object is equivalent to performing perspective distortion correction, and the cropped area that has undergone perspective distortion correction is output. However, the above method has the following problems: Due to the movement of the target object, the positions of the cropped area including the person object in the front and back frames of images are different, which further causes the displacement of the perspective distortion correction of the position points or pixel points in the two frames of output images to be different. Since adjacent two frames of anti-shake output images include the same or substantially the same content, the deformation degree of the same object between images is different, that is, the inter-frame consistency is lost, and a Jello phenomenon occurs, which reduces the quality of the video output.
[0535] 3002. Perform perspective distortion correction processing on the first cropped area to obtain a third cropped area.
[0536] In the embodiment, different from the perspective distortion correction on the whole region of the original image in the prior implementation, the perspective distortion correction is performed on the cropped region. Due to the movement of the target object, the positions of the sub-image including the target object in the front and back frame images are different, and further the displacement of the perspective distortion correction of the position points or pixel points in the output two frame images is different. However, the sizes of the cropped regions determined after the target recognition between the two frames are consistent. Therefore, if the perspective distortion correction is performed on the cropped region, the degree of the perspective distortion correction of the adjacent frames is consistent (because the distances of the position points or pixel points in the cropped regions of the two frames from the center point of the sub-image are consistent). And because the adjacent two frame cropped regions are images including the same or basically the same content, and the deformation degrees of the same object between the cropped regions are the same, that is, the inter-frame consistency is maintained, and further the obvious Jello phenomenon is avoided, and the quality of the video output is improved.
[0537] In the embodiment, the cropped regions of the front and back frames can include the same target object (for example, a person object). Because the object of the perspective distortion correction is the cropped region, the shape difference between the target objects deformed in the perspective distortion correction outputs of the front and back frames is small with respect to the target object. Specifically, the shape difference can be within a preset range. The preset range can be understood as being difficult to be distinguished by the naked eye of a person, or the Jello phenomenon between the front and back frames is difficult to be distinguished by the naked eye of a person.
[0538] In a possible implementation, the shooting interface of the camera includes a second control, the second control is used to indicate to start the perspective distortion correction; a second operation of the user on the second control is received; and the perspective distortion correction is performed on the sub-image in response to the second operation.
[0539] More specific description about step 3002 can be referred to the description of step 302 in the above embodiment, and the similar parts will not be described herein.
[0540] 3003, obtaining a third image according to the first image and the third cropped region.
[0541] Specific description about step 3003 can be referred to the description of step 303 in the above embodiment, and the similar parts will not be described herein.
[0542] 3004, determining the target object in the second image collected to obtain a second cropped region including the target object; wherein the second cropped region is related to target movement information, and the target movement information is used to indicate the movement of the target object in the process of collecting the first image and the second image by the terminal;
[0543] The specific description of step 3004 can refer to the description of step 301 in the above embodiment, and similar parts will not be repeated here.
[0544] 3005, performing perspective distortion correction processing on the second cropping region to obtain a third cropping region;
[0545] The specific description of step 3005 can refer to the description of step 305 in the above embodiment, and similar parts will not be repeated here.
[0546] 3006, obtaining a fourth image according to the second image and the third cropping region;
[0547] 3007, generating a target video according to the third image and the fourth image.
[0548] The specific description of step 3006 can refer to the description of step 306 in the above embodiment, and similar parts will not be repeated here.
[0549] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.
[0550] In a possible implementation, before the perspective distortion correction processing on the first cropping region, the method further includes:
[0551] Detecting that the terminal meets a distortion correction condition to enable a distortion correction function.
[0552] In a possible implementation, detecting that the terminal meets a distortion correction condition includes, but is not limited to, one or more of the following cases: case 1: detecting that the shooting magnification of the terminal is less than a first preset threshold.
[0553] In an implementation, whether to enable perspective distortion correction can be determined based on the shooting magnification of the terminal. In the case that the shooting magnification of the terminal is large (for example, in the full or partial magnification range of the telephoto shooting mode, or in the full or partial magnification range of the medium focal length shooting mode), the degree of perspective distortion of the image captured by the terminal is low. The so-called low degree of perspective distortion can be understood as the case that the human eye almost cannot distinguish that there is perspective distortion in the image captured by the terminal. Therefore, in the case that the shooting magnification of the terminal is large, perspective distortion correction does not need to be enabled.
[0554] In a possible implementation, when the shooting magnification is a~1, the terminal uses a front wide-angle camera or a rear wide-angle camera to capture a video stream in real time, and the first preset threshold is greater than 1, where a is the intrinsic magnification of the front wide-angle camera or the rear wide-angle camera, and the value range of a is 0.5~0.9; or,
[0555] When the photographing magnification is 1-b, the terminal adopts the rear main camera to collect a video stream in real time, b is the inherent magnification of the rear long-focus camera included in the terminal or is the maximum zoom value of the terminal, the first preset threshold is 1-15, and the value range of b is 3-15.
[0556] Case 2: A first enabling operation of starting the perspective distortion correction function is detected.
[0557] In a possible implementation, the first enabling operation includes a first operation on a first control on a photographing interface of the terminal, and the first control is used to indicate starting or stopping the perspective distortion correction, and the first operation is used to indicate starting the perspective distortion correction.
[0558] In an implementation, the photographing interface can include a control (referred to as a first control in the embodiments of the present application) used to indicate starting or stopping the perspective distortion correction. The user can trigger starting the perspective distortion correction by a first operation on the first control, that is, enable the perspective distortion correction by the first operation on the first control. Then, the terminal can detect the first enabling operation of starting the perspective distortion correction by the user, and the first enabling operation includes the first operation on the first control on the photographing interface of the terminal.
[0559] In an implementation, the control (referred to as a first control in the embodiments of the present application) used to indicate starting or stopping the perspective distortion correction is displayed on the photographing interface only when it is detected that the photographing magnification of the terminal is less than a first preset threshold.
[0560] Case 3: It is identified that a face exists in a photographing scene; or,
[0561] It is identified that a face exists in a photographing scene, and a distance between the face and the terminal is less than a preset value; or,
[0562] It is identified that a face exists in a photographing scene, and a pixel proportion of the face in an image corresponding to the photographing scene is greater than a preset proportion.
[0563] It is understood that when a human face exists in a shooting scene, the deformation degree of the human face caused by perspective distortion correction will be more visually obvious, and the smaller the distance between the human face and the terminal or the larger the pixel proportion of the human face in the image corresponding to the shooting scene, the larger the area of the human face in the image, and the more visually obvious the deformation degree of the human face caused by perspective distortion correction. Therefore, in the above scenario, perspective distortion correction needs to be enabled. In this embodiment, whether to enable perspective distortion correction is determined through the determination of the above conditions related to the human face in the shooting scene, so that the shooting scene in which perspective distortion will occur can be accurately judged, and perspective distortion correction processing is performed on the shooting scene in which perspective distortion will occur, while no perspective distortion correction processing is performed on the shooting scene in which perspective distortion will not occur, thereby achieving accurate processing of the image signal and saving power consumption.
[0564] The embodiment of the present application provides an image processing method, comprising: determining a target object in a first image collected to obtain a first clipping region comprising the target object; performing perspective distortion correction processing on the first clipping region to obtain a third clipping region; obtaining a third image according to the first image and the third clipping region; determining the target object in a second image collected to obtain a second clipping region comprising the target object; wherein the second clipping region is related to target movement information, and the target movement information is used to represent movement of the target object in the process of collecting the first image and the second image by the terminal; performing perspective distortion correction processing on the second clipping region to obtain a third clipping region; obtaining a fourth image according to the second image and the third clipping region, and the third image and the fourth image are used to generate a video; wherein the first image and the second image are a pair of original images collected by the terminal in time domain and adjacent to each other. Different from the perspective distortion correction on the whole region of the original image in the prior art, in the embodiment of the present application, the perspective distortion correction is performed on the clipping region. Due to the movement of the target object, the positions of the clipping region comprising the human object in the front and back frames of images are different, and then the displacement of the perspective distortion correction of the position points or pixel points in the output two frames of images is different. However, the size of the clipping region determined after the target recognition between the two frames is consistent. Therefore, if the perspective distortion correction is performed on the clipping region, the degree of perspective distortion correction of adjacent frames is consistent (because the distances of each position point or pixel point in the clipping region of the two frames relative to the center point of the clipping region are consistent). And because the adjacent two frames of clipping regions are images comprising the same or basically the same content, the deformation degree of the same object between images is the same, that is, the inter-frame consistency is maintained, and then the obvious Jello phenomenon does not occur, thereby improving the quality of the video output.
[0565] The application further provides an image display device, which can be a terminal device, with reference to Figure 35 , Figure 35 The application provides an image processing device, as shown in Figure 5a The image processing device 3300 comprises:
[0566] A shake processing module 3301 is configured to perform shake processing on the collected first image to obtain a first clipping region, and perform shake processing on the collected second image to obtain a second clipping region, wherein the second clipping region is related to the first clipping region and shake information, and the shake information is used to indicate the shake occurring in the process of collecting the first image and the second image by the terminal, and the first image and the second image are a pair of original images collected by the terminal in time sequence.
[0567] The shake processing module 3301 is specifically described in steps 301 and 304, which will not be repeated here.
[0568] A perspective distortion correction module 3302 is configured to perform perspective distortion correction processing on the first clipping region to obtain a third clipping region, and perform perspective distortion correction processing on the second clipping region to obtain a fourth clipping region.
[0569] The perspective distortion correction module 3302 is specifically described in steps 302 and 305, which will not be repeated here.
[0570] An image generation module 3303 is configured to obtain a third image according to the first image and the third clipping region, obtain a fourth image according to the second image and the fourth clipping region, and generate a target video according to the third image and the fourth image.
[0571] The image generation module 3303 is specifically described in steps 303, 306 and 307, which will not be repeated here.
[0572] In a possible implementation, the offset direction of the position of the second clipping region on the second image relative to the position of the first clipping region on the first image is opposite to the shake direction of the shake occurring in the process of collecting the first image and the second image by the terminal.
[0573] In a possible implementation, the first clipping region is used to indicate a first sub-region in the first image, and the second clipping region is used to indicate a second sub-region in the second image.
[0574] The first sub-region corresponds to first image content in the first image, and the second sub-region corresponds to second image content in the second image, and a similarity between the first image content and the second image content is greater than a similarity between the first image and the second image.
[0575] In a possible implementation, the apparatus further includes:
[0576] The first detection module 3304 is configured to, before the anti-shake processing module performs anti-shake processing on the collected first image, detect that the terminal satisfies a distortion correction condition, and enable a distortion correction function.
[0577] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0578] It is detected that a shooting magnification of the terminal is less than a first preset threshold.
[0579] In a possible implementation, when the shooting magnification is a~1, the terminal collects a video stream in real time by using a front wide-angle camera or a rear wide-angle camera, the first preset threshold is greater than 1, and the a is an inherent magnification of the front wide-angle camera or the rear wide-angle camera, and the a is in a range of 0.5 to 0.9; or,
[0580] When the shooting magnification is 1~b, the terminal collects a video stream in real time by using a rear main camera, the b is an inherent magnification of a rear telephoto camera included in the terminal or is a zoom maximum value of the terminal, the first preset threshold is 1 to 15, and the b is in a range of 3 to 15.
[0581] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0582] It is detected that a first enabling operation of a perspective distortion correction function is performed by a user, the first enabling operation includes a first operation on a first control on a shooting interface of the terminal, the first control is used to indicate that the perspective distortion correction is turned on or turned off, and the first operation is used to indicate that the perspective distortion correction is turned on.
[0583] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0584] It is identified that a face exists in a shooting scene; or,
[0585] It is identified that a face exists in a shooting scene, and a distance between the face and the terminal is less than a preset value; or,
[0586] It is identified that a face exists in a shooting scene, and a pixel proportion of the face in an image corresponding to the shooting scene is greater than a preset proportion.
[0587] In a possible implementation, the apparatus further includes:
[0588] The second detection module 3305 is configured to, before the anti-shake processing module performs anti-shake processing on the collected first image, detect that the terminal meets an anti-shake condition, and enable an anti-shake processing function.
[0589] In a possible implementation, the detection that the terminal meets the anti-shake condition includes:
[0590] The detection that a shooting magnification of the terminal is greater than an intrinsic magnification of a camera with the smallest magnification in the terminal.
[0591] In a possible implementation, the detection that the terminal meets the anti-shake condition includes:
[0592] The detection of a second enabling operation of a user to turn on the anti-shake processing function, the second enabling operation including a second operation on a second control on a shooting interface of the terminal, the second control being used to indicate turning on or turning off the anti-shake processing function, and the second operation being used to indicate turning on the anti-shake processing function.
[0593] In a possible implementation, the image generation module is specifically configured to obtain a first mapping relationship and a second mapping relationship, the first mapping relationship being used to represent a mapping relationship between each position point of the second cropped region and a corresponding position point in the second image, and the second mapping relationship being used to represent a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second cropped region.
[0594] The fourth image is obtained according to the second image, the first mapping relationship, and the second mapping relationship.
[0595] In a possible implementation, the image generation module is specifically configured to couple the first mapping relationship and the second mapping relationship to determine a target mapping relationship, the target mapping relationship including a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second image.
[0596] The fourth image is determined according to the second image and the target mapping relationship.
[0597] In a possible implementation, the first mapping relationship is different from the second mapping relationship.
[0598] In a possible implementation, the perspective distortion correction module is specifically configured to perform optical distortion correction processing on the first cropped region to obtain a corrected first cropped region, perform perspective distortion correction processing on the corrected first cropped region, or perform optical distortion correction processing on the first cropped region and perform perspective distortion correction processing on the corrected first cropped region.
[0599] The image generation module is specifically configured to perform optical distortion correction processing on the third cropped region to obtain a corrected third cropped region; obtain a third image according to the first image and the corrected third cropped region; or,
[0600] The perspective distortion correction module is specifically configured to perform optical distortion correction processing on the second cropped region to obtain a corrected second cropped region; perform perspective distortion correction processing on the corrected second cropped region; or,
[0601] The image generation module is specifically configured to perform optical distortion correction processing on the fourth cropped region to obtain a corrected fourth cropped region; and obtain a fourth image according to the second image and the corrected fourth cropped region.
[0602] An image processing device provided by an embodiment of the present application is applied to terminal real-time collection of a video stream; the device comprises: a shake reduction processing module, configured to perform shake reduction processing on a collected first image to obtain a first cropped region, and further configured to perform shake reduction processing on a collected second image to obtain a second cropped region; the second cropped region is related to the first cropped region and shake information; the shake information is used to indicate shake occurring in a process in which the terminal collects the first image and the second image; wherein the first image and the second image are a pair of original images adjacent in time domain and collected by the terminal; a perspective distortion correction module, configured to perform perspective distortion correction processing on the first cropped region to obtain a third cropped region, and further configured to perform perspective distortion correction processing on the second cropped region to obtain a fourth cropped region; an image generation module, configured to obtain a third image according to the first image and the third cropped region, and further configured to obtain a fourth image according to the second image and the fourth cropped region; and a target video is generated according to the third image and the fourth image. Since adjacent two frames of shake reduction output images including the same or basically the same content, the deformation degree of the same object between the images is the same, that is, the interframe consistency is maintained, and thus the Jello phenomenon does not occur, and the quality of video display is improved.
[0603] The present application further provides an image processing device, and the image display device can be a terminal device, which is described with reference to Figure 24a , Figure 30 The structure of an image processing device provided by an embodiment of the present application is shown in The image processing device 3400 comprises:
[0604] The object determining module 3401 is configured to determine a target object in a first image collected by the terminal to obtain a first clipping region including the target object, and determine the target object in a second image collected by the terminal to obtain a second clipping region including the target object. The second clipping region is related to target movement information, and the target movement information indicates movement of the target object during collection of the first image and the second image by the terminal. The first image and the second image are a pair of original images collected by the terminal in time sequence.
[0605] The specific description of the object determining module 3401 can refer to the description of steps 3001 and 3004, which will not be repeated here.
[0606] The perspective distortion correction module 3402 is configured to perform perspective distortion correction on the first clipping region to obtain a third clipping region, and perform perspective distortion correction on the second clipping region to obtain a third clipping region.
[0607] The specific description of the perspective distortion correction module 3402 can refer to the description of steps 3002 and 3005, which will not be repeated here.
[0608] The image generating module 3403 is configured to generate a third image according to the first image and the third clipping region, generate a fourth image according to the second image and the third clipping region, and generate a target video according to the third image and the fourth image.
[0609] The specific description of the image generating module 3403 can refer to the description of steps 3003, 3006 and 3007, which will not be repeated here.
[0610] In an optional implementation, the position of the target object in the second image is different from the position of the target object in the first image.
[0611] In an optional implementation, the apparatus further includes:
[0612] The first detection module 3404 is configured to detect that the terminal satisfies a distortion correction condition before the perspective distortion correction module performs perspective distortion correction on the first clipping region, and enable the distortion correction function.
[0613] In a possible implementation, the detection that the terminal satisfies the distortion correction condition includes:
[0614] It is detected that the shooting magnification of the terminal is less than a first preset threshold.
[0615] In a possible implementation, when the photographing magnification is a~1, the terminal adopts a front wide-angle camera or a rear wide-angle camera to collect a video stream in real time, the first preset threshold is greater than 1, where the a is an inherent magnification of the front wide-angle camera or the rear wide-angle camera, and the a is in a range of 0.5 to 0.9; or,
[0616] When the photographing magnification is 1~b, the terminal adopts a rear main camera to collect a video stream in real time, the b is an inherent magnification of a rear telephoto camera included in the terminal or is a zoom maximum value of the terminal, the first preset threshold is 1~15, and the b is in a range of 3 to 15.
[0617] In a possible implementation, the detecting that the terminal meets the distortion correction condition includes:
[0618] detecting a first enabling operation of a user to start a perspective distortion correction function, the first enabling operation including a first operation on a first control on a photographing interface of the terminal, the first control being used to indicate starting or stopping perspective distortion correction, and the first operation being used to indicate starting perspective distortion correction.
[0619] In a possible implementation, the detecting that the terminal meets the distortion correction condition includes:
[0620] identifying that a face exists in a photographing scene; or,
[0621] identifying that a face exists in a photographing scene and a distance between the face and the terminal is less than a preset value; or,
[0622] identifying that a face exists in a photographing scene and a pixel proportion of the face in an image corresponding to the photographing scene is greater than a preset proportion.
[0623] The embodiment of the present application provides an image processing device, which is applied to a terminal for collecting a video stream in real time; the device comprises: an object determining module, which is used for determining a target object in a first image collected by the terminal to obtain a first clipping area comprising the target object, and is also used for determining the target object in a second image collected by the terminal to obtain a second clipping area comprising the target object; wherein the second clipping area is related to target movement information, and the target movement information is used for representing movement of the target object in the process of collecting the first image and the second image by the terminal; wherein the first image and the second image are a pair of original images collected by the terminal in time domain and adjacent in time; a perspective distortion correction module, which is used for performing perspective distortion correction processing on the first clipping area to obtain a third clipping area, and is also used for performing perspective distortion correction processing on the second clipping area to obtain a third clipping area; an image generating module, which is used for obtaining a third image according to the first image and the third clipping area, and is also used for obtaining a fourth image according to the second image and the third clipping area; and a target video is generated according to the third image and the fourth image.
[0624] Different from the perspective distortion correction on the whole area of the original image in the prior art, the perspective distortion correction is performed on the clipping area in the embodiment of the present application. Due to the movement of the target object, the positions of the clipping area comprising the person object in the front and back frames of images are different, and then the displacement of the perspective distortion correction of the position points or pixel points in the output two frames of images is different. However, the size of the clipping area determined after the target recognition between the two frames is consistent. Therefore, if the perspective distortion correction is performed on the clipping area, the degree of the perspective distortion correction of the adjacent frames is consistent (because the distances of the position points or pixel points in the two frames of the clipping area to the center point of the clipping area are consistent). And because the adjacent two frames of the clipping area are images comprising the same or basically same content, the deformation degree of the same object between the images is the same, that is, the interframe consistency is maintained, and then the obvious Jello phenomenon does not appear, and the quality of the video output is improved.
[0625] Next, a terminal device provided by the embodiment of the present application is introduced, and the terminal device can be the image processing device in and , please refer to , A structural schematic diagram of a terminal device provided in an embodiment of the present application is shown in FIG. 3. The terminal device 3500 can be a virtual reality (VR) device, a mobile phone, a tablet, a notebook computer, a smart wearable device, etc., and is not limited herein. Specifically, the terminal device 3500 includes a receiver 3501, a transmitter 3502, a processor 3503, and a memory 3504 (wherein the number of processors 3503 in the terminal device 3500 can be one or more, for example, one processor is taken as an example in the present application), wherein the processor 3503 can include an application processor 35031 and a communication processor 35032. In some embodiments of the present application, the receiver 3501, the transmitter 3502, the processor 3503, and the memory 3504 can be connected through a bus or other means.
[0626] The memory 3504 can include a read-only memory and a random access memory, and provide the processor 3503 with instructions and data. A part of the memory 3504 can also include a non-volatile random access memory (NVRAM). The memory 3504 stores processor and operation instructions, executable modules or data structures, or a subset thereof, or an expanded set thereof, wherein the operation instructions can include various operation instructions for implementing various operations.
[0627] The processor 3503 controls the operation of the terminal device. In a specific application, various components of the terminal device are coupled together through a bus system, wherein the bus system can include a data bus, a power bus, a control bus, and a status signal bus, etc. in addition to the data bus. However, for the sake of clarity, all kinds of buses are referred to as a bus system in the figure.
[0628] The method disclosed in the embodiments of the present application can be applied to the processor 3503 or implemented by the processor 3503. The processor 3503 can be an integrated circuit chip having a signal processing capability. In implementation, the steps of the method disclosed above can be completed by an integrated logic circuit or an instruction in a form of software in the processor 3503. The processor 3503 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller. The processor 3503 can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components. The processor 3503 can implement or execute the methods, steps and logical block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code executed by the processor or a combination of hardware and software modules in the processor. The software module can reside in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or the like storage medium in the art. The storage medium is located in the storage 3504, and the processor 3503 reads information in the storage 3504 and combines the hardware to complete the steps of the methods disclosed above. Specifically, the processor 3503 can read information in the storage 3504 and combine the hardware to complete the steps related to data processing in steps 301 to 304 in the embodiments disclosed above and the steps related to data processing in steps 3001 to 3004 in the embodiments disclosed above.
[0629] The receiver 3501 can be configured to receive input digital or character information and generate signal input related to the relevant settings and function control of the terminal device. The transmitter 3502 can be configured to output digital or character information through the first interface. The transmitter 3502 can also be configured to send instructions to the disk group through the first interface to modify data in the disk group. The transmitter 3502 can further include a display device such as a display screen.
[0630] The embodiments of the present application also provide a computer program product including a computer program that, when running on a computer, causes the computer to perform the steps of the image processing method described in the embodiments of the present application. 、 and the corresponding embodiments described in the embodiments of the present application.
[0631] The embodiment of the present application further provides a computer readable storage medium, which stores a program for signal processing, and when the program is run on a computer, the computer is enabled to perform the steps of the image processing method in the method as described in the foregoing embodiment.
[0632] The image display device provided by the embodiment of the present application can be a chip, which comprises a processing unit, for example, a processor, and a communication unit, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute computer execution instructions stored in a storage unit, so as to enable the chip in the execution device to execute the data processing method described in the foregoing embodiment, or so as to enable the chip in the training device to execute the data processing method described in the foregoing embodiment. Alternatively, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0633] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiment. In addition, in the apparatus embodiment provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0634] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CPU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, the software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in a readable storage medium, such as floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, server or network device, etc.) execute the method described in various embodiments of the application.
[0635] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be in the form of computer program product entirely or partially.
[0636] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the flow or function described in the embodiments of the application is entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as server, data center, etc. integrated with one or more available media. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.
Claims
1. An image processing method, characterized by, The method is applied to a terminal for collecting a video stream in real time; the method comprises: performing anti-shake processing on a collected first image to obtain a first cropped region; performing perspective distortion correction processing on the first cropped region to obtain a third cropped region; obtaining a third image according to the first image and the third cropped region; performing anti-shake processing on a collected second image to obtain a second cropped region; the second cropped region is related to the first cropped region and shake information; the shake information is used to indicate shake occurring in a process in which the terminal collects the first image and the second image; performing perspective distortion correction processing on the second cropped region to obtain a fourth cropped region; obtaining a fourth image according to the second image and the fourth cropped region; generating a target video according to the third image and the fourth image; wherein the first image and the second image are a pair of original images collected by the terminal in time domain and adjacent in time; obtaining a fourth image according to the second image and the fourth cropped region comprises: obtaining a first mapping relationship and a second mapping relationship; the first mapping relationship is used to indicate a mapping relationship between each position point of the second cropped region and a corresponding position point in the second image; the second mapping relationship is used to indicate a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second cropped region; obtaining the fourth image according to the second image, the first mapping relationship and the second mapping relationship.
2. The method of claim 1, wherein, A shift direction of a position of the second cropped region on the second image relative to a position of the first cropped region on the first image is opposite to a shake direction of shake occurring in the process in which the terminal collects the first image and the second image.
3. The method according to claim 1 or 2, characterized in that, The first cropped region is used to indicate a first sub-region in the first image; the second cropped region is used to indicate a second sub-region in the second image; A similarity of first image content corresponding to the first sub-region in the first image and second image content corresponding to the second sub-region in the second image is greater than a similarity of the first image and the second image.
4. The method according to any one of claims 1 to 2, characterized in that, Before performing perspective distortion correction processing on the first cropped region, the method further comprises: detecting that the terminal satisfies a distortion correction condition to enable a distortion correction function.
5. The method of claim 4, wherein, The detection that the terminal satisfies the distortion correction condition comprises: detecting that a shooting magnification of the terminal is less than a first preset threshold.
6. The method of claim 5, wherein: when the shooting magnification is a~1, the terminal collects the video stream in real time by using a front wide-angle camera or a rear wide-angle camera, and the first preset threshold is greater than 1, wherein a is an inherent magnification of the front wide-angle camera or the rear wide-angle camera, and a value range of a is 0.5~0.9; or When the photographing magnification is 1~b, the terminal adopts a rear main camera to collect a video stream in real time, the b is an inherent magnification of a rear long-focus camera included in the terminal or is a zoom maximum value of the terminal, the first preset threshold is 1~15, and the b is in a range of 3~15.
7. The method of claim 4, wherein, The detection that the terminal meets the distortion correction condition includes: detecting a first enabling operation in which a user opens a perspective distortion correction function, the first enabling operation including a first operation on a first control on a photographing interface of the terminal, the first control being used to indicate that the perspective distortion correction is opened or closed, and the first operation being used to indicate that the perspective distortion correction is opened.
8. The method of claim 4, wherein, The detection that the terminal meets the distortion correction condition includes: identifying that a face exists in a photographing scene; or, identifying that a face exists in a photographing scene and a distance between the face and the terminal is less than a preset value; or, identifying that a face exists in a photographing scene and a pixel proportion of the face in an image corresponding to the photographing scene is greater than a preset proportion.
9. The method of any one of claims 1 to 2, wherein, Before the first image collected is subjected to the anti-shake processing, the method further includes: detecting that the terminal meets an anti-shake condition to enable an anti-shake processing function.
10. The method of claim 9, wherein, The detection that the terminal meets the anti-shake condition includes: detecting that a photographing magnification of the terminal is greater than an inherent magnification of a camera with a smallest magnification in the terminal.
11. The method of claim 9, wherein, The detection that the terminal meets the anti-shake condition includes: detecting a second enabling operation in which a user opens an anti-shake processing function, the second enabling operation including a second operation on a second control on a photographing interface of the terminal, the second control being used to indicate that the anti-shake processing function is opened or closed, and the second operation being used to indicate that the anti-shake processing function is opened.
12. The method of claim 1, wherein, The obtaining of the fourth image according to the second image, the first mapping relationship, and the second mapping relationship includes: coupling the first mapping relationship and the second mapping relationship to determine a target mapping relationship, the target mapping relationship including a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second image; determining the fourth image according to the second image and the target mapping relationship.
13. The method of claim 12, wherein, The first mapping relationship is different from the second mapping relationship.
14. The method of any one of claims 1 to 2, wherein, The perspective distortion correction processing of the first cropped region includes: optical distortion correction processing of the first cropped region to obtain a corrected first cropped region; and perspective distortion correction processing of the corrected first cropped region; or The obtaining of the third image according to the first image and the third cropped region includes: optical distortion correction processing of the third cropped region to obtain a corrected third cropped region; and the obtaining of the third image according to the first image and the corrected third cropped region; or The perspective distortion correction processing of the second cropped region includes: optical distortion correction processing of the second cropped region to obtain a corrected second cropped region; and perspective distortion correction processing of the corrected second cropped region; or The fourth image is obtained according to the second image and the fourth cropped region, and the obtaining comprises: performing optical distortion correction processing on the fourth cropped region to obtain a corrected fourth cropped region; and obtaining the fourth image according to the second image and the corrected fourth cropped region.
15. An image processing apparatus characterized by comprising: The device is applied to a terminal for collecting a video stream in real time, and the device comprises: a shake reduction processing module, configured to perform shake reduction processing on a collected first image to obtain a first cropped region, and perform shake reduction processing on a collected second image to obtain a second cropped region, wherein the second cropped region is related to the first cropped region and shake information, and the shake information is used to indicate shake occurring in a process in which the terminal collects the first image and the second image, wherein the first image and the second image are a pair of original images collected by the terminal and adjacent in time domain; a perspective distortion correction module, configured to perform perspective distortion correction processing on the first cropped region to obtain a third cropped region, and perform perspective distortion correction processing on the second cropped region to obtain a fourth cropped region; an image generation module, configured to obtain a third image according to the first image and the third cropped region, obtain a fourth image according to the second image and the fourth cropped region, and generate a target video according to the third image and the fourth image. The image generation module is specifically configured to obtain a first mapping relationship and a second mapping relationship, the first mapping relationship is used to indicate a mapping relationship between each position point of the second cropped region and a corresponding position point in the second image, and the second mapping relationship is used to indicate a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second cropped region. The fourth image is obtained according to the second image, the first mapping relationship and the second mapping relationship.
16. The apparatus of claim 15, wherein, An offset direction of a position of the second cropped region on the second image relative to a position of the first cropped region on the first image is opposite to a shake direction of shake occurring in the process in which the terminal collects the first image and the second image.
17. The apparatus of claim 15 or 16, wherein, The first cropped region is used to indicate a first sub-region in the first image, and the second cropped region is used to indicate a second sub-region in the second image. A similarity of first image content corresponding to the first sub-region in the first image and second image content corresponding to the second sub-region in the second image is greater than a similarity of the first image and the second image.
18. The apparatus of any one of claims 15 to 16, wherein, The device further comprises: a first detection module, configured to, before the shake reduction processing module performs shake reduction processing on a collected first image, detect that the terminal satisfies a distortion correction condition, and enable a distortion correction function.
19. The apparatus of claim 18, wherein, The detection that the terminal satisfies the distortion correction condition comprises: detecting that a shooting magnification of the terminal is less than a first preset threshold.
20. The device of claim 19, wherein When the shooting magnification is a~1, the terminal adopts a front wide-angle camera or a rear wide-angle camera to collect a video stream in real time, the first preset threshold is greater than 1, wherein the a is the inherent magnification of the front wide-angle camera or the rear wide-angle camera, and the a is in a range of 0.5~0.9; or, When the shooting magnification is 1~b, the terminal adopts a rear main camera to collect a video stream in real time, the b is the inherent magnification of a rear telephoto camera included in the terminal or is the maximum zoom value of the terminal, the first preset threshold is 1~15, and the b is in a range of 3~15.
21. The apparatus of claim 18, wherein, The detection that the terminal meets the distortion correction condition comprises: detecting a first enabling operation in which a user opens a perspective distortion correction function, the first enabling operation comprising a first operation on a first control on a shooting interface of the terminal, the first control being used to indicate that the perspective distortion correction is opened or closed, and the first operation being used to indicate that the perspective distortion correction is opened.
22. The apparatus of claim 18, wherein, The detection that the terminal meets the distortion correction condition comprises: identifying that a face exists in a shooting scene; or, identifying that a face exists in a shooting scene and a distance between the face and the terminal is less than a preset value; or, identifying that a face exists in a shooting scene and a pixel proportion of the face in an image corresponding to the shooting scene is greater than a preset proportion.
23. The apparatus of any one of claims 15 to 16, wherein, The apparatus further comprises: a second detection module configured to, before the anti-shake processing module performs anti-shake processing on the collected first image, detect that the terminal meets an anti-shake condition and enable an anti-shake processing function.
24. The apparatus of claim 23, wherein, The detection that the terminal meets the anti-shake condition comprises: detecting that a shooting magnification of the terminal is greater than an inherent magnification of a camera with the smallest magnification in the terminal.
25. The apparatus of claim 23, wherein, The detection that the terminal meets the anti-shake condition comprises: detecting a second enabling operation in which a user opens an anti-shake processing function, the second enabling operation comprising a second operation on a second control on a shooting interface of the terminal, the second control being used to indicate that the anti-shake processing function is opened or closed, and the second operation being used to indicate that the anti-shake processing function is opened.
26. The apparatus of claim 15, wherein, The image generation module is specifically configured to couple the first mapping relationship and the second mapping relationship to determine a target mapping relationship, the target mapping relationship comprising a mapping relationship between each position point of the fourth cropped region and a corresponding position point in the second image; determine the fourth image according to the second image and the target mapping relationship.
27. The apparatus of claim 26, wherein, The first mapping relationship and the second mapping relationship are different.
28. The apparatus of any one of claims 15 to 16, wherein, The perspective distortion correction module is specifically configured to perform optical distortion correction processing on the first cropped region to obtain a corrected first cropped region, and perform perspective distortion correction processing on the corrected first cropped region; or The image generation module is specifically configured to perform optical distortion correction processing on the third cropped region to obtain a corrected third cropped region, and obtain a third image according to the first image and the corrected third cropped region; or The perspective distortion correction module is specifically configured to perform optical distortion correction processing on the second cropped region to obtain a corrected second cropped region, and perform perspective distortion correction processing on the corrected second cropped region. The image generation module is specifically configured to perform optical distortion correction processing on the fourth cropped region to obtain a corrected fourth cropped region, and generate a fourth image according to the second image and the corrected fourth cropped region.
29. An image processing apparatus characterized by comprising: The device comprises a processor, a memory, a camera and a bus, wherein: The processor, the memory and the camera are connected through the bus; The camera is configured to collect a video in real time; The memory is configured to store computer programs or instructions; The processor is configured to call or execute the programs or instructions stored on the memory, and is further configured to call the camera to implement the method steps in any one of claims 1-14.
30. A computer readable storage medium comprising a program which, when executed on a computer, causes the computer to carry out the method of any one of claims 1 to 14.
31. A computer program product comprising instructions, wherein: When the computer program product is executed on the terminal, the terminal is caused to perform the method of any one of claims 1-14. When the computer program product is executed on the terminal, the terminal is caused to perform the method of any one of claims 1-14.
Citation Information
Patent Citations
Image processing device, image processing method, image processing program, and recording medium
CN104995908A
Video anti-shake method based on image content understanding
CN110602393A
Video stabilization
CN111133747A
Imaging apparatus, its control method, program and storage medium
JP2005195656A
User interface for capturing and managing visual medium
JP2021040300A