Image transformation method and apparatus
By obtaining the distance between the face and the camera during a selfie and performing 3D model distortion correction, the problem of facial perspective distortion in selfies was solved, improving the realism and aesthetics of selfie images.
Patent Information
- Application Number
- CN202011141042.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-28
- Filing Date
- 2020-10-22
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2040-10-22
AI Technical Summary
When taking selfies, the close proximity of the camera to the face causes perspective distortion, affecting the selfie quality. Current technology struggles to effectively eliminate this distortion while maintaining realism.
By obtaining the distance between the target person's face and the camera, and when the distance is less than a preset threshold, distortion correction is performed using a 3D model, including the establishment of a 3D model, perspective projection, and pixel displacement vector field calculation, to achieve a 3D transformation effect of the image.
It significantly improves the imaging effect in selfie scenarios, making the relative proportions and positions of facial features closer to the real appearance, and eliminating perspective distortion.
Smart Images

Figure CN113850726B_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, and more particularly to an image transformation method and apparatus. Background Technology
[0002] Photography has become an important medium for recording life, and in recent years, taking selfies with the front-facing camera of a mobile phone has become increasingly popular. However, due to the close distance between the camera and the face during selfies, the perspective distortion problem of "near objects appearing larger than distant objects" has become increasingly significant. For example, when shooting portraits at close range, the varying distances of different parts of the face from the camera often result in the nose appearing larger and the face being elongated, affecting the subjective effect of the portrait. Therefore, image transformation processing is needed, such as altering the distance, pose, and position of objects in the image, to eliminate the perspective distortion effect of "near objects appearing larger than distant objects" without affecting the realism of the portrait, thereby improving its aesthetic appeal.
[0003] The imaging process of a camera obtains a two-dimensional image of a three-dimensional object. Commonly used image processing algorithms are designed for two-dimensional images, but this makes it difficult to achieve the effect of realistic transformation of three-dimensional objects in the image obtained after transforming the two-dimensional image. Summary of the Invention
[0004] This application provides an image transformation method and apparatus to improve the imaging effect in selfie scenarios.
[0005] In a first aspect, this application provides an image transformation method, comprising: acquiring a first image of a target scene using a front-facing camera, the target scene including the face of a target person; acquiring a target distance between the face of the target person and the front-facing camera; when the target distance is less than a preset threshold, performing a first processing on the first image to obtain a second image; the first processing including distortion correction of the first image based on the target distance; wherein the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0006] The first image is captured by the device's front-facing camera during a selfie. This first image includes two scenarios: first, when the user hasn't triggered the shutter, the first image hasn't yet been captured on the image sensor and is merely a raw image obtained by the camera; second, when the user has triggered the shutter, the first image is a raw image already captured on the image sensor. Therefore, the second image also includes two scenarios. In the first scenario, the second image is a distorted version of the raw image captured by the camera, which is also not captured on the image sensor and serves only as a preview for the user. In the second scenario, the second image is a distorted version of the raw image already captured on the image sensor, which the device can save in its image library.
[0007] This application achieves a three-dimensional transformation effect of the image by using a three-dimensional model, and corrects the perspective distortion of the face of the target person in the close-up image. This makes the relative proportions and relative positions of the facial features of the target person after correction closer to the relative proportions and relative positions of the facial features of the target person, which can significantly improve the shooting imaging effect in selfie scenarios.
[0008] In one possible implementation, the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a designated part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera.
[0009] The target distance between the target person's face and the camera can be the distance between the foremost part of the target person's face (e.g., the nose) and the camera; or, the target distance between the target person's face and the camera can be the distance between a specific part of the target person's face (e.g., the eyes, mouth, or nose) and the camera; or, the target distance between the target person's face and the camera can be the distance between the center of the target person's face (e.g., the nose in a frontal view, or the cheekbone in a side view) and the camera. It should be noted that the definition of the target distance can vary depending on the specific circumstances of the first image, and this application does not impose any specific limitations on it.
[0010] In one possible implementation, obtaining the target distance between the target person's face and the front-facing camera includes: obtaining the screen ratio of the target person's face in the first image; and obtaining the target distance based on the screen ratio and the field of view (FOV) of the front-facing camera.
[0011] In one possible implementation, obtaining the target distance between the face of the target person and the front-facing camera includes: obtaining the target distance through a distance sensor, wherein the distance sensor includes a time-of-flight (TOF) sensor, a structured light sensor, or a binocular sensor.
[0012] This application can obtain the target distance by calculating the screen-to-body ratio of the face, that is, firstly, obtain the screen-to-body ratio of the target person's face in the first image (the ratio of the pixel area of the face to the pixel area of the first image), and then obtain the target distance based on the screen-to-body ratio and the FOV of the front-facing camera. Alternatively, the target distance can be obtained by measuring with a distance sensor. Other methods can also be used to obtain the target distance, and this application does not specifically limit these methods.
[0013] In one possible implementation, the preset threshold is less than 80 centimeters.
[0014] In one possible implementation, the preset threshold is 50 centimeters.
[0015] This application sets a threshold value. If the target distance between the subject's face and the front-facing camera is less than this threshold, the acquired first image containing the subject's face is considered distorted and requires distortion correction. The preset threshold value is within 80 centimeters; optionally, it can be set to 50 centimeters. It should be noted that the specific value of the preset threshold can vary depending on the performance of the front-facing camera, lighting conditions, etc., and this application does not impose specific limitations on it.
[0016] In one possible implementation, the second image includes a preview image or an image obtained after the shutter is triggered.
[0017] The second image can be a preview image captured by the front-facing camera. That is, before the shutter is triggered, the front-facing camera captures a first image facing the target scene area. The terminal then uses the method described above to correct perspective distortion in this first image to obtain the second image, which is then displayed on the terminal's screen. In this case, the second image seen by the user is a preview image that has already undergone perspective distortion correction. Alternatively, the second image can be an image obtained after the shutter is triggered. That is, after the shutter is triggered, the first image is the image captured on the terminal's image sensor. The terminal uses the method described above to correct perspective distortion in this first image to obtain the second image, which is then saved and displayed on the terminal's screen. In this case, the second image seen by the user is an image from the saved image library that has undergone perspective distortion correction.
[0018] In one possible implementation, the face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
[0019] When the first image is acquired, if the target distance between the subject's face and the front-facing camera is less than a preset threshold, the facial features in the first image may appear disproportionately large or small due to perspective distortion caused by the front-facing camera. This would result in changes in the size and stretching of the facial features, causing their relative proportions and positions to deviate from the true proportions and positions of the subject's face. The second image, obtained by correcting the perspective distortion of the first image, eliminates these disproportionate changes and stretching. Therefore, the relative proportions and positions of the facial features in the second image approach or even restore the true proportions and positions of the subject's face.
[0020] In one possible implementation, the distortion correction of the first image based on the target distance includes: fitting the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; and performing perspective distortion correction on the first image based on the depth information to obtain the second image.
[0021] In one possible implementation, the step of correcting perspective distortion of the first image based on the depth information to obtain the second image includes: establishing a first three-dimensional model of the target person's face; transforming the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the target person's face; obtaining a pixel displacement vector field of the target person's face based on the depth information, the first three-dimensional model, and the second three-dimensional model; and obtaining the second image based on the pixel displacement vector field of the target person's face.
[0022] A 3D model is built based on the target person's face, and then the 3D model is used to achieve a 3D transformation effect on the image. The pixel displacement vector field of the target person's face is obtained based on the correlation between the sampling points of the 3D model before and after the transformation, and then the transformed 2D image is obtained. This can realize perspective distortion correction of the target person's face in close-up images, so that the relative proportions and relative positions of the facial features of the corrected target person are closer to the relative proportions and relative positions of the facial features of the target person, which can significantly improve the shooting imaging effect in selfie scenarios.
[0023] In one possible implementation, a first coordinate set is obtained by perspective projection of the first 3D model based on the depth information, the first coordinate set including coordinate values corresponding to multiple pixels in the first 3D model; a second coordinate set is obtained by perspective projection of the second 3D model based on the depth information, the second coordinate set including coordinate values corresponding to multiple pixels in the second 3D model; the coordinate difference between the first coordinate value and the second coordinate value is calculated to obtain the pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to the first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the multiple identical pixels contained in the first 3D model and the second 3D model.
[0024] Secondly, this application provides an image transformation method, comprising: acquiring a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; displaying a distortion correction function menu; acquiring transformation parameters input by a user on the distortion correction function menu, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the face of the target person and the camera when the shooting terminal is shooting the face of the target person; performing a first processing on the first image to obtain a second image; the first processing including distortion correction on the first image according to the transformation parameters; wherein the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0025] This application corrects perspective distortion of the first image based on transformation parameters to obtain a second image. The face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image. That is, the relative proportions and relative positions of the facial features of the target person in the second image are closer to the relative proportions and relative positions of the facial features of the target person in the first image than the relative proportions and relative positions of the facial features of the target person.
[0026] In one possible implementation, the distortion of the target person's face in the first image is caused by the second terminal capturing the first image at a target distance less than a first preset threshold between the target person's face and the second terminal; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera.
[0027] In one possible implementation, the target distance is obtained by the screen-to-body ratio of the target person's face in the first image and the field of view (FOV) of the second terminal's camera; or, the target distance is obtained by the equivalent focal length in the EXIF information of the first image's exchangeable image file format.
[0028] In one possible implementation, the distortion correction function menu includes an option to adjust the equivalent simulated shooting distance; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the equivalent simulated shooting distance based on an instruction triggered by the user's operation of the control or slider in the option to adjust the equivalent simulated shooting distance.
[0029] In one possible implementation, when the distortion correction function menu is initially displayed, the value of the equivalent analog shooting distance in the option to adjust the equivalent analog shooting distance includes a default value or a pre-calculated value.
[0030] In one possible implementation, before displaying the distortion correction function menu, the method further includes: displaying a pop-up window when the target person's face is distorted, the pop-up window providing a selection control for whether to perform distortion correction; and responding to the user's operation when the user clicks the distortion correction control on the pop-up window.
[0031] In one possible implementation, before displaying the distortion correction function menu, the method further includes: when the target person's face is distorted, displaying a distortion correction control, the distortion correction control being used to open the distortion correction function menu; and responding to the user's operation when the user clicks the distortion correction control.
[0032] In one possible implementation, the distortion of the target person's face in the first image is caused by the second terminal capturing the first image when the camera's field of view (FOV) is greater than a second preset threshold, and the pixel distance between the target person's face and the edge of the FOV is less than a third preset threshold; wherein, the pixel distance includes the number of pixels between the foremost part of the target person's face and the edge of the FOV; or, the number of pixels between a specified part of the target person's face and the edge of the FOV; or, the number of pixels between the center position of the target person's face and the edge of the FOV.
[0033] In one possible implementation, the FOV is obtained from the EXIF information of the first image.
[0034] In one possible implementation, the second preset threshold is 90°, and the third preset threshold is one-quarter of the length or width of the first image.
[0035] In one possible implementation, the distortion correction function menu includes an option to adjust the displacement distance; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction and displacement distance based on the instructions triggered by the user's operation of the control or slider in the option to adjust the displacement distance.
[0036] In one possible implementation, the distortion correction function menu includes options for adjusting the relative position and / or relative proportion of facial features; the step of obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction, displacement distance, and / or facial feature size based on the instructions triggered by the user's operation of the controls or sliders in the options for adjusting the relative position and / or relative proportion of facial features.
[0037] In one possible implementation, the distortion correction function menu includes an option to adjust the angle; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction and adjustment angle according to the instruction triggered by the user's operation of the control or slider in the adjustment angle option; or, the distortion correction function menu includes an option to adjust facial expression; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining a new facial expression template according to the instruction triggered by the user's operation of the control or slider in the adjustment facial expression option; or, the distortion correction function menu includes an option to adjust motion; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining a new motion template according to the instruction triggered by the user's operation of the control or slider in the adjustment motion option.
[0038] In one possible implementation, the face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
[0039] In one possible implementation, the distortion correction of the first image based on the transformation parameters includes: fitting the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; and performing perspective distortion correction on the first image based on the depth information and the transformation parameters to obtain the second image.
[0040] In one possible implementation, the step of correcting perspective distortion of the first image based on the depth information and the transformation parameters to obtain the second image includes: establishing a first three-dimensional model of the target person's face; transforming the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the target person's face; obtaining a pixel displacement vector field of the target person's face based on the depth information, the first three-dimensional model, and the second three-dimensional model; and obtaining the second image based on the pixel displacement vector field of the target person's face.
[0041] In one possible implementation, obtaining the pixel displacement vector field of the target person's face based on the depth information, the first 3D model, and the second 3D model includes: performing perspective projection on the first 3D model based on the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to multiple pixels in the first 3D model; performing perspective projection on the second 3D model based on the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to multiple pixels in the second 3D model; calculating the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to the first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the multiple identical pixels contained in the first 3D model and the second 3D model.
[0042] Thirdly, this application provides an image transformation method applied to a first terminal. The method includes: acquiring a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; displaying a distortion correction function menu on a screen, the distortion correction function menu including one or more sliders and / or one or more controls; receiving a distortion correction instruction, the distortion correction instruction including transformation parameters generated when a user performs a touch operation on the one or more sliders and / or the one or more controls, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the face of the target person and the camera when the shooting terminal is shooting the face of the target person; performing a first processing on the first image according to the transformation parameters to obtain a second image; the first processing including distortion correction on the first image; wherein the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0043] Fourthly, this application provides an image transformation device, comprising: an acquisition module, configured to acquire a first image of a target scene using a front-facing camera, the target scene including the face of a target person; and acquire a target distance between the face of the target person and the front-facing camera; and a processing module, configured to perform a first processing on the first image to obtain a second image when the target distance is less than a preset threshold; the first processing includes distortion correction of the first image based on the target distance; wherein the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0044] In one possible implementation, the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a designated part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera.
[0045] In one possible implementation, the acquisition module is specifically used to acquire the screen ratio of the target person's face in the first image; and to obtain the target distance based on the screen ratio and the field of view (FOV) of the front-facing camera.
[0046] In one possible implementation, the acquisition module is specifically used to acquire the target distance through a distance sensor, which includes a time-of-flight (TOF) sensor, a structured light sensor, or a binocular sensor.
[0047] In one possible implementation, the preset threshold is less than 80 centimeters.
[0048] In one possible implementation, the second image includes a preview image or an image obtained after the shutter is triggered.
[0049] In one possible implementation, the face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
[0050] In one possible implementation, the processing module is specifically used to fit the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; and to perform perspective distortion correction on the first image based on the depth information to obtain the second image.
[0051] In one possible implementation, the processing module is specifically used to: establish a first three-dimensional model of the target person's face; transform the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the target person's face; obtain the pixel displacement vector field of the target person's face based on the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image based on the pixel displacement vector field of the target person's face.
[0052] In one possible implementation, the processing module is specifically configured to perform perspective projection on the first 3D model based on the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to multiple pixels in the first 3D model; perform perspective projection on the second 3D model based on the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to multiple pixels in the second 3D model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain a pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to the first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the multiple identical pixels contained in the first 3D model and the second 3D model.
[0053] Fifthly, this application provides an image transformation device, comprising: an acquisition module for acquiring a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; a display module for displaying a distortion correction function menu; the acquisition module is further configured to acquire transformation parameters input by a user on the distortion correction function menu, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the face of the target person and the camera when the shooting terminal is shooting the face of the target person; and a processing module for performing a first processing on the first image to obtain a second image; the first processing including distortion correction on the first image according to the transformation parameters; wherein the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0054] In one possible implementation, the distortion of the target person's face in the first image is caused by the second terminal capturing the first image at a target distance less than a first preset threshold between the target person's face and the second terminal; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera.
[0055] In one possible implementation, the target distance is obtained by the screen-to-body ratio of the target person's face in the first image and the field of view (FOV) of the second terminal's camera; or, the target distance is obtained by the equivalent focal length in the EXIF information of the first image's exchangeable image file format.
[0056] In one possible implementation, the distortion correction function menu includes an option to adjust the equivalent simulated shooting distance; the acquisition module is specifically used to acquire the equivalent simulated shooting distance based on the instruction triggered by the user's operation of the control or slider in the option to adjust the equivalent simulated shooting distance.
[0057] In one possible implementation, when the distortion correction function menu is initially displayed, the value of the equivalent analog shooting distance in the option to adjust the equivalent analog shooting distance includes a default value or a pre-calculated value.
[0058] In one possible implementation, the display module is further configured to display a pop-up window when the target person's face is distorted, the pop-up window providing a selection control for whether to perform distortion correction; and responding to the user's operation when the user clicks the distortion correction control on the pop-up window.
[0059] In one possible implementation, the display module is further configured to display a distortion correction control when the target person's face is distorted, the distortion correction control being used to open the distortion correction function menu; and to respond to the user's operation when the user clicks the distortion correction control.
[0060] In one possible implementation, the distortion of the target person's face in the first image is caused by the second terminal capturing the first image when the camera's field of view (FOV) is greater than a second preset threshold, and the pixel distance between the target person's face and the edge of the FOV is less than a third preset threshold; wherein, the pixel distance includes the number of pixels between the foremost part of the target person's face and the edge of the FOV; or, the number of pixels between a specified part of the target person's face and the edge of the FOV; or, the number of pixels between the center position of the target person's face and the edge of the FOV.
[0061] In one possible implementation, the FOV is obtained from the EXIF information of the first image.
[0062] In one possible implementation, the second preset threshold is 90°, and the third preset threshold is one-quarter of the length or width of the first image.
[0063] In one possible implementation, the distortion correction function menu includes an option to adjust the displacement distance; the acquisition module is further configured to acquire the adjustment direction and displacement distance based on the instructions triggered by the user's operation of the control or slider in the option to adjust the displacement distance.
[0064] In one possible implementation, the distortion correction function menu includes options for adjusting the relative position and / or relative proportion of facial features; the acquisition module is further configured to acquire the adjustment direction, displacement distance, and / or facial feature size based on instructions triggered by the user's operation of controls or sliders in the options for adjusting the relative position and / or relative proportion of facial features.
[0065] In one possible implementation, the distortion correction function menu includes an option to adjust the angle; the acquisition module is further configured to acquire the adjustment direction and adjustment angle based on instructions triggered by the user's operation of controls or sliders in the angle adjustment option; or, the distortion correction function menu includes an option to adjust facial expressions; the acquisition module is further configured to acquire new facial expression templates based on instructions triggered by the user's operation of controls or sliders in the facial expression adjustment option; or, the distortion correction function menu includes an option to adjust actions; the acquisition module is further configured to acquire new action templates based on instructions triggered by the user's operation of controls or sliders in the action adjustment option.
[0066] In one possible implementation, the face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
[0067] In one possible implementation, the processing module is specifically used to fit the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; and to perform perspective distortion correction on the first image based on the depth information and the transformation parameters to obtain the second image.
[0068] In one possible implementation, the processing module is specifically used to: establish a first three-dimensional model of the target person's face; transform the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the target person's face; obtain the pixel displacement vector field of the target person's face according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the target person's face.
[0069] In one possible implementation, the processing module is specifically configured to perform perspective projection on the first 3D model based on the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to multiple pixels in the first 3D model; perform perspective projection on the second 3D model based on the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to multiple pixels in the second 3D model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain a pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to the first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the multiple identical pixels contained in the first 3D model and the second 3D model.
[0070] In one possible implementation, the system further includes: a recording module; the acquisition module is further configured to acquire a recording instruction based on a user's trigger operation on the recording control; the recording module is configured to start recording the acquisition process of the second image based on the recording instruction until a stop recording instruction is received from a user's trigger operation on the stop recording control.
[0071] Sixthly, this application provides an image transformation device, comprising: an acquisition module for acquiring a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; a display module for displaying a distortion correction function menu on a screen, the distortion correction function menu including one or more sliders and / or one or more controls; receiving a distortion correction instruction, the distortion correction instruction including transformation parameters generated when a user performs a touch operation on the one or more sliders and / or the one or more controls, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the face of the target person and the camera when the shooting terminal shoots the face of the target person; and a processing module for performing a first processing on the first image according to the transformation parameters to obtain a second image; the first processing including distortion correction on the first image; wherein the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0072] In a seventh aspect, this application provides an apparatus comprising: one or more processors; a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in any one of the first to third aspects above.
[0073] Eighthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform the method described in any one of the first to third aspects above.
[0074] Ninthly, this application also provides a computer program product comprising computer program code, which, when run on a computer, causes the computer to perform the method described in any one of the first to third aspects.
[0075] In a tenth aspect, this application also provides an image processing method, the method comprising: acquiring a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; acquiring a target distance; the target distance being used to characterize the distance between the face of the target person and the shooting terminal when the first image is captured; performing a first processing on the first image to obtain a second image; wherein the first processing includes performing distortion correction on the face of the target person in the first image according to the target distance; the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0076] Eleventhly, this application also provides an image processing apparatus, the apparatus comprising: an acquisition module, configured to acquire a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; further configured to acquire a target distance; the target distance being used to characterize the distance between the face of the target person and the shooting terminal when the first image is captured; and a processing module, configured to perform a first processing on the face of the target person in the first image to obtain a second image; wherein the first processing includes distortion correction of the first image based on the target distance; and the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
[0077] According to aspect ten or eleven, in one possible design, the above-described method or apparatus can be applied to a first terminal. If the first terminal captures the first image in real time, the target distance can be understood as the distance between the first terminal and the target person, i.e., the shooting distance between the first terminal and the target person; if the first image is a retrieved historical image, the target distance can be understood as the distance between the second terminal and the target person in the first image when the first image is captured; it should be understood that the first terminal and the second terminal can be the same terminal or different terminals.
[0078] According to aspect ten or eleven, in one possible design, the first processing may also include more existing color processing or image processing methods, such as conventional ISP processing or specific filter processing.
[0079] According to the tenth or eleventh aspect, in one possible design, before performing the first processing on the first image to obtain the second image, the method further includes determining that one or more of the following conditions are met:
[0080] Condition 1: In the first image, the completeness of the target person's face is greater than a preset completeness threshold; or, the facial features of the target person are determined to be complete. This allows the terminal to recognize facial scenes that can undergo distortion correction.
[0081] Condition 2: The pitch angle of the target person's face deviating from a frontal view conforms to the first angle range, and the yaw angle of the target person's face deviating from a frontal view conforms to the second angle range; for example, the first angle range is within [-30°, 30°], and the second angle range is within [-30°, 30°]. That is, the facial image undergoing distortion correction needs to be as close to a frontal view as possible.
[0082] Condition 3: Face detection is performed on the current shooting scene, and only one valid face is detected; the valid face is the face of the target person. That is, distortion correction can be applied to only one subject.
[0083] Condition 4: Determine if the target distance is less than or equal to the preset distortion shooting distance. Shooting distance is usually a crucial factor in causing distortion of people. For example, the distortion shooting distance should not exceed 60cm, such as 60cm or 50cm.
[0084] Condition 5: Determine that the currently enabled camera on the terminal is the front-facing camera, which is used to capture the first image. Front-facing shooting is also a significant cause of human distortion.
[0085] Condition 6: The terminal's current shooting mode is the preset shooting mode, or distortion correction is enabled. This includes portrait shooting, default shooting, or other specific shooting modes; or the terminal has distortion correction enabled.
[0086] It should be understood that the satisfaction of one or more of the above conditions can serve as a decision condition for triggering distortion correction. If some conditions are not met, distortion correction may not be triggered. The above conditions and their free combinations are diverse and are not exhaustive in this invention. They can be flexibly defined according to actual design requirements, and their possible free combinations are also within the scope of protection of this invention. Accordingly, the above device also includes a judgment module for determining whether the above conditions are met.
[0087] According to aspect ten or eleven, in one possible design, acquiring the first image includes: capturing a single frame of an image of a target scene using a camera, and acquiring the first image based on the single frame; or capturing multiple frames of an image of a target scene using a camera, and synthesizing the first image based on the multiple frames; or retrieving the first image from images stored locally or in the cloud. The camera may be that of the aforementioned first terminal; and the camera may be a front-facing or rear-facing camera. Optionally, when the first terminal acquires the first image, the shooting mode may be preset to a pre-defined shooting mode or distortion correction may be enabled. Accordingly, this method may be specifically executed by the acquisition module.
[0088] According to the tenth or eleventh aspect, in one possible design, the distance between the target person's face and the shooting terminal is the distance between the foremost part, center position, eyebrows, eyes, nose, mouth, or ears of the target person's face and the shooting terminal. Obtaining the target distance includes: obtaining the screen-to-body ratio of the target person's face in the first image; obtaining the target distance based on the screen-to-body ratio and the field of view (FOV) of the first image; or, obtaining the target distance based on the Exchangeable Image File Format (EXIF) information of the first image; or, obtaining the target distance through a distance sensor, including a Time-of-Flight (TOF) sensor, a structured light sensor, or a binocular sensor. Accordingly, this method can specifically be executed by the acquisition module.
[0089] According to the tenth or eleventh aspect, in one possible design, distortion correction of the first image based on the target distance specifically includes: obtaining a correction distance; for the same target distance, the larger the value of the correction distance, the greater the degree of distortion correction; obtaining a region of interest; the region of interest includes at least one region among eyebrows, eyes, nose, mouth, or ears; and performing distortion correction on the first image based on the target distance, the region of interest, and the correction distance. Accordingly, this method can be specifically executed by a processing module.
[0090] According to aspect ten or eleven, in one possible design, obtaining the correction distance includes: obtaining the correction distance corresponding to the target distance based on a preset correspondence between shooting distance and correction distance. Accordingly, this method can be specifically executed by a processing module.
[0091] According to the tenth or eleventh aspect, in one possible design, obtaining the correction distance includes: displaying a distortion correction function menu; accepting a control adjustment instruction input by the user based on the distortion correction function menu; the control adjustment instruction being used to determine the correction distance. Displaying the distortion correction function menu includes: displaying a distortion correction control, the distortion correction control being used to open the distortion correction function menu; and displaying the distortion correction function menu in response to the user enabling the distortion correction control. The distortion correction function menu includes controls for determining the correction distance, and / or controls for selecting a region of interest, and / or controls for adjusting facial expressions, and / or controls for adjusting posture, and may be one or more combinations thereof, which are not listed here, and all possible methods should fall within the protection scope of this invention. Accordingly, the display method and operation method may be executed by a display module included in the device.
[0092] According to the tenth or eleventh aspect, in one possible design, distortion correction of the first image based on the target distance and the correction distance includes: fitting the face of the target person in the first image to a standard face model based on the target distance to obtain a first three-dimensional model; the first three-dimensional model is a three-dimensional model corresponding to the face of the target person; performing coordinate transformation on the first three-dimensional model based on the correction distance to obtain a second three-dimensional model; performing perspective projection on the first three-dimensional model based on the coordinate system of the first image to obtain a first set of projection points; performing perspective projection on the second three-dimensional model based on the coordinate system of the first image to obtain a second set of projection points; obtaining a displacement vector, the displacement vector being obtained by aligning the first set of projection points and the second set of projection points; and transforming the first image based on the displacement vector to obtain the second image. Accordingly, this method can be specifically executed by a processing module.
[0093] According to the tenth or eleventh aspect, in one possible design, distortion correction of the first image based on the target distance, the region of interest, and the correction distance includes: fitting the face of the target person in the first image to a standard face model based on the target distance to obtain a first three-dimensional model; the first three-dimensional model is a three-dimensional model corresponding to the face of the target person; performing coordinate transformation on the first three-dimensional model based on the correction distance to obtain a second three-dimensional model; performing perspective projection on the model region corresponding to the region of interest in the first three-dimensional model to obtain a first set of projection points; performing perspective projection on the model region corresponding to the region of interest in the second three-dimensional model to obtain a second set of projection points; obtaining a displacement vector, the displacement vector being obtained by aligning the first set of projection points and the second set of projection points; and transforming the first image based on the displacement vector to obtain the second image. Accordingly, this method can be specifically executed by a processing module.
[0094] According to aspect ten or eleven, in one possible design, when the target distance is within a first preset distance range, the smaller the target distance, the greater the intensity of the distortion correction; wherein the first preset distance range is less than or equal to a preset distortion shooting distance. The greater the angle at which the target person's face deviates from a frontal view, the smaller the intensity of the distortion correction.
[0095] In a twelfth aspect, the present invention also provides a terminal device, the terminal device comprising a memory, a processor, and a bus; the memory and the processor are connected via the bus; wherein the memory is used to store computer programs and instructions; and the processor is used to call the computer programs and instructions stored in the memory to cause the terminal device to execute any of the above-described possible design methods.
[0096] This invention provides a variety of flexible distortion correction processing scenarios and inventive concepts, facilitating users to quickly correct facial distortion while taking photos and browsing image libraries in real time. It also offers users many autonomous choices, along with novel interfaces and features, enhancing the user experience. Without violating natural laws, the optional implementation methods in this invention can be freely combined and transformed, and all of these should fall within the protection scope of this invention. Attached Figure Description
[0097] Figure 1 A schematic diagram illustrating an exemplary application architecture to which the image transformation method of this application is applicable is shown;
[0098] Figure 2 An exemplary structural diagram of terminal 200 is shown;
[0099] Figure 3 This is a flowchart of an embodiment of the image transformation method of this application;
[0100] Figure 4 An exemplary schematic diagram of a distance acquisition method is shown;
[0101] Figure 5 An exemplary schematic diagram illustrating the process of creating a 3D model of a human face is shown.
[0102] Figure 6 An exemplary schematic diagram illustrating the angular change of position movement is shown;
[0103] Figure 7 An exemplary schematic diagram of perspective projection is shown;
[0104] Figure 8a and Figure 8b The effects of face projection at object distances of 30cm and 55cm are shown as examples, respectively.
[0105] Figure 9 An exemplary schematic diagram of a pixel displacement vector dilation method is shown;
[0106] Figure 10 This is a flowchart of Embodiment 2 of the image transformation method of this application;
[0107] Figure 11 An exemplary schematic diagram of the distortion correction function menu is shown;
[0108] Figures 12a-12f The process of distortion correction performed by the terminal in a selfie scenario is illustrated by example;
[0109] Figures 13a-13n An example of a possible interface during the process of image distortion correction is shown;
[0110] Figure 14a This is a flowchart of an image processing method according to this application;
[0111] Figure 14b This is a schematic diagram of the structure of an image processing device according to this application;
[0112] Figure 15 This is a schematic diagram of the structure of an embodiment of the image transformation device of this application. Detailed Implementation
[0113] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0114] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0115] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0116] This application proposes an image transformation method that enables the transformed image, especially the target object in the image, to achieve a realistic transformation effect in three-dimensional space.
[0117] Figure 1 A schematic diagram illustrating an exemplary application architecture to which the image transformation method of this application is applicable is shown, such as... Figure 1 As shown, the framework includes an image acquisition module, an image processing module, and a display module. The image acquisition module is used to capture or acquire images to be processed, and can be, for example, a camera, a video camera, or other similar devices. The image processing module is used to transform the images to be processed to achieve distortion correction. This image processing module can be any device with image processing capabilities, such as a terminal, an image server, or any chip with image processing capabilities, such as a graphics processing unit (GPU) chip. The display module is used to display the images, and can be, for example, a monitor, a terminal screen, a television, a projector, or other similar devices.
[0118] In this application, the image acquisition module, image processing module, and display module can all be integrated into the same device. In this case, the processor of the device acts as a control module, controlling the image acquisition module, image processing module, and display module to perform their respective functions. Alternatively, the image acquisition module, image processing module, and display module can be independent devices. For example, the image acquisition module can be a camera or camcorder, and the image processing module and display module can be integrated into one device. In this case, the processor of the integrated device acts as a control module, controlling the image processing module and display module to perform their respective functions. The integrated device can also have wireless or wired transmission capabilities to receive images to be processed from the image acquisition module. The integrated device can also be equipped with an input interface to acquire images to be processed. Another example is that the image acquisition module uses a camera or camcorder, the image processing module uses a device with image processing capabilities, such as a mobile phone, tablet computer, or computer, and the display module uses a screen or television set. The three devices are connected wirelessly or wiredly to achieve image data transmission. For example, an image acquisition module and an image processing module are integrated into one device. This integrated device has the ability to acquire and process images, such as a mobile phone or tablet computer. The processor of this integrated device acts as a control module, controlling the image acquisition module and the image processing module to perform their respective functions. This device can also have wireless or wired transmission capabilities to transmit images to the display module. This integrated device can also be equipped with an output interface to transmit images through the output interface.
[0119] It should be noted that the above application architecture can also be implemented using other hardware and / or software methods, and this application does not impose specific limitations on it.
[0120] The aforementioned image processing module is the core module of this application. The device that includes the image processing module can be a terminal (such as a mobile phone, tablet, etc.), a wearable device with wireless communication function (such as a smartwatch), a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, etc. This application does not limit it.
[0121] Figure 2 A schematic diagram of the terminal 200 is shown.
[0122] Terminal 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, a time-of-flight (TOF) sensor 280N, a structured light sensor 280O, and a binocular sensor 280P, etc.
[0123] It is understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the terminal 200. In other embodiments of this application, the terminal 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0124] Processor 210 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0125] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0126] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 includes a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0127] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0128] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 210 may include multiple I2C buses. The processor 210 can couple to the touch sensor 280K, charger, flash, camera 293, etc., through different I2C bus interfaces. For example, the processor 210 can couple to the touch sensor 280K through the I2C interface, enabling the processor 210 and the touch sensor 280K to communicate through the I2C bus interface, thereby realizing the touch function of the terminal 200.
[0129] The I2S interface can be used for audio communication. In some embodiments, the processor 210 may include multiple I2S buses. The processor 210 can be coupled to the audio module 270 via the I2S bus to enable communication between the processor 210 and the audio module 270. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0130] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 270 and the wireless communication module 260 can be coupled via the PCM bus interface. In some embodiments, the audio module 270 can also transmit audio signals to the wireless communication module 260 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0131] The UART interface is a universal serial data bus used for asynchronous communication. This bus can include a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 210 and the wireless communication module 260. For example, the processor 210 communicates with the Bluetooth module in the wireless communication module 260 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the UART interface to enable music playback through Bluetooth headphones.
[0132] The MIPI interface can be used to connect the processor 210 to peripheral devices such as the display screen 294 and the camera 293. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 210 and the camera 293 communicate via the CSI interface to enable the shooting function of the terminal 200. The processor 210 and the display screen 294 communicate via the DSI interface to enable the display function of the terminal 200.
[0133] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 210 to a camera 293, a display screen 294, a wireless communication module 260, an audio module 270, a sensor module 280, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0134] USB port 230 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, or USB Type-C port. USB port 230 can be used to connect a charger to charge terminal 200, and can also be used for data transfer between terminal 200 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other terminals, such as AR devices.
[0135] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the terminal 200. In other embodiments of this application, the terminal 200 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0136] The charging management module 240 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 receives charging input from the wired charger via the USB interface 230. In some wireless charging embodiments, the charging management module 240 receives wireless charging input via the wireless charging coil of the terminal 200. While charging the battery 242, the charging management module 240 can also supply power to the terminal via the power management module 241.
[0137] The power management module 241 connects the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240, providing power to the processor 210, internal memory 221, display screen 294, camera 293, and wireless communication module 260, etc. The power management module 241 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 241 may also be located within the processor 210. In other embodiments, the power management module 241 and the charging management module 240 may be located in the same device.
[0138] The wireless communication function of terminal 200 can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor and baseband processor, etc.
[0139] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 200 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0140] The mobile communication module 250 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the terminal 200. The mobile communication module 250 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 250 may be housed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 may be housed in the same device.
[0141] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 270A, receiver 270B, etc.) or displays images or videos through the display screen 294. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 210 and may be housed in the same device as the mobile communication module 250 or other functional modules.
[0142] The wireless communication module 260 can provide solutions for wireless communication applications on the terminal 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 210. The wireless communication module 260 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0143] In some embodiments, antenna 1 of terminal 200 is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, enabling terminal 200 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0144] Terminal 200 implements display functions through a GPU, display screen 294, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0145] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, terminal 200 may include one or N displays 294, where N is a positive integer greater than 1.
[0146] Terminal 200 can perform shooting functions through ISP, camera 293, video codec, GPU, display 294 and application processor.
[0147] The ISP (Image Signal Processor) is used to process data fed back from the camera 293. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 293.
[0148] Camera 293 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, terminal 200 may include one or N cameras 293, where N is a positive integer greater than 1. One or more cameras 293 may be located on the front of terminal 200, such as in the center of the top of the screen, which can be understood as the front-facing camera of the terminal. Corresponding to a binocular sensor device, there may also be two front-facing cameras. Alternatively, one or more cameras 293 may be located on the back of terminal 200, such as in the upper left corner of the back of the terminal, which can be understood as the rear-facing camera of the terminal.
[0149] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal 200 selects a frequency point, the DSP can perform Fourier transforms on the frequency energy.
[0150] Video codecs are used to compress or decompress digital video. Terminal 200 may support one or more video codecs. Thus, terminal 200 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0151] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in terminals, such as image recognition, facial recognition, speech recognition, and text understanding.
[0152] The external storage interface 220 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal 200. The external storage card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0153] Internal memory 221 can be used to store computer executable program code, which includes instructions. Internal memory 221 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of terminal 200 (such as audio data, phonebook, etc.). Furthermore, internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 210 executes various functional applications and data processing of terminal 200 by running instructions stored in internal memory 221 and / or instructions stored in memory located in the processor.
[0154] Terminal 200 can implement audio functions, such as music playback and recording, through audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and application processor.
[0155] The audio module 270 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 may be located in the processor 210, or some functional modules of the audio module 270 may be located in the processor 210.
[0156] The speaker 270A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The terminal 200 can listen to music or make hands-free calls through the speaker 270A.
[0157] The receiver 270B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal 200 receives a phone call or voice message, the receiver 270B can be brought close to the listener's ear to hear the voice.
[0158] Microphone 270C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 270C, inputting the sound signal into microphone 270C. Terminal 200 may have at least one microphone 270C. In some embodiments, terminal 200 may have two microphones 270C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal 200 may have three, four, or more microphones 270C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0159] The headphone jack 270D is used to connect wired headphones. The headphone jack 270D can be a USB 230 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0160] Pressure sensor 280A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 280A can be disposed on display screen 294. There are many types of pressure sensors 280A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 280A, the capacitance between the electrodes changes. Terminal 200 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 294, terminal 200 detects the intensity of the touch operation based on pressure sensor 280A. Terminal 200 can also calculate the touch position based on the detection signal from pressure sensor 280A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example: when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0161] The gyroscope sensor 280B can be used to determine the motion attitude of the terminal 200. In some embodiments, the gyroscope sensor 280B can determine the angular velocity of the terminal 200 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 280B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 280B detects the angle of the terminal 200's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the terminal 200 through reverse movement, thus achieving image stabilization. The gyroscope sensor 280B can also be used in navigation and motion-sensing game scenarios.
[0162] The barometric pressure sensor 280C is used to measure air pressure. In some embodiments, the terminal 200 calculates altitude using the air pressure value measured by the barometric pressure sensor 280C to assist in positioning and navigation.
[0163] The magnetic sensor 280D includes a Hall sensor. The terminal 200 can use the magnetic sensor 280D to detect the opening and closing of the flip cover. In some embodiments, when the terminal 200 is a flip phone, the terminal 200 can detect the opening and closing of the flip cover using the magnetic sensor 280D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.
[0164] The accelerometer 280E can detect the magnitude of acceleration of the terminal 200 in various directions (typically three axes). When the terminal 200 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the terminal's posture and can be applied to applications such as screen orientation switching and pedometers.
[0165] A distance sensor 280F is used to measure distance. Terminal 200 can measure distance via infrared or laser. In some embodiments, during a shooting scene, terminal 200 can utilize the distance sensor 280F for distance measurement to achieve fast focusing.
[0166] The proximity sensor 280G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The terminal 200 emits infrared light outward through the LED. The terminal 200 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 200. When insufficient reflected light is detected, the terminal 200 can determine that there is no object near the terminal 200. The terminal 200 can use the proximity sensor 280G to detect when the user holds the terminal 200 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 280G can also be used in holster mode and pocket mode for automatic unlocking and screen locking.
[0167] The ambient light sensor 280L is used to sense the ambient light intensity. The terminal 200 can adaptively adjust the brightness of the display screen 294 according to the sensed ambient light intensity. The ambient light sensor 280L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 280L can also work with the proximity sensor 280G to detect whether the terminal 200 is in a pocket to prevent accidental touches.
[0168] The fingerprint sensor 280H is used to collect fingerprints. The terminal 200 can use the characteristics of the collected fingerprints to achieve fingerprint unlocking, access to application locks, fingerprint photography, fingerprint answering of incoming calls, etc.
[0169] Temperature sensor 280J is used to detect temperature. In some embodiments, terminal 200 uses the temperature detected by temperature sensor 280J to execute a temperature processing strategy. For example, when the temperature reported by temperature sensor 280J exceeds a threshold, terminal 200 reduces the performance of the processor located near temperature sensor 280J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is below another threshold, terminal 200 heats battery 242 to prevent abnormal shutdown of terminal 200 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, terminal 200 boosts the output voltage of battery 242 to prevent abnormal shutdown due to low temperature.
[0170] Touch sensor 280K, also known as a "touch device," can be located on display screen 294. The touch sensor 280K and display screen 294 together form a touchscreen, also known as a "touchscreen." Touch sensor 280K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 294. In other embodiments, touch sensor 280K may also be located on the surface of terminal 200, in a different position than display screen 294.
[0171] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 280M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 280M can also be incorporated into headphones to form bone conduction headphones. The audio module 270 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 280M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 280M to realize heart rate detection functionality.
[0172] Buttons 290 include a power button, volume buttons, etc. Buttons 290 can be mechanical buttons or touch-sensitive buttons. Terminal 200 can receive button input and generate key signal inputs related to user settings and function control of the terminal 200.
[0173] Motor 291 can generate vibration alerts. Motor 291 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.). Motor 291 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 294. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0174] Indicator 292 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0175] The SIM card interface 295 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to make contact with and separate from the terminal 200. The terminal 200 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 295 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 295 is also compatible with different types of SIM cards. The SIM card interface 295 is also compatible with external memory cards. The terminal 200 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the terminal 200 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 200 and cannot be separated from the terminal 200.
[0176] Figure 3 This is a flowchart of an embodiment of the image transformation method of this application, as shown below. Figure 3 As shown, the method in this embodiment can be applied to Figure 1 The application architecture shown can have the following execution entities: Figure 2 The terminal shown. This image transformation method may include:
[0177] Step 301: Obtain the first image of the target scene using the front-facing camera.
[0178] The first image in this application is captured by the front-facing camera of the terminal in a scenario where the user is taking a selfie. Typically, the FOV of the front-facing camera is set to 70°–110°, preferably 90°. The area directly in front of the front-facing camera is the target scene, which includes the face of the target person (i.e., the user).
[0179] Optionally, the first image may be a preview image captured by the front-facing camera of the terminal and displayed on the screen, at which time the shutter has not been triggered and the first image has not yet been formed on the image sensor; or, the first image may be an image captured by the terminal but not displayed on the screen, at which time the shutter has not been triggered and the first image has not been formed on the sensor; or, the first image may be an image captured by the terminal and formed on the image sensor after the shutter is triggered.
[0180] It should be noted that the first image can also be obtained through the terminal's rear camera, and there are no specific limitations on this.
[0181] Step 302: Obtain the target distance between the target person's face and the front-facing camera.
[0182] The target distance between the target person's face and the front-facing camera can be the distance between the foremost part of the target person's face (e.g., the nose) and the front-facing camera; or, the target distance between the target person's face and the front-facing camera can be the distance between a specific part of the target person's face (e.g., the eyes, mouth, or nose) and the front-facing camera; or, the target distance between the target person's face and the front-facing camera can be the distance between the center position of the target person's face (e.g., the nose in a frontal view, or the cheekbone in a side view) and the front-facing camera. It should be noted that the definition of the above target distance can depend on the specific circumstances of the first image, and this application does not impose a specific limitation on it.
[0183] In one possible implementation, the terminal can obtain the target distance by calculating the screen ratio of the face. That is, first obtain the screen ratio of the target person's face in the first image (the ratio of the pixel area of the face to the pixel area of the first image), and then obtain the distance based on the screen ratio and the field of view (FOV) of the front camera. Figure 4 An exemplary schematic diagram of a distance acquisition method is shown, such as... Figure 4 As shown, assuming the average face length is 20cm and width is 15cm, based on the face's screen-to-body ratio P, the true area of the entire field of view of the first image at the target distance D between the target person's face and the front-facing camera can be estimated as S = 20 × 15 / P cm. 2 The diagonal length L of the first image can be obtained from its aspect ratio. For example, when the aspect ratio of the first image is 1:1, the diagonal length is L = S. 0.5 Based on the imaging relationship shown in the figure above, the target distance D = L / (2×tan(0.5×FOV)).
[0184] In one possible implementation, the terminal can also obtain the target distance via a distance sensor. That is, when a user takes a selfie, the distance between the front-facing camera and the face of the person in front can be measured using the distance sensor on the terminal. This distance sensor could include, for example, a time-of-flight (TOF) sensor, a structured light sensor, or a binocular sensor.
[0185] Step 303: When the target distance is less than a preset threshold, the first image is processed to obtain the second image. The first processing includes distortion correction of the first image based on the target distance.
[0186] In selfies, the distance between the face and the front-facing camera is usually short, leading to perspective distortion due to the "near objects appear larger, far objects appear smaller" effect. For example, when the distance between the subject's face and the front-facing camera is too small, the varying distances from different parts of the face to the camera can cause the nose to appear larger and the face to appear elongated. Conversely, when the distance is greater, these issues may be mitigated. Therefore, this application sets a threshold value. When the distance between the subject's face and the front-facing camera is less than this threshold, the first image containing the subject's face is considered distorted and requires distortion correction. This preset threshold value is within 80 centimeters, and optionally, it can be set to 50 centimeters. It should be noted that the specific value of the preset threshold can be determined based on the performance of the front-facing camera, lighting conditions, etc., and this application does not impose specific limitations on it.
[0187] The terminal corrects the perspective distortion of the first image to obtain a second image. The face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image. That is, the relative proportions and relative positions of the facial features of the target person in the second image are closer to the relative proportions and relative positions of the facial features of the target person in the first image than the relative proportions and relative positions of the facial features of the target person.
[0188] As mentioned above, when the first image is acquired, the distance between the target person's face and the front-facing camera is less than a preset threshold. Due to the "near-to-far" perspective distortion problem of the front-facing camera, the facial features in the first image may appear disproportionately large or stretched, causing their relative proportions and positions to deviate from the true proportions and positions of the target person's face. The second image, obtained by correcting the perspective distortion of the first image, eliminates these disproportionate changes and stretching of the target person's facial features. Therefore, the relative proportions and positions of the facial features in the second image approach or even restore the true proportions and positions of the target person's face.
[0189] Optionally, the second image can be a preview image acquired by the front-facing camera. That is, before the shutter is triggered, the front-facing camera can acquire a first image facing the target scene area (this first image is displayed on the screen as a preview image, or the first image has not appeared on the screen). The terminal uses the method described above to correct the perspective distortion of the first image to obtain the second image, which is then displayed on the terminal's screen. At this time, the second image seen by the user is a preview image that has already undergone perspective distortion correction; the shutter has not been triggered, and the second image has not yet been imaged on the image sensor. Alternatively, the second image can also be an image obtained after the shutter is triggered. That is, after the shutter is triggered, the first image is an image imaged on the terminal's image sensor. The terminal uses the method described above to correct the perspective distortion of the first image to obtain the second image, saves the second image, and displays it on the terminal's screen. At this time, the second image seen by the user is an image that has undergone perspective distortion correction and has been saved in the image library.
[0190] In this application, the process of obtaining a second image by correcting perspective distortion of a first image may include: first, fitting the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; then, correcting perspective distortion of the first image based on the depth information to obtain the second image. The standard face model is a pre-created facial model including facial features, with shape transformation coefficients, expression transformation coefficients, etc. Adjusting these coefficient values can change the expression or shape of the face model. D represents the target distance value. The terminal assumes that the standard 3D face model is placed directly in front of the camera, and the target distance between the standard 3D face model and the camera is D. Specified points on the standard 3D face model (e.g., the tip of the nose, the center point of the face model, etc.) are mapped to the origin O of the 3D coordinate system. A two-dimensional projection point set A of the feature points on the standard 3D face model is obtained through perspective projection, and a point set B of the two-dimensional feature points of the target person's face in the first image is obtained. Each point in point set A has a corresponding unique matching point in point set B. The sum F of the two-dimensional coordinate distance differences of all matching points is calculated. To obtain a realistic 3D model of the target person's face, the planar position (i.e., the specified point is moved up and down, left and right away from the origin O), shape, and relative proportions of facial features of the standard 3D face model can be adjusted multiple times to minimize or even approach zero the sum of distance differences F, meaning that point set A and point set B nearly coincide one-to-one. Based on the above method, a realistic 3D model of the target person's face corresponding to the first image can be obtained. Using the coordinates (x, y, z) of any pixel on the realistic 3D model relative to the origin O, and the target distance D between the origin O and the camera, the coordinates (x, y, z + D) of any pixel on the realistic 3D model of the target person's face relative to the camera can be obtained. Then, through perspective projection, the depth information of any pixel on the target person's face in the first image can be obtained. Optionally, a TOF sensor, structured light sensor, or other methods can be used to directly acquire depth information.
[0191] Based on the above fitting process, the terminal can establish a first 3D model of the target person's face. This fitting process can employ methods such as facial feature point fitting, deep learning fitting, and TOF camera depth fitting. The first 3D model can be presented as a 3D point cloud, corresponding to the 2D image of the target person's face, exhibiting the same facial expressions, relative proportions of facial features, and relative positions.
[0192] For example, a three-dimensional model of the target person's face is established by fitting facial feature points. Figure 5 An exemplary schematic diagram illustrating the process of creating a 3D model of a human face is shown, such as... Figure 5As shown, the process first extracts feature points of facial features and contours from the input image. Then, using a deformable basic 3D face model, based on the correspondence between the 3D key points of the face model and the obtained feature points, and according to fitting parameters (target distance, FOV coefficient, etc.) and fitting optimization terms (face rotation, translation, scaling parameters, face model deformation coefficient, etc.), a realistic face shape is fitted. Since the corresponding positions of the facial contour feature points on the 3D face model change with the angle of the face, different correspondences under different angles can be selected, and multiple iterations of fitting are used to ensure fitting accuracy. Finally, the fitted output is the 3D spatial coordinates of each point on the 3D point cloud that fits the target person's face. The distance from the camera to the nose is used to build a 3D face model. A standard model is placed at this distance and projected. The coordinate distance difference between the two (photo and 2D projection) is minimized, and the projection is continuously repeated to obtain a 3D head. The depth information is a vector. Fitting is not only about fitting the shape, but also about fitting the model's spatial position along the x and y axes.
[0193] The terminal transforms the pose and / or shape of the first 3D model of the target person's face to obtain a second 3D model of the target person's face. Optionally, the terminal performs coordinate transformation on the first 3D model based on the correction distance to obtain the second 3D model.
[0194] Adjusting the pose and / or shape of the first 3D model, and performing coordinate transformations on the first 3D model, all require adjustments based on the distance between the target person's face and the front-facing camera. Since the distance between the target person's face and the front-facing camera in the first image is less than a preset threshold, the first 3D model can be moved backward to obtain a second 3D model of the target person's face. Adjustments based on the 3D model can simulate adjustments to the target person's face in the real world; therefore, the adjusted second 3D model can show the result of the target person's face being moved backward relative to the front-facing camera.
[0195] In one possible implementation, the first 3D model is moved backward along the extension line between the first 3D model and the camera to a position twice the original distance to obtain the second 3D model. That is, if the spatial coordinates of the original first 3D model are (x, y, z), the position of the second 3D model is (2x, 2y, 2z). In other words, there exists a mapping relationship of 2x.
[0196] In one possible implementation, when the second 3D model is obtained by moving the first 3D model, angle compensation is performed on the second 3D model.
[0197] In the real world, when a person's face moves backward relative to the front-facing camera, the angle of the person's face relative to the camera may change. If you do not want to retain the angle change caused by the backward movement, you can perform angle compensation. Figure 6 An exemplary schematic diagram illustrating the angular change of position movement is shown, such as... Figure 6 As shown, initially, the target person's face is less than a preset threshold away from the front-facing camera. Its angle relative to the front-facing camera is α, its vertical distance to the camera is tz1, and its horizontal distance is tx. When the target person's face moves backward, its angle relative to the front-facing camera changes to β, its vertical distance to the camera changes to tz2, while its horizontal distance remains tx. Therefore, the changing angle of the target person's face relative to the front-facing camera... If the first 3D model is simply moved backward, the 2D image obtained based on the second 3D model will have an angular change due to the backward movement. Therefore, the second 3D model needs to be angle compensated, that is, rotated by an angle of Δθ so that the final 2D image obtained maintains the same posture as the 2D image before the backward movement (i.e., the face of the target person in the original image).
[0198] The terminal obtains the pixel displacement vector field of the target object based on the first 3D model and the second 3D model.
[0199] Optionally, after obtaining the two 3D models before and after the transformation (the first 3D model and the second 3D model mentioned above), this application performs perspective projection on the first 3D model based on the depth information to obtain a first coordinate set, which includes 2D coordinate values obtained by projecting multiple sampling points in the first 3D model. Then, it performs perspective projection on the second 3D model based on the depth information to obtain a second coordinate set, which includes 2D coordinate values obtained by projecting multiple sampling points in the second 3D model.
[0200] Optionally, after obtaining the two 3D models before and after the transformation (the first 3D model and the second 3D model mentioned above), based on the coordinate system of the first image, the first 3D model is subjected to perspective projection to obtain a first set of projection points; based on the coordinate system of the first image, the second 3D model is subjected to perspective projection to obtain a second set of projection points; the first set of projection points includes two-dimensional coordinate values obtained by projecting multiple sampling points in the first 3D model. The second set of projection points includes two-dimensional coordinate values obtained by projecting multiple sampling points in the second 3D model.
[0201] Projection is a method of transforming three-dimensional coordinates into two-dimensional coordinates. Common projection methods include orthographic projection and perspective projection. Taking perspective projection as an example, the basic perspective projection model consists of two parts: a viewpoint E and a view plane P. The viewpoint E is not on the view plane P. The viewpoint E can be considered as the position of the camera. The view plane P is the two-dimensional plane used to render the perspective view of the three-dimensional target object. Figure 7 An exemplary schematic diagram of perspective projection is shown, such as... Figure 7 As shown, for any point X in the real world, a ray originating from viewpoint E and passing through point X is constructed. The intersection point Xp of this ray and the view plane P is the perspective projection of point X. Objects in the three-dimensional world can be considered as being composed of a set of points {Xi}. Thus, rays Ri originating from viewpoint E and passing through point Xi are constructed. The set of intersection points of these rays Ri with the view plane P is the two-dimensional projection of the three-dimensional object at viewpoint E.
[0202] Based on the above principle, each sampling point of the 3D model is projected to obtain the corresponding pixel on a 2D plane. These pixels on the 2D plane can be represented by a coordinate value in the 2D plane, thus obtaining a coordinate set corresponding to the sampling points of the 3D model. This application can obtain a first coordinate set corresponding to the first 3D model and a second coordinate set corresponding to the second 3D model.
[0203] Optionally, to maintain the consistent size and position of the target person's face, appropriate alignment points and scaling scales can be selected to translate or scale the coordinate values in the second coordinate set. The principle of translation or scaling adjustment can be to minimize the displacement of the target person's face edge relative to the surrounding area, which includes the background area or the edge of the field of view. For example, when the lowest point of the first coordinate set is located at the boundary of the edge of the field of view, while the lowest point of the second coordinate set is too large and deviates from the boundary, translating the second coordinate set makes the lowest point coincide with the boundary of the edge of the field of view.
[0204] The pixel displacement vector field of the target person's face is obtained by calculating the coordinate difference between the first coordinate value and the second coordinate value. The first coordinate value is the coordinate value corresponding to the first sampling point in the first coordinate set, and the second coordinate value is the coordinate value corresponding to the first sampling point in the second coordinate set. The first sampling point is any point in the multiple identical point clouds contained in the first 3D model and the second 3D model.
[0205] The second 3D model is obtained by transforming the pose and / or shape of the first 3D model. Therefore, both models contain a large number of identical sampling points, and some even contain completely identical sampling points. This means that many sets of coordinate values in the first and second coordinate sets correspond to the same sampling point; that is, a sampling point corresponds to one coordinate value in the first coordinate set and also to one coordinate value in the second coordinate set. The coordinate difference between the first coordinate value and the translated or scaled second coordinate value is calculated, specifically by calculating the difference in the x-axis and y-axis coordinates of the first and second coordinate values to obtain the coordinate difference of the first pixel. By calculating the coordinate differences of all identical sampling points in the first and second 3D models, the pixel displacement vector field of the target person's face can be obtained. This pixel displacement vector field is composed of the coordinate differences of each sampling point.
[0206] In one possible implementation, when the coordinate difference between the coordinate value of the pixel at the edge position of the target person's face in the second coordinate set and the coordinate value of the pixel in the surrounding area is greater than a preset threshold, the coordinate value in the second coordinate set is translated or scaled, and the surrounding area is adjacent to the target person's face.
[0207] To maintain the consistent size and position of the target person's face, when the edge position of the target person's face is displaced too much (which can be measured by a preset threshold) and causes background distortion, an appropriate alignment point and scaling scale can be selected to translate or scale the coordinate values in the second coordinate set. Figure 8a and Figure 8b The effects of face projection at object distances of 30cm and 55cm are shown as examples, respectively.
[0208] The terminal obtains the transformed image based on the pixel displacement vector field of the target person's face.
[0209] In one possible implementation, the target person's face, the edge region of the field of view, and the background region are constrained and corrected by an algorithm based on the pixel displacement vector field of the target person's face to obtain the transformed image. The edge region of the field of view is the strip region located at the edge of the image, and the background region is the other regions in the image except for the target person's face and the edge region of the field of view.
[0210] The image is divided into three regions: one is the area occupied by the face of the target person, another is the edge region of the image (i.e., the edge region of the field of view), and the third is the background region (i.e., the background part outside the face of the target person, which does not include the edge region of the image).
[0211] This application can determine the initial image matrices corresponding to the face, field of view edge region, and background region of the target person based on the pixel displacement vector field of the target person's face; construct constraint terms corresponding to the face, field of view edge region, and background region of the target person, and construct regularization constraint terms for the image; obtain the pixel displacement matrices corresponding to the face, field of view edge region, and background region of the target person based on the constraint terms and regularization constraint terms corresponding to the face, field of view edge region, and background region of the target person, and the weight coefficients corresponding to each constraint term; and obtain the transformed image through color mapping based on the initial image matrices corresponding to the face, field of view edge region, and background region of the target person and the pixel displacement matrices corresponding to the face, field of view edge region, and background region of the target person.
[0212] In one possible implementation, the pixel displacement vector field of the target person's face is expanded using an interpolation algorithm to obtain the pixel displacement vector field of the mask region, which includes the target person's face; the mask region, the field of view edge region, and the background region are then constrained and corrected using an algorithm based on the pixel displacement vector field of the mask region to obtain the transformed image.
[0213] The interpolation algorithm may include assigning the pixel displacement vector of the first sampling point to the second sampling point as the pixel displacement vector of the second sampling point. The second sampling point is any sampling point located outside the face region of the target person but within the mask region. The first sampling point is the pixel point on the boundary contour of the target person's face that is closest to the second pixel point.
[0214] Figure 9 An exemplary schematic diagram of a pixel displacement vector dilation method is shown, such as... Figure 9 As shown, the target region is the face of a person, and the mask region is the area of the head of a person. The mask region includes the face of the person. The pixel displacement vector field of the face is extended to the entire mask region using the above interpolation algorithm to obtain the pixel displacement vector field of the mask region.
[0215] The image is divided into four regions: one is the region occupied by the face; another is a region related to the face, which can change accordingly with the pose and / or shape of the target object; for example, when the face rotates, the head will also rotate. The face and the above-mentioned region form the mask region. The third is the edge region of the image (i.e., the edge of the field of view); and the fourth is the background region (i.e., the background part outside the target object, which does not include the edge region of the image).
[0216] This application determines the initial image matrices corresponding to the mask region, the field of view edge region, and the background region based on the pixel displacement vector field of the mask region; constructs constraint terms corresponding to the mask region, the field of view edge region, and the background region, and constructs regularization constraint terms for the image; obtains the pixel displacement matrices corresponding to the mask region, the field of view edge region, and the background region based on the constraint terms and regularization constraint terms corresponding to the mask region, the field of view edge region, and the background region, as well as the weight coefficients corresponding to each constraint term; and obtains the transformed image through color mapping based on the initial image matrices corresponding to the mask region, the field of view edge region, and the background region, and the pixel displacement matrices corresponding to the mask region, the field of view edge region, and the background region.
[0217] The following describes the constraints for the mask region, background region, and field of view edge region, as well as the regularization constraints used for the global image.
[0218] (1) The constraint term corresponding to the mask region is used to constrain the target image matrix corresponding to the mask region in the image to be an image matrix after geometric transformation using the pixel displacement vector field of the mask region in the previous step, so as to correct the distortion of the mask region. This geometric transformation represents a spatial mapping, that is, transforming a pixel displacement vector field into another image matrix. The geometric transformation in this application can be at least one of image translation, image scaling, and image rotation.
[0219] For ease of description, the constraints corresponding to the mask region can be simply referred to as mask constraints. When there are multiple target objects in the image, different target objects can correspond to different mask constraints.
[0220] The mask constraint term can be denoted as Term1, and the formula for Term1 is as follows:
[0221] Term1(i,j)=SUM (i,j)∈HeadRegionk ||M0(i,j)+Dt(i,j)-Func1 k [M1(i,j)]||
[0222] Specifically, for regions located in the head region (i.e., (i,j)... ∈HeadRegionk Let M0(i,j) be the image matrix of pixels, and M1(i,j) = [u1(i,j), v1(i,j)] be the coordinates of the target object after shape preservation. T Dt(i,j) represents the displacement matrix corresponding to M0(i,j), k represents the k-th mask region of the image, and Func1 k Let ||k|| represent the geometric transformation function corresponding to the k-th mask region, and ||...|| represent the vector 2 norm.
[0223] The mask constraint term Term1(i,j) must ensure that under the action of the translation matrix Dt(i,j), the image matrix M0(i,j) tends to be an appropriate geometric transformation of M1(i,j), including at least one of the transformation operations of image rotation, image translation and image scaling.
[0224] The geometric transformation function Func1 corresponding to the k-th mask region k This indicates that all points within the k-th mask region share the same geometric transformation function Func1. k Different mask regions correspond to different geometric transformation functions. The geometric transformation function is Func1. k This can be specifically expressed as:
[0225]
[0226] Where, ρ 1k θ represents the scaling factor for the k-th mask region. 1k TX represents the rotation angle of the k-th mask region. 1k and TY 1k These represent the lateral and longitudinal displacements of the k-th mask region, respectively.
[0227] Term1(i,j) can be specifically represented as:
[0228]
[0229] Here, du(i,j) and dv(i,j) are the unknowns that need to be solved, and this term should be kept as small as possible when solving the constraint equations.
[0230] (2) The constraint terms corresponding to the field of view edge region are used to constrain the pixels in the initial image matrix corresponding to the field of view edge region in the image to be displaced along the edge of the image or to the outside of the image, so as to maintain or expand the field of view edge region.
[0231] For ease of description, the constraint terms corresponding to the edge region of the field of view can be simply referred to as the field of view edge constraint terms. The field of view edge constraint terms can be denoted as Term3, and the formula for Term3 is as follows:
[0232] Term3(i,j)=SUM (i,j)∈EdgeRegion ||M0(i,j)+Dt(i,j)-Func3 (i,j) [M0(i,j)]||
[0233] Where M0(i,j) represents the region located at the edge of the field of view (i.e., (i,j)). ∈EdgeRegion The image coordinates of the pixel M0(i,j), where Dt(i,j) represents the displacement matrix corresponding to M0(i,j), and Func3 (i,j)Let M0(i,j) denote the displacement function, and ||…|| denote the vector 2 norm.
[0234] The field of view edge constraint term must ensure that under the action of the displacement matrix Dt(i,j), the image matrix M0(i,j) tends to be an appropriate displacement of the coordinate value M0(i,j). The displacement rule is to move only along the edge region or appropriately to the outside of the edge region, and avoid moving to the inside of the edge region. The advantage of doing this is that it can minimize the loss of image information caused by subsequent rectangular cropping, and can even gain and expand the image content of the field of view edge region.
[0235] Suppose a pixel A is located at the edge of the image's field of view, with image coordinates [u0, v0]. T The tangential vector of point A along the field of view boundary is denoted as y(u0,v0), and the normal vector towards the outside of the image is denoted as x(u0,v0). When the boundary region is known, x(u0,v0) and y(u0,v0) are also known. Then Func3 (i,j) This can be specifically expressed as:
[0236]
[0237] In this context, α(u0(i,j),v0(i,j)) must be restricted to be no less than 0 to ensure that the point does not shift inwards towards the edge of the field of view. The sign of β(u0(i,j),v0(i,j)) is not restricted. α(u0(i,j),v0(i,j)) and β(u0(i,j),v0(i,j)) are intermediate unknowns that do not need to be explicitly solved.
[0238] Term3(i,j) can be specifically represented as:
[0239]
[0240] Here, du(i,j) and dv(i,j) are the unknowns that need to be solved, and this term should be kept as small as possible when solving the constraint equations.
[0241] (3) The constraint term corresponding to the background region is used to constrain the pixel points in the image matrix corresponding to the background region in the image to be displaced, and the first vector corresponding to the pixel point before displacement and the second vector corresponding to the pixel point after displacement should be kept as parallel as possible, so as to make the image content in the background region smooth and continuous and to make the image content in the background region that runs through the human image continuous and consistent in human vision; wherein, the first vector represents the vector between the pixel point before displacement and the neighboring pixel point corresponding to the pixel point before displacement; the second vector represents the vector between the pixel point after displacement and the neighboring pixel point corresponding to the pixel point after displacement.
[0242] For ease of description, the constraint terms corresponding to the background area can be simply referred to as background constraint terms, which can be denoted as Term4. The formula for Term4 is as follows:
[0243] Term4(i,j)=SUM (i,j)∈BkgRegion {Func4 (i,j) (M0(i,j),M0(i,j)+Dt(i,j))}
[0244] Where M0(i,j) represents the region located in the background (i,j) ∈BkgRegion The image coordinates of the pixel M0(i,j), where Dt(i,j) represents the displacement matrix corresponding to M0(i,j), and Func4 (i,j) Let M0(i,j) denote the displacement function, and ||…|| denote the vector 2 norm.
[0245] The background constraint term must ensure that the coordinate value M0(i,j) under the action of the displacement matrix Dt(i,j) tends to be an appropriate displacement of the coordinate value M0(i,j). In this application, each pixel in the background region can be divided into different control domains. This application does not limit the size, shape, or number of control domains. In particular, for background pixels located at the boundary between the target object and the background, their control domain must extend across the target object to the other end of the target object. Assume there exists a background pixel A and its control domain pixel set {Bi}, where the control domain is the neighborhood of point A. The control domain of A extends across the intermediate mask region to the other end of the mask region. Bi represents the neighboring pixels of A, which are moved to A′ and {B′i} respectively after displacement. The displacement rule is that the background constraint term will restrict the vectors ABi and A′B′i to keep their directions as parallel as possible. The advantage of doing so is that it can ensure a smooth transition between the target object and the background region, and the image content running through the human figure in the background region can be continuous and consistent in human vision, avoiding distortion or hollow streaks in the background image. Func4 (i,j) This can be specifically expressed as:
[0246]
[0247] in:
[0248]
[0249]
[0250]
[0251] Where angle[] represents the angle between two vectors, and vec1 represents the foreground and background points [i,j] after correction. TGiven the vector formed by the points within its control domain, vec2 represents the corrected background point [i,j]. T SUM is the vector formed by a point within its control domain after correction. (i+di,j+dj)∈CtrlRegion This represents the summation of the angles between all vectors within the control domain.
[0252] Term4(i,j) can be specifically represented as:
[0253] Term4(i,j)=SUM (i,j)∈BkgRegion {SUM (i+di,j+dj)∈CtrlRegion {angle[vec1,vec2]}}
[0254] Here, du(i,j) and dv(i,j) are the unknowns that need to be solved, and this term should be kept as small as possible when solving the constraint equations.
[0255] (4) The regularization constraint term is used to constrain the difference between the displacement matrices of any two adjacent pixels in the displacement matrices corresponding to the background region, mask region and field edge region of the image to be less than a preset threshold, so as to make the global image content of the image smooth and continuous.
[0256] The regularity constraint term can be denoted as Term5, and the formula for Term5 is as follows:
[0257] Term5(i,j)=SUM (i,j)∈AllRegion {Func5 (i,j) (Dt(i,j))}
[0258] For the entire image range (i.e., (i,j)) ∈AllRegion For pixel M0(i,j), the regularization constraint must ensure that the translation matrix Dt(i,j) of adjacent pixels is smooth and continuous to avoid excessively large local jumps. The constraint principle is that for point [i,j]... T The difference between the displacement at a given point and the displacements of its neighboring points (i+di,j+dj) should be as small as possible (i.e., less than a certain threshold). Func5 (i,j) This can be specifically expressed as:
[0259]
[0260] Term5(i,j) can be specifically represented as:
[0261]
[0262] Here, du(i,j) and dv(i,j) are the unknowns that need to be solved, and this term should be kept as small as possible when solving the constraint equations.
[0263] Based on each constraint term and its corresponding weight coefficient, the displacement matrix corresponding to each region is obtained.
[0264] Specifically, weight coefficients can be set for the constraints and regularization constraints of the mask region, the edge region of the field of view, and the background region. Constraint equations can be established based on each constraint and its corresponding weight coefficients. Solving these constraint equations will yield the offset of each position point in each region.
[0265] Suppose the coordinate matrix of the image after algorithmic constraint correction (also known as the target image matrix) is Mt(i,j), where Mt(i,j) = [ut(i,j), vt(i,j)]. T Its displacement matrix compared to the image matrix M0(i,j) is Dt(i,j), where Dt(i,j) = [du(i,j),dv(i,j)]. T In other words:
[0266] Mt(i,j)=M0(i,j)+Dt(i,j)
[0267] ut(i,j)=u0(i,j)+du(i,j)
[0268] vt(i,j)=v0(i,j)+dv(i,j)
[0269] Assign weight coefficients to each constraint term and construct the constraint equations as follows:
[0270] Dt(i,j)=(du(i,j),dv(i,j))
[0271] =argmin(α1(i,j)×Term1(i,j)+α2(i,j)×Term2(i,j)+α3(i,j)×Term3(i,j)+α4(i,j)×Term4(i,j)+α5(i,j)×Term5(i,j))
[0272] Where α1(i,j)~α5(i,j) are the weight coefficients (weight matrices) corresponding to Term1~Term5 respectively.
[0273] By solving the constraint equation using the least squares method, gradient descent method, or various improved algorithms, the displacement matrix Dt(i,j) of each pixel in the image is finally obtained. Based on this displacement matrix Dt(i,j), the transformed image can be obtained.
[0274] This application achieves a three-dimensional transformation effect of the image by using a three-dimensional model, and corrects the perspective distortion of the face of the target person in the close-up image. This makes the relative proportions and relative positions of the facial features of the target person after correction closer to the relative proportions and relative positions of the facial features of the target person, which can significantly improve the shooting imaging effect in selfie scenarios.
[0275] In one possible implementation, for video recording scenarios, the terminal can adopt... Figure 3 The method in the illustrated embodiment performs distortion correction processing on multiple image frames in the recorded video to obtain a distortion-corrected video. The terminal can directly play the distortion-corrected video on the screen, or the terminal can use a split-screen method, displaying the uncorrected video in one area and the corrected video in another area of the screen. (See also...) Figure 12f .
[0276] Figure 10 This is a flowchart of Embodiment 2 of the image transformation method of this application, as follows: Figure 10 As shown, the method in this embodiment can be applied to Figure 1 The application architecture shown can have the following execution entities: Figure 2 The terminal shown. This image transformation method may include:
[0277] Step 1001: Obtain the first image.
[0278] In this application, the first image is stored in the image library of the second terminal. The first image can be a photograph taken by the second terminal or a frame from a video taken by the second terminal. This application does not specifically limit the method of obtaining the first image.
[0279] It should be noted that the first terminal and the second terminal that perform distortion correction processing on the first image can be the same device or different devices.
[0280] In one possible implementation, the second terminal, acting as the device for acquiring the first image, can be any device with shooting capabilities, such as a camera or camcorder. The second terminal stores the captured image locally or in the cloud. The first terminal, acting as the processing device for distortion correction of the first image, can be any device with image processing capabilities, such as a mobile phone, computer, or tablet computer. The first terminal can receive the first image from the second terminal or the cloud via wired or wireless communication. Alternatively, the first terminal can acquire the first image captured by the second terminal via a storage medium (such as a USB flash drive).
[0281] In one possible implementation, the first terminal has both shooting and image processing functions, such as a mobile phone or tablet computer. The first terminal obtains the first image from a local image library, or the first terminal takes a picture and obtains the first image based on the instruction triggered by the shutter being pressed.
[0282] In one possible implementation, the first image includes the face of the target person. Distortion in the target person's face in the first image is caused by the target distance between the second terminal and the target person's face being less than a first preset threshold when the second terminal captures the first image. That is, the target distance between the target person's face and the camera is small when the second terminal captures the first image. Typically, when using a front-facing camera for selfies, the distance between the face and the camera is small, resulting in a "near-large, far-small" perspective distortion problem. For example, when the target distance between the target person's face and the front-facing camera is too small, the difference in distance from different parts of the face to the camera may cause the nose to appear larger and the face to appear elongated in the image. Conversely, when the distance between the target person's face and the front-facing camera is larger, these problems may be mitigated. Therefore, this application sets a threshold, considering that when the distance between the target person's face and the front-facing camera is less than this threshold, the acquired first image containing the target person's face is distorted and needs distortion correction. The preset threshold ranges from 80 centimeters, and optionally, it can be set to 50 centimeters. It should be noted that the specific value of the preset threshold may vary depending on the performance of the front camera, the lighting conditions, etc., and this application does not impose any specific restrictions on it.
[0283] The target distance between the target person's face and the camera can be the distance between the foremost part of the target person's face (e.g., the nose) and the camera; or, the target distance between the target person's face and the camera can be the distance between a specific part of the target person's face (e.g., the eyes, mouth, or nose) and the camera; or, the target distance between the target person's face and the camera can be the distance between the center of the target person's face (e.g., the nose in a frontal view, or the cheekbone in a side view) and the camera. It should be noted that the definition of the target distance can vary depending on the specific circumstances of the first image, and this application does not impose any specific limitations on it.
[0284] The terminal can obtain the target distance based on the screen-to-body ratio of the target person's face in the first image and the FOV of the second terminal's camera. It should be noted that when the second terminal includes multiple cameras, after capturing the first image, the information of the camera that captured the first image is recorded in the exchangeable image file format (EXIF) information. Therefore, the FOV of the second terminal mentioned above refers to the FOV of the camera recorded in the EXIF information. The principle is the same as step 302 above, and will not be repeated here. The FOV can be obtained from the FOV in the EXIF information of the first image, or calculated from the equivalent focal length in the EXIF information, for example, fov = 2.0 × atan(43.27 / 2f), where 43.27 is the diagonal length of 135mm film, and f represents the equivalent focal length. The terminal can also obtain the target distance based on the target shooting distance saved in the EXIF information of the first image. EXIF is specifically designed for digital camera photos. It records the attribute information and shooting data of digital photos. The terminal can directly read the target shooting distance, FOV, or equivalent focal length at the time the first image was taken from the EXIF information, and thus obtain the aforementioned target distance. The principle can also refer to step 302 above, and will not be repeated here.
[0285] In one possible implementation, the first image includes the face of the target person. Distortion in the target person's face in the first image is caused by the second terminal capturing the first image when the camera's field of view (FOV) exceeds a second preset threshold, and the pixel distance between the target person's face and the edge of the FOV is less than a third preset threshold. If the terminal's camera is a wide-angle camera, distortion will also occur when the target person's face is at the edge of the camera's FOV, while distortion will be reduced or even eliminated when the target person's face is in the middle of the FOV. Therefore, this application sets two thresholds: when the terminal's camera's FOV exceeds the corresponding threshold, and the pixel distance between the target person's face and the edge of the FOV is less than the corresponding threshold, the acquired first image containing the target person's face is distorted and requires distortion correction. The threshold corresponding to the FOV is 90°, and the threshold corresponding to the pixel distance is one-quarter of the length or width of the first image. It should be noted that the specific value of the threshold can be determined depending on the camera's performance, shooting lighting, etc., and this application does not specifically limit it.
[0286] The pixel distance can be the number of pixels between the foremost part of the target person's face and the boundary of the first image; or, the pixel distance can be the number of pixels between a specified part of the target person's face and the boundary of the first image; or, the pixel distance can be the number of pixels between the center of the target person's face and the boundary of the first image.
[0287] Step 1002: Display the distortion correction function menu.
[0288] When the target person's face is distorted, the terminal displays a pop-up window. This pop-up provides a selection control for whether to perform distortion correction, for example... Figure 13c As shown; or the terminal displays a distortion correction control, which is used to open the distortion correction function menu, for example. Figure 12d As shown. When the user clicks the "Yes" control or the distortion correction control, the distortion correction function menu is displayed in response to the user's operation.
[0289] This application provides a distortion correction function menu, which includes options for changing transformation parameters, such as adjusting the equivalent simulated shooting distance, adjusting the displacement distance, adjusting the relative position and / or proportion of facial features, etc. These options can be adjusted using sliders or controls; that is, the transformation parameter's value can be changed by adjusting the position of a slider, or its selected value can be determined by triggering a control. When the distortion correction function menu is initially displayed, the transformation parameter values for each option can be default values or pre-calculated values. For example, the slider corresponding to the equivalent simulated shooting distance can initially be set to a value of 0. Figure 11 As shown, or, the terminal obtains an equivalent simulated shooting distance adjustment amount based on the image transformation algorithm, and displays the slider at the position corresponding to the adjustment amount, such as... Figure 13e As shown.
[0290] Figure 11 An exemplary schematic diagram of the distortion correction function menu is shown, such as... Figure 11As shown, the function menu includes four areas: the left side of the function menu displays the first image, which is placed under a coordinate axis including xyz axes; the upper right side of the function menu includes two controls, one for saving the image and the other for starting the image recording transformation process, which can be triggered to start the corresponding operation; the lower right side of the function menu includes expression controls and action controls. The expression controls offer six selectable expression templates: remain unchanged, happy, sad, disgusted, surprised, and angry. When the user clicks the corresponding expression control, the expression template is selected. The terminal can then change the relative proportions and positions of the facial features in the left-side image according to the selected expression template. For example, when happy, the eyes usually shrink and the corners of the mouth curve upward; when surprised, the eyes and mouth usually widen; when angry, the eyebrows usually furrow and the corners of the eyes and mouth curve downward. The action templates available under the action control include six actions: no action, nodding, shaking the head left and right, Xinjiang dance head shaking, blinking, and laughing. When the user clicks the corresponding action control, the action template is selected. The terminal can change the facial features of the person in the left part of the image multiple times according to the action sequence based on the selected action template. For example, nodding typically involves two actions: looking down and looking up; shaking the head left and right typically involves two actions: shaking the head to the left and shaking the head to the right; blinking typically involves two actions: closing the eyes and opening the eyes. Below the function menu are four sliders, corresponding to distance, left / right head shaking, up / down head shaking, and clockwise head turning. The distance slider adjusts the distance between the face and the camera by moving it left or right; the left / right head shaking slider adjusts the direction (left or right) and angle of the head shaking; the up / down head shaking slider adjusts the direction (up or down) and angle of the head shaking; and the clockwise head turning slider adjusts the angle of the head turning clockwise or counterclockwise. Users can adjust the corresponding transformation parameter values by controlling one or more of these sliders or controls. The terminal obtains the corresponding transformation parameter values after detecting the user's operation.
[0291] Step 1003: Obtain the transformation parameters entered by the user in the distortion correction function menu.
[0292] The transformation parameters are obtained based on the user's actions on the options included in the distortion correction function menu (e.g., based on the user's dragging of the slider associated with the transformation parameters; and / or based on the user's triggering of the control associated with the transformation parameters).
[0293] In one possible implementation, the distortion correction function menu includes an option to adjust the displacement distance; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction and displacement distance based on the instructions triggered by the user's operation of the control or slider in the option to adjust the displacement distance.
[0294] In one possible implementation, the distortion correction function menu includes options for adjusting the relative position and / or relative proportion of facial features; and obtains transformation parameters input by the user on the distortion correction function menu, including: obtaining adjustment direction, displacement distance and / or facial feature size based on instructions triggered by the user's operation of controls or sliders in the options for adjusting the relative position and / or relative proportion of facial features.
[0295] In one possible implementation, the distortion correction function menu includes an option to adjust the angle; the transformation parameters input by the user on the distortion correction function menu are obtained, including: obtaining the adjustment direction and adjustment angle based on the instructions triggered by the user's operation of the control or slider in the adjustment angle option.
[0296] In one possible implementation, the distortion correction function menu includes options for adjusting facial expressions; obtaining transformation parameters input by the user on the distortion correction function menu includes: obtaining a new facial expression template based on instructions triggered by the user's operation of controls or sliders in the options for adjusting facial expressions.
[0297] In one possible implementation, the distortion correction function menu includes options for adjusting the action; obtaining transformation parameters input by the user on the distortion correction function menu includes: obtaining a new action template based on instructions triggered by the user's operation of controls or sliders in the adjustment action options.
[0298] Step 1004: Perform a first process on the first image to obtain a second image. The first process includes distortion correction of the first image according to the transformation parameters.
[0299] The terminal corrects the perspective distortion of the first image according to the transformation parameters to obtain the second image. The face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image. That is, the relative proportions and relative positions of the facial features of the target person in the second image are closer to the relative proportions and relative positions of the facial features of the target person in the first image than the relative proportions and relative positions of the facial features of the target person.
[0300] The second image obtained by correcting perspective distortion in the first image can eliminate the changes in size and stretching of the facial features of the target person. Therefore, the relative proportions and relative positions of the facial features of the target person in the second image are close to or even restored to the relative proportions and relative positions of the real appearance of the target person's face.
[0301] In one possible implementation, the terminal can also obtain a recording command based on the user's triggering operation of the start recording control, and then start recording the process of acquiring the second image. At this time, the start recording control becomes a stop recording control. When the user triggers the stop recording control, the terminal can receive the stop recording command and then stop recording.
[0302] The recording process in this application may include at least two of the following scenarios:
[0303] One scenario involves a user clicking the "Start Recording" control before distortion correction begins on a first image. The terminal responds to this command by activating screen recording, recording the terminal's screen. The user then clicks or drags the slider on the distortion correction menu to set transformation parameters. As the user interacts with the image, the terminal performs distortion correction or other transformations on the first image based on these parameters, displaying the processed second image on the screen. The entire process, from manipulating the distortion correction menu controls to the transformation from the first to the second image, is recorded by the terminal's screen recording function. When the user clicks the "Stop Recording" control, the terminal stops recording, and the terminal has acquired the video of this entire process.
[0304] Another scenario involves a user clicking the "Start Recording" control before distortion correction begins on a video. The terminal responds to this command by activating screen recording, recording the terminal's screen. The user then clicks the control or drags the slider on the distortion correction menu to set transformation parameters. As the user interacts with these parameters, the terminal performs distortion correction or other transformations on multiple image frames in the video, displaying the processed video on the screen. The entire process, from manipulating the distortion correction menu controls to playing the processed video, is recorded by the terminal's screen recording function. When the user clicks the "Stop Recording" control, the terminal closes the screen recording function, and the terminal then captures the video of the entire process.
[0305] In one possible implementation, the terminal can also obtain a storage instruction based on the user's trigger operation on the save image control, and then store the currently obtained image into the image library.
[0306] In this application, the process of obtaining the second image by performing perspective distortion correction on the first image can be referred to the description of step 303, with the difference being that: in addition to performing image transformations for the target person's face shifting backward, this application can also perform image transformations for other transformations of the target person's face, such as head shaking, changing facial expressions, changing actions, etc. Based on this, besides performing distortion correction on the first image, arbitrary transformations can also be performed on the first image. Even if the first image does not exhibit distortion, the target person's face can be transformed according to the transformation parameters entered by the user in the distortion correction function menu, making the target person's face in the transformed image closer to the appearance under the user-selected transformation parameters.
[0307] In one possible implementation, in step 1001, in addition to acquiring the first image containing the face of the target person, annotation data of the first image can also be acquired, such as image segmentation mask, bounding box positions, feature point positions, etc. In step 1003, after obtaining the transformed second image, the annotation data can be transformed according to the same transformation method as in step 1003 to obtain a new annotation file.
[0308] This approach can augment the image library with labeled data, generating various transformed images from a single image. The transformed images do not require manual re-labeling, and their effects are natural and realistic, with significant differences from the original image, which is beneficial for deep learning training.
[0309] In one possible implementation, step 1003 specifies the transformation method of the first 3D model, for example, specifying that the first 3D model is shifted to the left or right by about 6cm (the distance between the left and right eyes of a human face). After obtaining the transformed second image, the first image and the second image are used as inputs for the left and right eyes of the VR device, respectively, to achieve a 3D display effect of the target subject.
[0310] This allows ordinary 2D images or videos to be converted into VR input sources, achieving a 3D display effect of the target person's face. Because the depth of the 3D model itself is taken into account, the entire target object can achieve a stereoscopic effect.
[0311] This application achieves 3D image transformation effects using a 3D model. It corrects perspective distortion in close-up images of a target person's face from a photo library, making the relative proportions and positions of the facial features more closely resemble those of the actual target person, thus altering the imaging effect. Furthermore, by applying transformation parameters input by the user, diverse facial transformations can be achieved, enabling virtual image transformation.
[0312] It should be noted that the above embodiments establish a 3D model of the face of a target person in the image, thereby achieving distortion correction of the face. The image transformation method provided in this application can also be applied to the distortion correction of any other target object, the difference being that the object of establishing the 3D model changes from the face of the target person to the target object. Therefore, this application does not specifically limit the object of image transformation.
[0313] The image transformation method of this application is described below with two specific embodiments.
[0314] Example 1 Figures 12a-12f The example illustrates the process of distortion correction performed by the terminal in a selfie scenario.
[0315] like Figure 12a As shown, the user clicks the camera icon on the terminal's desktop to open the camera.
[0316] like Figure 12b As shown, the camera is set to the rear camera by default, and the image of the target scene captured by the rear camera is displayed on the terminal screen. Users can switch to the front camera by clicking the camera switch control in the photo function menu.
[0317] like Figure 12c As shown, the terminal's screen displays the first image of the target scene captured by the front-facing camera, which includes the user's face.
[0318] When the distance between the face of the target person in the first image and the front-facing camera is less than a preset threshold...
[0319] In one scenario, the terminal employs the above-mentioned... Figure 3 The image transformation method in the illustrated embodiment corrects the distortion of the first image to obtain a second image. Since no shutter is triggered, the second image is displayed as a preview image on the terminal screen. Figure 12d As shown, the terminal displays the words "distortion correction" on the screen and a close control. If the user does not need to display the corrected preview image, they can click the close control. After receiving the corresponding instruction, the terminal will display the first image as the preview image on the screen.
[0320] In another case, such as Figure 12e As shown, when the user clicks the shutter control on the photo function menu, the terminal uses the above-mentioned... Figure 3 The method in the illustrated embodiment corrects the distortion of the photographed image (first image) to obtain a second image, which is then saved into the image library.
[0321] In the third case, such as Figure 12f As shown, the terminal adopts the above-mentioned Figure 3In the method shown in the embodiment, the distortion of the first image is corrected to obtain the second image. Since the shutter is not triggered, the terminal displays both the first image and the second image as preview images on the terminal screen, and the user can intuitively see the difference between the first image and the second image.
[0322] Example 2 Figures 13a-13h An example is shown of the process of distortion correction for images in a picture library.
[0323] like Figure 13a As shown, the user clicks the photo library icon on the terminal's desktop to open the photo library.
[0324] like Figure 13b As shown, the user selects the first image in the image library that needs distortion correction. When this first image was captured, the distance between the target person's face and the camera was less than a preset threshold. This distance information can be obtained through the above... Figure 10 The distance acquisition method shown in the embodiment is obtained and will not be described again here.
[0325] like Figure 13c As shown, when the terminal detects distortion in the face of a target person in the currently displayed image, a pop-up window appears on the screen. This window displays the message "The image is distorted. Do you want to perform distortion correction?" and below this message are "Yes" and "No" controls. When the user clicks "Yes," the distortion correction function menu is displayed. It should be noted that... Figure 13c An example of a pop-up window that allows users to choose whether to perform distortion correction is provided, but this does not limit the interface or content of the pop-up window. For example, the text content, font size, and font of the text displayed on the pop-up window, as well as the content of the two controls corresponding to "Yes" and "No", can all be implemented in other ways. This application does not impose specific limitations on the implementation method of the pop-up window.
[0326] When the terminal detects distortion in the face of a target person in the currently displayed image, it can also... Figure 12d The screen displays a control used to toggle the distortion correction function menu on or off. It should be noted that... Figure 12d An example of a trigger control for the distortion correction function menu is provided, but this does not limit the interface or content of the pop-up window. For example, the position of the control and the content on the control can be implemented in other ways. This application does not specifically limit the implementation method of the control.
[0327] The distortion correction function menu can be as follows: Figure 11 or Figure 13i As shown.
[0328] When the distortion correction function menu is initially displayed, the transformation parameter values for each option can be either default values or pre-calculated values. For example, the slider corresponding to distance can initially be set to a value of 0. Figure 11 As shown, or, the terminal obtains a distance adjustment amount based on the image transformation algorithm, and displays the slider at the position corresponding to the adjustment amount, such as... Figure 13d As shown.
[0329] like Figure 13e As shown, users can adjust the distance by dragging the corresponding slider. Figure 13f As shown, users can also select the desired emoticon (happy) by triggering a control. This process can be referenced above. Figure 10 The methods described in the illustrated embodiments will not be repeated here.
[0330] like Figure 13g As shown, when the user is satisfied with the transformation result, the image saving control can be triggered. When the terminal receives the image saving instruction, it will save the currently obtained image into the image library.
[0331] like Figure 13h As shown, before selecting transformation parameters, the user can trigger the start recording control. After receiving the start recording command, the terminal starts screen recording and records the process of acquiring the second image. During this process, the start recording control changes to a stop recording control. When the user triggers the stop recording control, the terminal stops screen recording.
[0332] Example 3, Figure 13i Other examples of the distortion correction function menu are shown as examples.
[0333] like Figure 13i As shown, the distortion correction function menu includes controls for selecting facial features and adjusting their positions. The user first selects the nose (the oval control in front of the nose is black), then clicks the control indicating enlargement. Each click adjusts the size of the nose according to the set step size. Figure 10 The method in the illustrated embodiment enlarges the nose of the target person.
[0334] It should be noted that this application Figure 11 , Figures 12a-12f , Figures 13a-13h , Figure 13i The embodiments shown are all examples, but they do not constitute a limitation on the terminal's desktop, distortion correction function menu, camera interface, etc., and this application does not make any specific limitations in these regard.
[0335] In conjunction with the foregoing embodiments, the present invention can also provide an implementation method for distortion correction. (Participate) Figure 14a The method may include the following steps:
[0336] Step 401: Obtain a first image, which includes the face of the target person; wherein the face of the target person in the first image is distorted.
[0337] Optionally, the method for acquiring the first image includes, but is not limited to, one of the following methods:
[0338] Method 1: Capture a frame of image of the target scene using a camera, and obtain the first image based on the frame of image.
[0339] Method 2: Capture multiple frames of images of the target scene using a camera, and synthesize a first image based on the multiple frames of images;
[0340] Method 3: Retrieve the first image from images stored locally or in the cloud. For example, an image selected by the user in the terminal's gallery that includes the face of the target person (i.e., the user), which is usually taken at close range.
[0341] It should be understood that the image acquisition in methods 1 and 2 above pertains to real-time photography, which can occur when the shutter button is clicked. At this time, one or more images are acquired and processed by an ISP (Internet Service Provider) algorithm, such as black level correction, single-frame / multi-frame denoising, demosaic, and white balance, resulting in a single frame, i.e., the first image. Before the user clicks the shutter button, a preview stream of the shooting scene is displayed in the viewfinder. For example, in a selfie scenario, the first image is acquired by the front-facing camera of the terminal. Optionally, the FOV (Field of View) of the front-facing camera is set to 70°–110°, preferably 90°. Optionally, the front-facing camera is a wide-angle lens, and the FOV can be greater than 90°, such as 100°, 115°, etc., which are not exhaustively listed or limited in this invention. A depth sensor, such as structured light or TOF, can also be installed next to the front-facing camera. The pixel count of the camera sensor can use mainstream industry pixel values, such as 10M, 12M, 13M, 20M, etc.
[0342] Optionally, to save power, distortion correction can be omitted for the preview image. After the shutter is pressed, step 401 is executed to acquire the image, and distortion correction is performed on the acquired image as in step 403. Thus, the preview image and the target image obtained by the user after clicking the shutter in step 403 may not be the same.
[0343] It should be understood that the camera can be either front-facing or rear-facing. The reasons for facial distortion can also be found in the relevant descriptions in the foregoing embodiments.
[0344] In some embodiments, the target person is a person located in the central or near-central area of the image in the preview image. For a target person of interest to the user, the person is usually facing or nearly facing the camera during shooting, thus the target person is typically located in the central area of the preview image. In other embodiments, the number of pixels or screen ratio of the target person is greater than a preset threshold. Target people of interest to the user are usually facing the camera and close to it, especially in the case of a front-facing camera selfie, thus the target person is close to the central area in the preview image and has an area greater than the preset threshold. In other embodiments, the target person is the person with the smallest depth in the central area of the image in the preview image. When the person in the central area of the image in the preview image includes multiple individuals with different depths, the target person is the one with the smallest depth. In some embodiments, the terminal defaults to including only one target person. It is understood that the shooting terminal can determine the target person in various ways, and this application embodiment does not specifically limit the method. In the preview stream of the photo, the terminal can select the target person or face frame using a circle, rectangle, or person frame or face frame.
[0345] Step 402: Obtain the target distance; the target distance is used to characterize the distance between the face of the target person and the shooting terminal when the first image is captured. It can also be understood as the shooting distance between the shooting terminal and the target person when the first image is captured.
[0346] Optionally, the distance between the target person's face and the shooting terminal can be, but is not limited to, the distance between the foreground, center, eyebrows, eyes, nose, mouth, or ears of the target person's face and the shooting terminal. The shooting terminal can be the current terminal, or another terminal used to capture the first image, or the current terminal.
[0347] Specifically, the methods for obtaining the target distance include, but are not limited to, the following:
[0348] Method 1: Obtain the screen-to-body ratio of the target person's face in the first image; calculate the target distance based on the screen-to-body ratio and the field of view (FOV) of the first image; where the FOV of the first image is related to the camera's FOV and zoom level. This method can be applied to images captured in real-time or retrieved from historical images.
[0349] Method 2: The target distance is obtained through the EXIF information of the first image's exchangeable image file format; especially for images not captured in real time, such as historical images, the target distance can be determined based on the shooting distance information or related information recorded in the EXIF information.
[0350] Method 3: Obtain the target distance using a distance sensor, including a Time-of-Flight (TOF) sensor, a structured light sensor, or a binocular sensor. This method is suitable for real-time captured images. In this case, the current capturing terminal needs to have a distance sensor.
[0351] The above methods have been explained in the foregoing embodiments and will not be repeated here.
[0352] Step 403: Perform a first processing on the first image to obtain a second image; wherein, the first processing includes distortion correction of the first image based on the target distance; the face of the target person in the second image is closer to the true appearance of the target person's face than the face of the target person in the first image.
[0353] The first image is processed to obtain the second image. The present invention provides, but is not limited to, the following two processing approaches.
[0354] Approach 1: Obtain the correction distance; perform distortion correction on the first image based on the target distance and the correction distance. It should be understood that, for the same target distance, a larger correction distance indicates a greater degree of distortion correction. The correction distance needs to be greater than the target distance.
[0355] This approach is primarily used for distortion correction of the entire human face.
[0356] Optionally, the methods for obtaining the correction distance include, but are not limited to, the following two optional methods:
[0357] Method 1: Obtain the correction distance corresponding to the target distance based on a preset correspondence between shooting distance and correction distance. For example, the correspondence can be a 1:N ratio between the target distance and the correction distance; N can be a value greater than 1, including but not limited to 2, 2.5, 3, 4, 5, etc. Optionally, the correspondence can also be a pre-set correspondence table.
[0358] The following two tables serve as examples:
[0359] Table 1
[0360]
[0361]
[0362] Table 2
[0363] Shooting distance Correction distance …… …… 15cm 15cm 20cm 30cm 30cm 50cm 40cm 55cm 50cm 60cm 60cm 65cm 70cm 70cm …… ……
[0364] Optionally, Table 1 indicates that the correction intensity is strongest when the shooting distance is 25cm-35cm, and no correction can be performed when the shooting distance is greater than 50cm. Table 2 indicates that the correction intensity is relatively strong when the shooting distance is 20-40cm, and no correction can be performed when the shooting distance is greater than 70cm. The ratio of correction distance to shooting distance expresses the correction intensity at that shooting distance; the larger the ratio, the stronger the correction intensity. Furthermore, considering the overall distribution of shooting distances, the corresponding correction intensity will vary for different shooting distances.
[0365] Method 2: Display the distortion correction function menu; accept control adjustment commands input by the user based on the distortion correction function menu; the control adjustment commands are used to determine the correction distance. The control adjustment commands can explicitly determine the correction distance; or they can be based on the increment of the target distance, with the correction distance further calculated by the terminal. Specifically, display the distortion correction control, which is used to open the distortion correction function menu; respond to the user's enabling of the distortion correction control, display the distortion correction function menu; the distortion correction function menu may include controls for determining the correction distance and / or controls for selecting regions of interest, and / or controls in the options for adjusting facial expressions. Specifically, the form and use of the menu can be referred to in the aforementioned embodiments. Figure 11 , 13a -13h related description. Optionally, some distortion correction function menus can also be found in the appendix. Figure 13i-13n It should be understood that in actual implementation, the distortion correction function menu may include more or fewer content or elements, as shown in the example image, and this invention does not limit this. Optionally, the distortion correction menu may be displayed or rendered in the first image, or it may occupy a portion of the screen space along with the first image; this invention does not limit this.
[0366] Based on idea 1, methods for distortion correction of the first image may include:
[0367] A first 3D model is obtained by fitting the face of the target person in the first image to a standard face model based on the target distance. The first 3D model is the 3D model corresponding to the face of the target person, representing the shape, position, and posture of the target person's real face in 3D space. A second 3D model is obtained by performing coordinate transformation on the first 3D model based on the correction distance. A first set of projection points is obtained by performing perspective projection on the first 3D model based on the coordinate system of the first image. A second set of projection points is obtained by performing perspective projection on the second 3D model based on the coordinate system of the first image. A displacement vector (field) is obtained by aligning the first set of projection points and the second set of projection points. The first image is transformed based on the displacement vector (field) to obtain the second image.
[0368] The second three-dimensional model is obtained by performing coordinate transformation on the first three-dimensional model based on the correction distance;
[0369] One possible implementation is as follows: for a given shooting distance, there is a corresponding correction distance, and the ratio of the correction distance to the shooting distance is k; if the spatial coordinates of any pixel on the first 3D model are (x, y, z), then the coordinates of that point mapped onto the second 3D model are (kx, ky, kz). The value of k can be determined through the correspondence between the shooting distance and the correction distance in Method 1 above; it can also be determined through the adjustment of the distortion correction control in Method 2 above. Optionally, refer to... Figure 13k Users can control the distance changes in the preview interface. The slider in the 13k can be used to increase or decrease the distance relative to the current shooting distance; for example, moving it to the right increases the distance, and moving it to the left decreases it. The terminal determines the correction distance based on the shooting distance of the current image and the user's actions on the slider control. The midpoint or the left starting point of the progress bar indicates that the correction distance is the same as the shooting distance of the first image. In a simpler example, such as... Figure 13l As shown, the interface includes a more direct control for adjusting the distortion correction intensity, where k is the ratio of the correction distance to the shooting distance. For example, from left to right, the k value gradually increases, and the distortion correction intensity gradually becomes stronger; the midpoint of the progress bar can be k=1, or the starting point on the left side of the progress bar can be k=1. It should be understood that for cases where k is less than 1, interesting distortion magnification can be achieved.
[0370] Another possible implementation is as follows: For a given shooting distance, there is a corresponding correction distance. These shooting and correction distances are three-dimensional vectors, which can be represented as (dx1, dy1, dz1) and (dx2, dy2, dz2), respectively. If the spatial coordinates of any pixel on the first three-dimensional model are (x, y, z), then the coordinates mapped to the second three-dimensional model are (k1*x, k2*y, k3*z), where k1 = dx2 / dx1, k2 = dy2 / dy1, and k3 = dz2 / dz1. The aforementioned three-dimensional vectors and spatial coordinates correspond to the same reference coordinate system, and the coordinate axes are also corresponding. The values of k1, k2, and k3 can be determined through the correspondence between shooting distance and correction distance in Method 1 above. In this case, k1, k2, and k3 can be equal. The values of k1, k2, and k3 can also be determined by adjusting the distortion correction control in Method 2 above. Optionally, refer to... Figure 13m Users can control the distance changes in the preview interface. The slider in the 13m setting can be used to increase or decrease the distance relative to the current shooting distance in three dimensions (x, y, z) of a preset coordinate system. For example, moving right increases the distance, and moving left decreases it. The terminal can determine the correction distance based on the shooting distance of the current image and the user's operation on the slider control. The midpoint or the left starting point of the progress bar indicates that the correction distance is the same as the shooting distance of the first image. In a simpler example, such as... Figure 13n As shown, the interface includes more direct controls for adjusting the distortion correction intensity. The values of k1, k2, and k3 are the ratios of the correction distance to the shooting distance in the three dimensions. For example, for the same progress bar, from left to right, the values of k1, k2, or k3 gradually increase, indicating a gradually stronger distortion correction intensity in that dimension. The midpoint of the progress bar can be k1, k2, or k3 = 1, or the starting point on the left side of the progress bar can be k1, k2, or k3 = 1. It should be understood that for cases where k1, k2, or k3 is less than 1, interesting distortion amplification in a specific dimension can be achieved.
[0371] Optionally, the coordinate system based on the first image can be understood as a coordinate system established with the center point of the first image as the origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis. Assuming the width and height of the image are W and H, the coordinates of the upper right corner can be (W / 2, H / 2).
[0372] Specifically, some of the related operations can be referred to the relevant descriptions in the foregoing embodiments.
[0373] Approach 2: Obtain the correction distance; obtain the region of interest (ROI); the ROI includes at least one of the following: eyebrows, eyes, nose, mouth, or ears; perform distortion correction on the first image based on the target distance, the ROI, and the correction distance. It should be understood that for the same target distance, a larger correction distance indicates a greater degree of distortion correction. The correction distance needs to be greater than the target distance.
[0374] This approach is primarily used for distortion correction of regions of interest in a human face. The method for obtaining the correction distance can refer to methods 1 and 2 in approach 1.
[0375] At this point, the method for distortion correction of the first image may include:
[0376] A first 3D model is obtained by fitting the face of the target person in the first image to a standard face model based on the target distance; the first 3D model is the 3D model corresponding to the face of the target person; a second 3D model is obtained by performing coordinate transformation on the first 3D model based on the correction distance; a first set of projection points is obtained by performing perspective projection on the model region corresponding to the region of interest in the first 3D model; a second set of projection points is obtained by performing perspective projection on the model region corresponding to the region of interest in the second 3D model; a displacement vector (field) is obtained by aligning the first set of projection points and the second set of projection points; the first image is transformed based on the displacement vector (field) to obtain the second image.
[0377] The specific method for transforming the first 3D model into a second 3D model based on the correction distance can be found in the relevant description in Idea 1, and will not be repeated here. It is worth noting that the distortion correction menu includes a region of interest option; see [link to relevant documentation] for details. Figure 13j-13m The app offers facial feature options; it can respond to user selections and correct distortions in areas of interest, such as eyebrows, eyes, nose, mouth, ears, or facial contours.
[0378] Regarding the two approaches mentioned above, when the target distance is within a first preset distance range, the smaller the target distance (i.e., the closer the actual shooting distance), the greater the required distortion correction intensity. This necessitates determining a relatively higher correction distance for algorithmic processing to achieve a better correction effect. The first preset distance range is less than or equal to a preset distortion shooting distance. For example, the first preset distance range could include [30cm, 40cm] or [25cm, 45cm]. Furthermore, the greater the angle at which the target person's face deviates from a frontal view, the less required distortion correction intensity. This means a relatively smaller correction distance needs to be determined for algorithmic processing to achieve a better correction effect.
[0379] Specifically, some of the related operations can be referred to the relevant descriptions in the foregoing embodiments.
[0380] In practice, not all images require distortion correction. Therefore, after acquiring the first image, distortion correction can be triggered by determining certain conditions. These trigger conditions include, but are not limited to, conditions where one or more of the following conditions are met:
[0381] Condition 1: In the first image, the completeness of the target person's face is greater than a preset completeness threshold; or, the facial features of the target person are determined to be complete. This allows the terminal to recognize facial scenes that can undergo distortion correction.
[0382] Condition 2: The pitch angle of the target person's face deviating from a frontal view conforms to the first angle range, and the yaw angle of the target person's face deviating from a frontal view conforms to the second angle range; for example, the first angle range is within [-30°, 30°], and the second angle range is within [-30°, 30°]. That is, the facial image undergoing distortion correction needs to be as close to a frontal view as possible.
[0383] Condition 3: Face detection is performed on the current shooting scene, and only one valid face is detected; the valid face is the face of the target person. That is, distortion correction can be applied to only one subject.
[0384] Condition 4: Determine if the target distance is less than or equal to the preset distortion shooting distance. Shooting distance is usually a crucial factor in causing distortion of people. For example, the distortion shooting distance should not exceed 60cm, such as 60cm or 50cm.
[0385] Condition 5: Determine that the currently enabled camera on the terminal is the front-facing camera, which is used to capture the first image. Front-facing shooting is also a significant cause of human distortion.
[0386] Condition 6: The terminal's current shooting mode is the preset shooting mode, or distortion correction is enabled. This includes portrait shooting, default shooting, or other specific shooting modes; or the terminal has distortion correction enabled.
[0387] This invention provides an image processing method for distortion correction. It is not only applicable to real-time shooting but also helps users perform quick, convenient, and fun distortion correction and editing while viewing an image library, providing a superior user experience.
[0388] Accordingly, this application also provides an image processing apparatus, see below. Figure 14b The device 1400 includes:
[0389] The acquisition module 1401 is used to acquire a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; and is also used to acquire a target distance; the target distance is used to characterize the distance between the face of the target person and the shooting terminal when the first image is captured.
[0390] In the specific implementation process, the acquisition module 1401 is specifically used to execute the methods mentioned in steps 401 and 402, as well as the methods that can be equivalently replaced.
[0391] The processing module 1402 is used to perform a first processing on the face of the target person in the first image to obtain a second image; wherein, the first processing includes distortion correction of the first image according to the target distance; the face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image.
[0392] In the specific implementation process, the processing module 1402 is specifically used to execute the method mentioned in step 403 and the equivalent replacement method.
[0393] Furthermore, the device 1400 may also include:
[0394] The judgment module 1403 is used to determine whether one or more of the conditions 1-6 mentioned above are true.
[0395] Display module 1404 is used to display the various interfaces and images described above.
[0396] The recording module 1405 is used to record segments of the distortion correction process as described above, or part or all of the process of the user using the camera or editing images, in order to generate interesting videos or GIFs.
[0397] The specific method embodiments described above, as well as the explanations, descriptions, and extensions of the technical features in the embodiments, are also applicable to the execution of the method in the device, and will not be elaborated upon in the device embodiments.
[0398] The present invention has numerous implementation methods and application scenarios, which cannot be listed one by one in detail. Without violating the laws of nature, the optional implementation methods in the present invention can be freely combined and transformed, and these should all fall within the protection scope of the present invention.
[0399] Figure 15 This is a schematic diagram of the structure of an embodiment of the image transformation device of this application, as shown below. Figure 15 As shown, the device in this embodiment can be applied to Figure 2 The terminal shown. The image transformation module includes: an acquisition module 1501, a processing module 1502, a recording module 1503, and a display module 1504. Among them,
[0400] In a selfie scenario, the acquisition module 1501 is used to acquire a first image of a target scene using a front-facing camera, the target scene including the face of a target person; and to acquire a target distance between the face of the target person and the front-facing camera. The processing module 1502 is used to perform a first processing on the first image to obtain a second image when the target distance is less than a preset threshold; the first processing includes distortion correction of the first image based on the target distance; wherein the face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image.
[0401] In one possible implementation, the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a designated part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera.
[0402] In one possible implementation, the acquisition module 1501 is specifically used to acquire the screen ratio of the target person's face in the first image; and to obtain the target distance based on the screen ratio and the field of view (FOV) of the front-facing camera.
[0403] In one possible implementation, the acquisition module 1501 is specifically used to acquire the target distance through a distance sensor, which includes a time-of-flight (TOF) sensor, a structured light sensor, or a binocular sensor.
[0404] In one possible implementation, the preset threshold is less than 80 centimeters.
[0405] In one possible implementation, the second image includes a preview image or an image obtained after the shutter is triggered.
[0406] In one possible implementation, the face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
[0407] In one possible implementation, the processing module 1502 is specifically used to fit the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; and to perform perspective distortion correction on the first image based on the depth information to obtain the second image.
[0408] In one possible implementation, the processing module 1502 is specifically used to establish a first three-dimensional model of the target person's face; transform the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the target person's face; obtain the pixel displacement vector field of the target person's face based on the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image based on the pixel displacement vector field of the target person's face.
[0409] In one possible implementation, the processing module 1502 is specifically configured to perform perspective projection on the first 3D model based on the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to multiple pixels in the first 3D model; perform perspective projection on the second 3D model based on the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to multiple pixels in the second 3D model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to the first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the multiple identical pixels contained in the first 3D model and the second 3D model.
[0410] In a scenario where distortion correction of an image is performed based on user-inputted transformation parameters, an acquisition module 1501 is used to acquire a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; a display module 1504 is used to display a distortion correction function menu; the acquisition module 1501 is also used to acquire transformation parameters input by the user on the distortion correction function menu, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the face of the target person and the camera when the shooting terminal is shooting the face of the target person; a processing module 1502 is used to perform a first processing on the first image to obtain a second image; the first processing includes distortion correction of the first image according to the transformation parameters; wherein the face of the target person in the second image is closer to the true appearance of the target person's face than the face of the target person in the first image.
[0411] In one possible implementation, the distortion of the target person's face in the first image is caused by the second terminal capturing the first image at a target distance less than a first preset threshold between the target person's face and the second terminal; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera.
[0412] In one possible implementation, the target distance is obtained by the screen-to-body ratio of the target person's face in the first image and the field of view (FOV) of the second terminal's camera; or, the target distance is obtained by the equivalent focal length in the EXIF information of the first image's exchangeable image file format.
[0413] In one possible implementation, the distortion correction function menu includes an option to adjust the equivalent simulated shooting distance; the acquisition module 1501 is specifically used to acquire the equivalent simulated shooting distance based on the instruction triggered by the user's operation of the control or slider in the option to adjust the equivalent simulated shooting distance.
[0414] In one possible implementation, when the distortion correction function menu is initially displayed, the value of the equivalent analog shooting distance in the option to adjust the equivalent analog shooting distance includes a default value or a pre-calculated value.
[0415] In one possible implementation, the display module 1504 is further configured to display a pop-up window when the face of the target person is distorted, the pop-up window being used to provide a selection control for whether to perform distortion correction; and responding to the instruction generated by the user operation when the user clicks the distortion correction control on the pop-up window.
[0416] In one possible implementation, the display module 1504 is further configured to display a distortion correction control when the face of the target person is distorted, the distortion correction control being used to open the distortion correction function menu; and to respond to the instructions generated by the user's operation when the user clicks the distortion correction control.
[0417] In one possible implementation, the distortion of the target person's face in the first image is caused by the second terminal capturing the first image when the camera's field of view (FOV) is greater than a second preset threshold, and the pixel distance between the target person's face and the edge of the FOV is less than a third preset threshold; wherein, the pixel distance includes the number of pixels between the foremost part of the target person's face and the edge of the FOV; or, the number of pixels between a specified part of the target person's face and the edge of the FOV; or, the number of pixels between the center position of the target person's face and the edge of the FOV.
[0418] In one possible implementation, the FOV is obtained from the EXIF information of the first image.
[0419] In one possible implementation, the second preset threshold is 90°, and the third preset threshold is one-quarter of the length or width of the first image.
[0420] In one possible implementation, the distortion correction function menu includes an option to adjust the displacement distance; the acquisition module 1501 is further configured to acquire the adjustment direction and displacement distance based on the instructions triggered by the user's operation of the control or slider in the option to adjust the displacement distance.
[0421] In one possible implementation, the distortion correction function menu includes options for adjusting the relative position and / or relative proportion of facial features; the acquisition module 1501 is further configured to acquire the adjustment direction, displacement distance, and / or facial feature size based on instructions triggered by the user's operation of controls or sliders in the options for adjusting the relative position and / or relative proportion of facial features.
[0422] In one possible implementation, the distortion correction function menu includes an option to adjust the angle; the acquisition module 1501 is further configured to acquire the adjustment direction and adjustment angle based on instructions triggered by the user's operation of the controls or sliders in the angle adjustment option; or, the distortion correction function menu includes an option to adjust facial expressions; the acquisition module 1501 is further configured to acquire a new facial expression template based on instructions triggered by the user's operation of the controls or sliders in the facial expression adjustment option; or, the distortion correction function menu includes an option to adjust actions; the acquisition module 1501 is further configured to acquire a new action template based on instructions triggered by the user's operation of the controls or sliders in the action adjustment option.
[0423] In one possible implementation, the face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
[0424] In one possible implementation, the processing module 1502 is specifically used to fit the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the target person's face; and to perform perspective distortion correction on the first image based on the depth information and the transformation parameters to obtain the second image.
[0425] In one possible implementation, the processing module 1502 is specifically used to establish a first three-dimensional model of the target person's face; transform the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the target person's face; obtain the pixel displacement vector field of the target person's face according to the depth information, the first three-dimensional model and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the target person's face.
[0426] In one possible implementation, the processing module 1502 is specifically configured to perform perspective projection on the first 3D model based on the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to multiple pixels in the first 3D model; perform perspective projection on the second 3D model based on the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to multiple pixels in the second 3D model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to the first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the multiple identical pixels contained in the first 3D model and the second 3D model.
[0427] In one possible implementation, the system further includes: a recording module 1503; the acquisition module is further configured to acquire a recording instruction based on a user's trigger operation on the recording control; the recording module 1503 is configured to start recording the acquisition process of the second image based on the recording instruction until a stop recording instruction is received from a user's trigger operation on the stop recording control.
[0428] In one possible implementation, the display module 1504 is configured to display a distortion correction function menu on the screen, the distortion correction function menu including one or more sliders and / or one or more controls; receive a distortion correction instruction, the distortion correction instruction including transformation parameters generated when the user performs touch operation on the one or more sliders and / or the one or more controls, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the target person's face and the camera when the shooting terminal is shooting the target person's face; the processing module 1502 is configured to perform a first processing on the first image according to the transformation parameters to obtain a second image; the first processing includes distortion correction on the first image; wherein, the target person's face in the second image is closer to the real appearance of the target person's face than the target person's face in the first image.
[0429] The apparatus of this embodiment can be used to perform Figure 3 or Figure 10 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.
[0430] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or implemented by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0431] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0432] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0433] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0434] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0435] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0436] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0437] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0438] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An image processing method, characterized in that, The method is applied to a first terminal, and the method includes: A first image is acquired, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; the distortion of the face of the target person in the first image is caused by the target distance between the second terminal and the target person's face being less than a first preset threshold when the second terminal captures the first image; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera; Display the distortion correction function menu; Obtain the transformation parameters input by the user on the distortion correction function menu. The transformation parameters include at least the equivalent simulated shooting distance, which is used to simulate the distance between the target person's face and the camera when the shooting terminal is shooting the target person's face. The first image is processed to obtain a second image; the first processing includes distortion correction of the first image according to the transformation parameters, including: fitting the face of the target person in the first image to a standard face model according to the target distance to obtain the depth information of the target person's face; and performing perspective distortion correction of the first image according to the depth information and the transformation parameters to obtain the second image; wherein the face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image.
2. The method according to claim 1, characterized in that, The target distance is obtained by the screen ratio of the target person's face in the first image and the FOV of the second terminal's camera; or, the target distance is obtained by the equivalent focal length in the EXIF information of the first image's exchangeable image file format.
3. The method according to claim 1 or 2, characterized in that, The distortion correction function menu includes an option to adjust the equivalent simulated shooting distance; The process of obtaining the transformation parameters input by the user on the distortion correction function menu includes: The equivalent simulated shooting distance is obtained based on the instruction triggered by the user's operation of the control or slider in the option to adjust the equivalent simulated shooting distance.
4. The method according to claim 3, characterized in that, When the distortion correction function menu is initially displayed, the value of the equivalent analog shooting distance in the option to adjust the equivalent analog shooting distance includes a default value or a pre-calculated value.
5. The method according to claim 1 or 2, characterized in that, Before displaying the distortion correction function menu, it also includes: When the face of the target person is distorted, a pop-up window is displayed, which provides a selection control for whether to perform distortion correction. When the user clicks the distortion correction control on the pop-up window, it responds to the instructions generated by the user's operation.
6. The method according to claim 1 or 2, characterized in that, Before displaying the distortion correction function menu, it also includes: When the face of the target person is distorted, a distortion correction control is displayed, which is used to open the distortion correction function menu; When the user clicks the distortion correction control, it responds to the instructions generated by the user's operation.
7. The method according to claim 1 or 2, characterized in that, The distortion of the target person's face in the first image is caused by the second terminal capturing the first image when the camera's field of view (FOV) is greater than a second preset threshold, and the pixel distance between the target person's face and the edge of the FOV is less than a third preset threshold; wherein, the pixel distance includes the number of pixels between the foremost part of the target person's face and the edge of the FOV; or, the number of pixels between a specified part of the target person's face and the edge of the FOV; or, the number of pixels between the center position of the target person's face and the edge of the FOV.
8. The method according to claim 7, characterized in that, The FOV is obtained from the EXIF information of the first image.
9. The method according to claim 7, characterized in that, The second preset threshold is 90°, and the third preset threshold is one-quarter of the length or width of the first image.
10. The method according to claim 7, characterized in that, The distortion correction function menu includes an option to adjust the displacement distance; The process of obtaining the transformation parameters input by the user on the distortion correction function menu includes: The adjustment direction and displacement distance are obtained based on the instructions triggered by the user's operation of the controls or sliders in the option to adjust the displacement distance.
11. The method according to claim 1 or 2, characterized in that, The distortion correction function menu includes options for adjusting the relative position and / or relative proportion of facial features; The process of obtaining the transformation parameters input by the user on the distortion correction function menu includes: The adjustment direction, displacement distance, and / or facial feature size are obtained based on the instructions triggered by the user's operation of the controls or sliders in the options for adjusting the relative position and / or relative proportion of facial features.
12. The method according to claim 1 or 2, characterized in that, The distortion correction function menu includes an option to adjust the angle; obtaining the transformation parameters input by the user on the distortion correction function menu includes: The adjustment direction and angle are obtained based on the user's input of the controls or sliders in the angle adjustment options; or, The distortion correction function menu includes options for adjusting facial expressions; obtaining the transformation parameters input by the user on the distortion correction function menu includes: A new emoji template is retrieved based on the user's actions using the controls or sliders in the emoji adjustment options; or, The distortion correction function menu includes options for adjusting actions; obtaining the transformation parameters input by the user on the distortion correction function menu includes: A new action template is obtained based on the instructions triggered by the user's operation of the controls or sliders in the options for adjusting the action.
13. The method according to claim 1 or 2, characterized in that, The face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: The relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the target person's facial features than those in the first image; and / or, The relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person in the first image than the relative positions of the facial features of the target person in the first image.
14. The method according to claim 1, characterized in that, The step of correcting perspective distortion of the first image based on the depth information and the transformation parameters to obtain the second image includes: Establish a first 3D model of the target person's face; The pose and / or shape of the first three-dimensional model are transformed according to the transformation parameters to obtain a second three-dimensional model of the target person's face; The pixel displacement vector field of the target person's face is obtained based on the depth information, the first three-dimensional model, and the second three-dimensional model; The second image is obtained based on the pixel displacement vector field of the target person's face.
15. The method according to claim 14, characterized in that, The step of obtaining the pixel displacement vector field of the target person's face based on the depth information, the first 3D model, and the second 3D model includes: A first coordinate set is obtained by performing perspective projection on the first three-dimensional model based on the depth information. The first coordinate set includes coordinate values corresponding to multiple pixels in the first three-dimensional model. A second coordinate set is obtained by performing perspective projection on the second three-dimensional model based on the depth information. The second coordinate set includes coordinate values corresponding to multiple pixels in the second three-dimensional model. The pixel displacement vector field of the target object is obtained by calculating the coordinate difference between the first coordinate value and the second coordinate value. The first coordinate value includes the coordinate value corresponding to the first pixel in the first coordinate set, and the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set. The first pixel includes any one of the multiple identical pixels contained in the first three-dimensional model and the second three-dimensional model.
16. An image transformation method, characterized in that, The method is applied to a first terminal, and the method includes: A first image is acquired, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; the distortion of the face of the target person in the first image is caused by the target distance between the second terminal and the target person's face being less than a first preset threshold when the second terminal captures the first image; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera; A distortion correction function menu is displayed on the screen, the distortion correction function menu including one or more sliders and / or one or more controls; The system receives a distortion correction instruction, which includes transformation parameters generated when the user performs a touch operation on the one or more sliders and / or the one or more controls. The transformation parameters include at least an equivalent simulated shooting distance, which is used to simulate the distance between the target person's face and the camera when the shooting terminal is shooting the target person's face. The first image is processed according to the transformation parameters to obtain a second image; the first processing includes distortion correction of the first image, including: fitting the face of the target person in the first image to a standard face model according to the target distance to obtain the depth information of the target person's face; and performing perspective distortion correction on the first image according to the depth information and the transformation parameters to obtain the second image; wherein the face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image.
17. An image transformation device, characterized in that, include: An acquisition module is used to acquire a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; the distortion of the face of the target person in the first image is caused by the target distance between the second terminal and the second terminal being less than a first preset threshold when the first image is captured; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera; The display module is used to display the distortion correction function menu; The acquisition module is also used to acquire the transformation parameters input by the user on the distortion correction function menu. The transformation parameters include at least the equivalent simulated shooting distance, which is used to simulate the distance between the target person's face and the camera when the shooting terminal is shooting the target person's face. A processing module is configured to perform a first processing on the first image to obtain a second image; the first processing includes distortion correction of the first image according to the transformation parameters, including: fitting the face of the target person in the first image to a standard face model according to the target distance to obtain the depth information of the target person's face; and performing perspective distortion correction on the first image according to the depth information and the transformation parameters to obtain the second image; wherein the face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image.
18. The apparatus according to claim 17, characterized in that, The target distance is obtained by the screen ratio of the target person's face in the first image and the FOV of the second terminal's camera; or, the target distance is obtained by the equivalent focal length in the EXIF information of the first image's exchangeable image file format.
19. The apparatus according to claim 17 or 18, characterized in that, The distortion correction function menu includes an option to adjust the equivalent simulated shooting distance; the acquisition module is specifically used to acquire the equivalent simulated shooting distance based on the instructions triggered by the user's operation of the controls or sliders in the option to adjust the equivalent simulated shooting distance.
20. The apparatus according to claim 19, characterized in that, When the distortion correction function menu is initially displayed, the value of the equivalent analog shooting distance in the option to adjust the equivalent analog shooting distance includes a default value or a pre-calculated value.
21. The apparatus according to claim 17 or 18, characterized in that, The display module is also used to display a pop-up window when the face of the target person is distorted, the pop-up window being used to provide a selection control for whether to perform distortion correction; when the user clicks the control to perform distortion correction on the pop-up window, the module responds to the command generated by the user's operation.
22. The apparatus according to claim 17 or 18, characterized in that, The display module is also used to display a distortion correction control when the face of the target person is distorted. The distortion correction control is used to open the distortion correction function menu. When the user clicks the distortion correction control, it responds to the command generated by the user operation.
23. The apparatus according to claim 17 or 18, characterized in that, The distortion of the target person's face in the first image is caused by the second terminal capturing the first image when the camera's field of view (FOV) is greater than a second preset threshold, and the pixel distance between the target person's face and the edge of the FOV is less than a third preset threshold; wherein, the pixel distance includes the number of pixels between the foremost part of the target person's face and the edge of the FOV; or, the number of pixels between a specified part of the target person's face and the edge of the FOV; or, the number of pixels between the center position of the target person's face and the edge of the FOV.
24. The apparatus according to claim 23, characterized in that, The FOV is obtained from the EXIF information of the first image.
25. The apparatus according to claim 23, characterized in that, The second preset threshold is 90°, and the third preset threshold is one-quarter of the length or width of the first image.
26. The apparatus according to claim 23, characterized in that, The distortion correction function menu includes an option to adjust the displacement distance; the acquisition module is also used to acquire the adjustment direction and displacement distance based on the instructions triggered by the user's operation of the controls or sliders in the option to adjust the displacement distance.
27. The apparatus according to claim 17 or 18, characterized in that, The distortion correction function menu includes options for adjusting the relative position and / or relative proportion of facial features; the acquisition module is also used to acquire the adjustment direction, displacement distance and / or facial feature size based on the instructions triggered by the user's operation of the controls or sliders in the options for adjusting the relative position and / or relative proportion of facial features.
28. The apparatus according to claim 17 or 18, characterized in that, The distortion correction function menu includes an option to adjust the angle; the acquisition module is further configured to acquire the adjustment direction and adjustment angle based on the user's operation of the controls or slider in the angle adjustment option; or, the distortion correction function menu includes an option to adjust facial expressions; the acquisition module is further configured to acquire a new facial expression template based on the user's operation of the controls or slider in the facial expression adjustment option; or, the distortion correction function menu includes an option to adjust actions; the acquisition module is further configured to acquire a new action template based on the user's operation of the controls or slider in the action adjustment option.
29. The apparatus according to claim 17 or 18, characterized in that, The face of the target person in the second image is closer to the actual appearance of the target person's face than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.
30. The apparatus according to claim 17, characterized in that, The processing module is specifically used to establish a first three-dimensional model of the target person's face; transform the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the target person's face; obtain the pixel displacement vector field of the target person's face according to the depth information, the first three-dimensional model and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the target person's face.
31. The apparatus according to claim 30, characterized in that, The processing module is specifically used to perform perspective projection on the first three-dimensional model based on the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to multiple pixels in the first three-dimensional model; and to perform perspective projection on the second three-dimensional model based on the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to multiple pixels in the second three-dimensional model. The pixel displacement vector field of the target object is obtained by calculating the coordinate difference between the first coordinate value and the second coordinate value. The first coordinate value includes the coordinate value corresponding to the first pixel in the first coordinate set, and the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set. The first pixel includes any one of the multiple identical pixels contained in the first three-dimensional model and the second three-dimensional model.
32. An image transformation device, characterized in that, include: An acquisition module is used for a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; the distortion of the face of the target person in the first image is caused by the second terminal capturing the first image and the target distance between the face of the target person and the second terminal being less than a first preset threshold; wherein the target distance includes the distance between the foremost part of the face of the target person and the front-facing camera; or, the distance between a specified part of the face of the target person and the front-facing camera; or, the distance between the center of the face of the target person and the front-facing camera; The display module is used to display a distortion correction function menu on the screen, the distortion correction function menu including one or more sliders and / or one or more controls; and to receive distortion correction instructions, the distortion correction instructions including transformation parameters generated when the user performs touch operation on the one or more sliders and / or the one or more controls, the transformation parameters including at least an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the target person's face and the camera when the shooting terminal is shooting the target person's face; The processing module is configured to perform a first processing on the first image according to the transformation parameters to obtain a second image; the first processing includes distortion correction of the first image, including: fitting the face of the target person in the first image to a standard face model according to the target distance to obtain the depth information of the target person's face; and performing perspective distortion correction on the first image according to the depth information and the transformation parameters to obtain the second image; wherein the face of the target person in the second image is closer to the real appearance of the target person's face than the face of the target person in the first image.
33. A terminal device, characterized in that, The terminal device includes a memory and one or more processors; wherein... The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-16.
34. The terminal device according to claim 33, characterized in that, The terminal device further includes an antenna system, which, under the control of the processor, transmits and receives wireless communication signals to achieve wireless communication with a mobile communication network; the mobile communication network includes one or more of the following: GSM network, CDMA network, 3G network, 4G network, FDMA, TDMA, PDC, TACS, AMPS, WCDMA, TDSCDMA, WIFI, and LTE network.
35. A computer-readable storage medium, characterized in that, Includes a computer program, which, when executed on a computer, causes the computer to perform the method of any one of claims 1-16.
36. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer or processor, causes the computer or processor to perform the method of any one of claims 1-16.
37. An image processing method, characterized in that, The method includes: A first image is acquired, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; the distortion of the face of the target person in the first image is caused by the target distance between the second terminal and the target person's face being less than a first preset threshold when the second terminal captures the first image; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera; Obtain the target distance; the target distance is used to characterize the distance between the face of the target person and the shooting terminal when the first image is captured; The first image is processed to obtain a second image; wherein, the first processing includes distortion correction of the face of the target person in the first image based on the target distance, including: fitting the face of the target person in the first image to a standard face model based on the target distance to obtain the depth information of the face of the target person; and correcting the perspective distortion of the first image based on the depth information and transformation parameters to obtain the second image; the face of the target person in the second image is closer to the real appearance of the face of the target person than the face of the target person in the first image.
38. The method according to claim 37, characterized in that, Before performing the first processing on the first image to obtain the second image, the method further includes: It is determined that the completeness of the target person's face in the first image is greater than a preset completeness threshold; or, it is determined that the facial features of the target person are complete.
39. The method according to claim 37, characterized in that, Before performing the first processing on the first image to obtain the second image, the method further includes: It is determined that in the first image, the pitch angle of the target person's face deviating from the frontal face conforms to a first angle range, and the yaw angle of the target person's face deviating from the frontal face conforms to a second angle range.
40. The method according to claim 37, characterized in that, Before performing the first processing on the first image to obtain the second image, the method further includes: Face detection is performed on the current shooting scene, and only one valid face is detected; the valid face is the face of the target person.
41. The method according to claim 37, characterized in that, Before performing the first processing on the first image to obtain the second image, the method further includes: The target distance is determined to be less than or equal to a preset distortion shooting distance.
42. The method according to any one of claims 37-41, characterized in that, The acquisition of the first image includes: A first image is obtained by capturing a frame of an image of a target scene using a camera; or... A first image is synthesized by capturing multiple frames of images of a target scene using a camera; or... Retrieve the first image from images stored locally or in the cloud.
43. The method according to claim 42, characterized in that, The camera in question is a front-facing camera.
44. The method according to claim 42, characterized in that, Before the camera captures images of the target scene, the method further includes setting the shooting mode to a preset shooting mode.
45. The method according to any one of claims 37-41, characterized in that, When the target distance is within a first preset distance range, the smaller the target distance, the greater the intensity of the distortion correction; wherein, the first preset distance range is less than or equal to a preset distortion shooting distance.
46. The method according to any one of claims 37-41, characterized in that, The greater the angle at which the target person's face deviates from a frontal view, the weaker the intensity of the distortion correction.
47. The method according to any one of claims 37-41, characterized in that, The distance between the target person's face and the shooting terminal is defined as the distance between the foremost part, center position, eyebrows, eyes, nose, mouth, or ears of the target person's face and the shooting terminal.
48. The method according to any one of claims 37-41, characterized in that, The acquisition of the target distance includes: Obtain the screen ratio of the target person's face in the first image; The target distance is obtained based on the screen ratio and the field of view (FOV) of the first image; or, The target distance is obtained using the EXIF information of the first image.
49. The method according to any one of claims 37-41, characterized in that, The acquisition of the target distance includes: The distance to the target is obtained by a distance sensor, including a time-of-flight (TOF) sensor, a structured light sensor, or a binocular sensor.
50. The method according to any one of claims 37-41, characterized in that, The distortion correction of the first image based on the target distance includes: Obtain the correction distance; for the same target distance, the larger the value of the correction distance, the greater the degree of distortion correction; The first image is subjected to distortion correction based on the target distance and the correction distance.
51. The method according to any one of claims 37-41, characterized in that, The distortion correction of the first image based on the target distance includes: Obtain the correction distance; for the same target distance, the larger the value of the correction distance, the greater the degree of distortion correction; Obtain the region of interest; the region of interest includes at least one of the following: eyebrows, eyes, nose, mouth, or ears; The first image is subjected to distortion correction based on the target distance, the region of interest, and the correction distance.
52. The method according to claim 50, characterized in that, The acquisition of the correction distance includes: Based on the preset correspondence between shooting distance and correction distance, the correction distance corresponding to the target distance is obtained.
53. The method according to claim 50, characterized in that, The acquisition of the correction distance includes: Display the distortion correction function menu; Accept control adjustment instructions entered by the user based on the distortion correction function menu; the control adjustment instructions are used to determine the correction distance.
54. The method according to claim 50, characterized in that, The distortion correction of the first image based on the target distance and the correction distance includes: Based on the target distance, the face of the target person in the first image is fitted with a standard face model to obtain a first three-dimensional model; the first three-dimensional model is the three-dimensional model corresponding to the face of the target person. Based on the correction distance, the first three-dimensional model is transformed into a second three-dimensional model; Based on the coordinate system of the first image, the first three-dimensional model is subjected to perspective projection to obtain the first set of projection points; Based on the coordinate system of the first image, the second three-dimensional model is subjected to perspective projection to obtain a second set of projection points; The displacement vector is obtained by aligning the first set of projection points and the second set of projection points. The first image is transformed according to the displacement vector to obtain the second image.
55. The method according to claim 51, characterized in that, The distortion correction of the first image based on the target distance, the region of interest, and the correction distance includes: Based on the target distance, the face of the target person in the first image is fitted with a standard face model to obtain a first three-dimensional model; the first three-dimensional model is the three-dimensional model corresponding to the face of the target person. Based on the correction distance, the first three-dimensional model is transformed into a second three-dimensional model; The first set of projection points is obtained by performing perspective projection on the model region corresponding to the region of interest in the first three-dimensional model. The second set of projection points is obtained by performing perspective projection on the model region corresponding to the region of interest in the second three-dimensional model. The displacement vector is obtained by aligning the first set of projection points and the second set of projection points. The first image is transformed according to the displacement vector to obtain the second image.
56. The method according to any one of claims 39-41, characterized in that, The first angle range is within [-30°, 30°], and the second angle range is within [-30°, 30°].
57. The method according to claim 41, characterized in that, The distortion shooting distance is no more than 60cm.
58. The method according to claim 45, characterized in that, The first preset distance range includes [30cm, 50cm].
59. The method according to claim 50, characterized in that, The correction distance is greater than the target distance.
60. The method according to any one of claims 37-41, characterized in that, Before performing the first processing on the first image to obtain the second image, the method further includes: triggering the shutter.
61. The method according to claim 53, characterized in that, The display distortion correction function menu includes: Display distortion correction controls, which are used to open the distortion correction function menu; In response to the user enabling the distortion correction control, the distortion correction function menu is displayed.
62. The method according to claim 53, characterized in that, The distortion correction function menu includes controls for determining the correction distance, and / or controls for selecting the region of interest, and / or controls for adjusting facial expressions, and / or controls for adjusting posture.
63. An image processing apparatus, characterized in that, The device includes: An acquisition module is configured to acquire a first image, the first image including the face of a target person; wherein the face of the target person in the first image is distorted; the distortion of the face of the target person in the first image is caused by the second terminal capturing the first image at a target distance less than a first preset threshold; wherein the target distance includes the distance between the foremost part of the target person's face and the front-facing camera; or, the distance between a specified part of the target person's face and the front-facing camera; or, the distance between the center of the target person's face and the front-facing camera; and is further configured to acquire a target distance; the target distance is used to characterize the distance between the target person's face and the capturing terminal when capturing the first image; A processing module is configured to perform a first processing on the first image to obtain a second image; wherein the first processing includes distortion correction of the face of a target person in the first image based on the target distance, including: fitting the face of the target person in the first image to a standard face model based on the target distance to obtain depth information of the face of the target person; and performing perspective distortion correction on the first image based on the depth information and transformation parameters to obtain the second image; the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.
64. The apparatus according to claim 63, characterized in that, The device further includes a judgment module for determining whether one or more of the following conditions are met: Condition 1: In the first image, the completeness of the target person's face is greater than a preset completeness threshold; or, the facial features of the target person are determined to be complete. Condition 2: The pitch angle of the target person's face deviating from the frontal face conforms to the first angle range, and the yaw angle of the target person's face deviating from the frontal face conforms to the second angle range; Condition 3: Face detection is performed on the current shooting scene, and only one valid face is detected; the valid face is the face of the target person. Condition 4: Determine that the target distance is less than or equal to the preset distortion shooting distance.
65. The apparatus according to claim 63 or 64, characterized in that, The acquisition module is specifically used for: A first image is obtained by capturing a frame of an image of a target scene using a camera; or... A first image is synthesized by capturing multiple frames of images of a target scene using a camera; or... Retrieve the first image from images stored locally or in the cloud.
66. The apparatus according to claim 63 or 64, characterized in that, The camera in question is a front-facing camera.
67. The apparatus according to claim 63 or 64, characterized in that, The device also includes a setting module for setting the shooting mode to a preset shooting mode before the camera captures images of the target scene.
68. The apparatus according to claim 63 or 64, characterized in that, When the target distance is within a first preset distance range, the smaller the target distance, the greater the intensity of the distortion correction; wherein, the first preset distance range is less than or equal to a preset distortion shooting distance.
69. The apparatus according to claim 63 or 64, characterized in that, The greater the angle at which the target person's face deviates from a frontal view, the weaker the intensity of the distortion correction.
70. The apparatus according to claim 63 or 64, characterized in that, The distance between the target person's face and the shooting terminal is defined as the distance between the foremost part, center position, eyebrows, eyes, nose, mouth, or ears of the target person's face and the shooting terminal.
71. The apparatus according to claim 63 or 64, characterized in that, The acquisition module is specifically used for: Obtain the screen ratio of the target person's face in the first image; The target distance is obtained based on the screen ratio and the field of view (FOV) of the first image; or, The target distance is obtained using the EXIF information of the first image; or, The distance to the target is obtained by a distance sensor, including a time-of-flight (TOF) sensor, a structured light sensor, or a binocular sensor.
72. The apparatus according to claim 63 or 64, characterized in that, The processing module is specifically used for: Obtain the correction distance; for the same target distance, the larger the value of the correction distance, the greater the degree of distortion correction; The first image is subjected to distortion correction based on the target distance and the correction distance.
73. The apparatus according to claim 63 or 64, characterized in that, The processing module is specifically used for: Obtain the correction distance; for the same target distance, the larger the value of the correction distance, the greater the degree of distortion correction; Obtain the region of interest; the region of interest includes at least one of the following: eyebrows, eyes, nose, mouth, or ears; The first image is subjected to distortion correction based on the target distance, the region of interest, and the correction distance.
74. The apparatus according to claim 72, characterized in that, The processing module is specifically used for: Based on the preset correspondence between shooting distance and correction distance, the correction distance corresponding to the target distance is obtained.
75. The apparatus according to claim 72, characterized in that, The device also includes a display module for displaying the distortion correction function menu; The processing module is specifically used to: accept control adjustment instructions input by the user based on the distortion correction function menu; the control adjustment instructions are used to determine the correction distance.
76. The apparatus according to claim 72, characterized in that, The processing module is specifically used for: Based on the target distance, the face of the target person in the first image is fitted with a standard face model to obtain a first three-dimensional model; the first three-dimensional model is the three-dimensional model corresponding to the face of the target person. Based on the correction distance, the first three-dimensional model is transformed into a second three-dimensional model; Based on the coordinate system of the first image, the first three-dimensional model is subjected to perspective projection to obtain the first set of projection points; Based on the coordinate system of the first image, the second three-dimensional model is subjected to perspective projection to obtain a second set of projection points; The displacement vector is obtained by aligning the first set of projection points and the second set of projection points. The first image is transformed according to the displacement vector to obtain the second image.
77. The apparatus according to claim 73, characterized in that, The processing module is specifically used for: Based on the target distance, the face of the target person in the first image is fitted with a standard face model to obtain a first three-dimensional model; the first three-dimensional model is the three-dimensional model corresponding to the face of the target person. Based on the correction distance, the first three-dimensional model is transformed into a second three-dimensional model; The first set of projection points is obtained by performing perspective projection on the model region corresponding to the region of interest in the first three-dimensional model. The second set of projection points is obtained by performing perspective projection on the model region corresponding to the region of interest in the second three-dimensional model. The displacement vector is obtained by aligning the first set of projection points and the second set of projection points. The first image is transformed according to the displacement vector to obtain the second image.
78. The apparatus according to claim 64, characterized in that, The first angle range is within [-30°, 30°], and the second angle range is within [-30°, 30°].
79. The apparatus according to claim 64, characterized in that, The distortion shooting distance is no more than 60cm.
80. The apparatus according to claim 68, characterized in that, The first preset distance range includes [30cm, 50cm].
81. The apparatus according to claim 72, characterized in that, The correction distance is greater than the target distance.
82. The apparatus according to claim 63 or 64, characterized in that, The device further includes a triggering module for triggering the shutter before the processing module performs a first processing on the first image to obtain a second image.
83. The apparatus according to claim 75, characterized in that, The display module is specifically used for: Display distortion correction controls, which are used to open the distortion correction function menu; In response to the user enabling the distortion correction control, the distortion correction function menu is displayed.
84. The apparatus according to claim 63 or 64, characterized in that, The distortion correction function menu includes controls for determining the correction distance, and / or controls for selecting the region of interest, and / or controls for adjusting facial expressions, and / or controls for adjusting posture.
85. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a bus; the memory and the processor are connected via the bus; wherein... The memory is used to store computer programs and instructions; The processor is used to invoke the computer program and instructions stored in the memory, so that the terminal device performs the method as described in any one of claims 37-62.
86. The terminal device according to claim 85, characterized in that, The terminal device also includes a camera, which is used to acquire images under the control of the processor.
Citation Information
Patent Citations
Method and apparatus for face image correction
CN105405104A
Wide-angle lens 3D distortion correction method and device, terminal and storage medium
CN110189269A
Screen display control method and electronic equipment
CN111049973A