Image transformation method and apparatus

The image transformation method addresses the issue of perspective distortion in close-up selfies by measuring and correcting facial distances using three-dimensional modeling, resulting in images with more accurate facial proportions.

CN113850709BActive Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010600182.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-28
Publication Date
2025-07-15
Estimated Expiration
2040-06-28

AI Technical Summary

Technical Problem

When taking selfies, the distance between the front camera and the face is close, resulting in facial perspective distortion, which affects the selfie effect, especially the authenticity of facial features and position.

Method used

By obtaining the distance between the face and the camera, using a three-dimensional model for distortion correction, establishing a three-dimensional model of the target person's face, performing perspective distortion correction, and restoring the relative proportion and position of the facial features.

Benefits of technology

It significantly improves the shooting and imaging effect in self-portrait scenes, so that the relative proportion and position of facial features of a person are closer to the true appearance, and eliminates the problem of distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850709B_ABST
    Figure CN113850709B_ABST
Patent Text Reader

Abstract

The present application provides an image transformation method and apparatus. The image transformation method of the present application includes: obtaining a first image of a target scene through a front camera, where the target scene includes the face of a target person; obtaining a target distance between the face of the target person and the front camera; when the target distance is less than a preset threshold, performing a first process on the first image to obtain a second image; the first process includes performing distortion correction on the first image according to the target distance; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image. The present application can improve the shooting imaging effect in a self-shooting scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image processing technology, and in particular, to an image transformation method and apparatus. Background Art

[0002] Photography has become an important medium for recording life. In recent years, "selfies" taken with the front camera of mobile phones have become increasingly popular. However, due to the relatively short distance between the camera and the face during selfies, the perspective distortion problem of the face, where "objects closer appear larger and those farther appear smaller", has become increasingly prominent. For example, when taking a portrait at close range, due to the difference in the distance between different parts of the face and the camera, the nose usually appears larger, and at the same time, the face is elongated, affecting the subjective effect of the portrait. Therefore, image transformation processing is required. For example, through transformation of the distance, pose, position, etc. of the target in the image, while ensuring that the authenticity of the portrait is not affected, the perspective distortion effect of "objects closer appear larger and those farther appear smaller" is eliminated, and the aesthetic degree of the portrait is improved.

[0003] The imaging process of the camera results in a two-dimensional image of a three-dimensional object. Commonly used image processing algorithms are for two-dimensional images, but this makes it difficult to achieve the effect of real transformation of three-dimensional objects for the image obtained after two-dimensional image transformation. Summary of the Invention

[0004] This application provides an image transformation method and apparatus to improve the imaging effect in the selfie scenario.

[0005] In a first aspect, this application provides an image transformation method, including: obtaining a first image of a target scene through a front camera, where the target scene includes the face of a target person; obtaining a target distance between the face of the target person and the front camera; when the target distance is less than a preset threshold, performing a first process on the first image to obtain a second image; the first process includes performing distortion correction on the first image according to the target distance; where the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0006] The first image is captured by the front camera of the terminal in the scenario where the user takes a selfie. The first image includes two situations. The first situation is that when the user does not trigger the shutter, the first image has not yet been imaged on the image sensor, and is only an original picture captured by the camera; the other situation is that the user triggers the shutter, and the first image is the original image that has been imaged on the image sensor. Therefore, relative to the first image, the second image also includes two situations. Corresponding to the former situation, the second image is a corrected picture obtained by performing distortion correction based on the original picture captured by the camera. The picture is also not imaged on the image sensor and is only used as a preview image for the user; corresponding to the latter situation, the second image is a corrected image obtained by performing distortion correction based on the original image that has been imaged on the image sensor. The terminal can save the corrected image and store it in the picture library.

[0007] The present application achieves a three-dimensional transformation effect of an image with the help of a three-dimensional model, and corrects the perspective distortion of the face of a target person in an image taken at a close distance, so that the relative proportions and relative positions of the facial features of the target person after correction are closer to the relative proportions and relative positions of the facial features of the target person, thereby significantly improving the shooting imaging effect in selfie scenes.

[0008] In one possible implementation, the target distance includes the distance between the frontmost part of the face of the target person and the front camera; or, the distance between a specified part of the face of the target person and the front camera; or, the distance between the center position of the face of the target person and the front camera.

[0009] The target distance between the target person's face and the camera may be the distance between the frontmost part of the target person's face (e.g., the nose) and the camera; or, the target distance between the target person's face and the camera may be the distance between a designated part of the target person's face (e.g., the eyes, mouth, or nose, etc.) and the camera; or, the target distance between the target person's face and the camera may be the distance between the center position of the target person's face (e.g., the nose of the target person in a front view, or the position of the cheekbones in a side view of the target person, etc.) and the camera. It should be noted that the definition of the above target distance may depend on the specific situation of the first image, and this application does not make any specific limitation on this.

[0010] In a possible implementation, obtaining the target distance between the face of the target person and the front camera includes: obtaining a screen-to-body ratio of the face of the target person in the first image; and obtaining the target distance according to the screen-to-body ratio and a field of view FOV of the front camera.

[0011] In a possible implementation, obtaining the target distance between the face of the target person and the front camera includes: obtaining the target distance through a distance sensor, where the distance sensor includes a Time-of-Flight (TOF) sensor, a structured light sensor, or a binocular sensor.

[0012] This application can obtain the above target distance by calculating the face screen occupation ratio, that is, first obtaining the screen occupation ratio of the face of the target person in the first image (the ratio of the pixel area of the face to the pixel area of the first image), and then obtaining the above target distance based on this screen occupation ratio and the Field of View (FOV) of the front camera. The above target distance can also be measured by a distance sensor. Other methods can also be used to obtain this target distance, and this application does not make specific limitations in this regard.

[0013] In a possible implementation, the preset threshold is less than 80 centimeters.

[0014] In a possible implementation, the preset threshold is 50 centimeters.

[0015] This application sets a threshold. When it is considered that the target distance between the face of the target person and the front camera is less than this threshold, the first image containing the face of the target person is distorted and needs to be corrected for distortion. The value range of this preset threshold is within 80 centimeters. Optionally, this preset threshold can be set to 50 centimeters. It should be noted that the specific value of the preset threshold can be determined according to the performance of the front camera, shooting light, etc., and this application does not make specific limitations in this regard.

[0016] In a possible implementation, the second image includes a preview image or an image obtained after triggering the shutter.

[0017] The second image can be a preview image obtained by the front camera. That is, before the shutter is triggered, the front camera can obtain the first image facing the target scene area. The terminal uses the above method to perform perspective distortion correction on this first image to obtain the second image, and displays the second image on the screen of the terminal. At this time, the second image seen by the user is a preview image that has been corrected for perspective distortion. The second image can also be an image obtained after triggering the shutter. That is, after the shutter is triggered, the first image is the image formed on the image sensor of the terminal. The terminal uses the above method to perform perspective distortion correction on the first image to obtain the second image, saves the second image, and displays it on the screen of the terminal. At this time, the second image seen by the user is an image in the picture library that has been corrected for perspective distortion and has been saved.

[0018] In a possible implementation, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the face of the target person than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the face of the target person than the relative positions of the facial features of the target person in the first image.

[0019] When the first image is acquired, the target distance between the face of the target person and the front camera is less than a preset threshold. It is very likely that due to the perspective distortion problem of the front camera where "objects appear larger when closer and smaller when farther", the sizes of the facial features of the face of the target person in the first image change, stretch, etc., and the relative proportions and relative positions of its facial features deviate from the relative proportions and relative positions of the true appearance of the face of the target person. The second image obtained by performing perspective distortion correction on the first image can eliminate the above-mentioned situations such as the size change and stretching of the facial features of the face of the target person, so that the relative proportions and relative positions of the facial features of the face of the target person in the second image approach, or even return to, the relative proportions and relative positions of the true appearance of the face of the target person.

[0020] In a possible implementation, the performing distortion correction on the first image according to the target distance includes: fitting the face of the target person in the first image with a standard face model according to the target distance to obtain the depth information of the face of the target person; performing perspective distortion correction on the first image according to the depth information to obtain the second image.

[0021] In a possible implementation, the performing perspective distortion correction on the first image according to the depth information to obtain the second image includes: establishing a first three-dimensional model of the face of the target person; transforming the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the face of the target person; obtaining a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; obtaining the second image according to the pixel displacement vector field of the face of the target person.

[0022] A three-dimensional model is established based on the face of a target person, and then a three-dimensional transformation effect of an image is realized through the three-dimensional model. A pixel displacement vector field of the face of the target person is obtained based on the association of sampling points between the three-dimensional models before and after the transformation, and then a transformed two-dimensional image can be obtained, which can realize perspective distortion correction of the face of the target person in an image taken at a close distance, so that the relative proportions and relative positions of the facial features of the face of the target person after correction are closer to the relative proportions and relative positions of the facial features of the face of the target person, and the imaging effect in the selfie scenario can be significantly improved.

[0023] In a possible implementation manner, a perspective projection is performed on the first three-dimensional model according to the depth information to obtain a first coordinate set, where the first coordinate set includes coordinate values corresponding to a plurality of pixels in the first three-dimensional model; a perspective projection is performed on the second three-dimensional model according to the depth information to obtain a second coordinate set, where the second coordinate set includes coordinate values corresponding to a plurality of pixels in the second three-dimensional model; a coordinate difference between the first coordinate value and the second coordinate value is calculated to obtain the pixel displacement vector field of the target object, where the first coordinate value includes the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel includes any one of a plurality of identical pixels included in the first three-dimensional model and the second three-dimensional model.

[0024] In a second aspect, the present application provides an image transformation method, including: obtaining a first image, where the first image includes the face of a target person; where, the face of the target person in the first image is distorted; displaying a distortion correction function menu; obtaining transformation parameters input by a user on the distortion correction function menu, where the transformation parameters at least include an equivalent simulated shooting distance, and the equivalent simulated shooting distance is used to simulate the distance between the face of the target person and the camera when the shooting terminal shoots the face of the target person; performing a first process on the first image to obtain a second image; the first process includes performing distortion correction on the first image according to the transformation parameters; where, the face of the target person in the second image is closer to the real appearance of the face of the target person than the face of the target person in the first image.

[0025] The present application performs perspective distortion correction on the first image according to the transformation parameters to obtain a second image, where the face of the target person in the second image is closer to the real appearance of the face of the target person than the face of the target person in the first image, that is, the relative proportions and relative positions of the facial features of the target person in the second image are closer to the relative proportions and relative positions of the facial features of the face of the target person than the relative proportions and relative positions of the facial features of the target person in the first image.

[0026] In a possible implementation, the distortion of the face of the target person in the first image is caused by the fact that the target distance between the face of the target person and the second terminal is less than a first preset threshold when the second terminal captures the first image; wherein, the target distance includes the distance between the foremost part on the face of the target person and the front camera; or, the distance between a specified part on the face of the target person and the front camera; or, the distance between the center position on the face of the target person and the front camera.

[0027] In a possible implementation, the target distance is obtained based on the screen occupation ratio of the face of the target person in the first image and the FOV of the camera of the second terminal; or, the target distance is obtained based on the equivalent focal length in the EXIF information of the first image in the Exchangeable Image File Format.

[0028] In a possible implementation, the distortion correction function menu includes an option for adjusting the equivalent simulated shooting distance; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the equivalent simulated shooting distance according to an instruction triggered by the user's operation on a control or slider in the option for adjusting the equivalent simulated shooting distance.

[0029] In a possible implementation, when the distortion correction function menu is initially displayed, the value of the equivalent simulated shooting distance in the option for adjusting the equivalent simulated shooting distance includes a default value or a value obtained by pre-calculation.

[0030] In a possible implementation, before displaying the distortion correction function menu, it further includes: when the face of the target person is distorted, displaying a pop-up window, where the pop-up window is used to provide a selection control for whether to perform distortion correction; when the user clicks the control for performing distortion correction on the pop-up window, responding to the instruction generated by the user's operation.

[0031] In a possible implementation, before displaying the distortion correction function menu, it further includes: when the face of the target person is distorted, displaying a distortion correction control, where the distortion correction control is used to open the distortion correction function menu; when the user clicks the distortion correction control, responding to the instruction generated by the user's operation.

[0032] In a possible implementation, the distortion of the face of the target person in the first image is caused by the fact that when the second terminal captures the first image, the field of view (FOV) of the camera is greater than a second preset threshold, and the pixel distance between the face of the target person and the edge of the FOV is less than a third preset threshold; wherein the pixel distance includes the number of pixels between the forefront part on the face of the target person and the edge of the FOV; or, the number of pixels between a specified part on the face of the target person and the edge of the FOV; or, the number of pixels between the center position on the face of the target person and the edge of the FOV.

[0033] In a possible implementation, the FOV is obtained through the EXIF information of the first image.

[0034] In a possible implementation, the second preset threshold is 90°, and the third preset threshold is one-fourth of the length or width of the first image.

[0035] In a possible implementation, the distortion correction function menu includes an option to adjust the displacement distance; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction and displacement distance according to the instruction triggered by the user's operation on the control or slider in the option to adjust the displacement distance.

[0036] In a possible implementation, the distortion correction function menu includes options to adjust the relative positions and / or relative proportions of facial features; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction, displacement distance, and / or facial feature sizes according to the instruction triggered by the user's operation on the control or slider in the option to adjust the relative positions and / or relative proportions of facial features.

[0037] In a possible implementation, the distortion correction function menu includes an option to adjust the angle; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining the adjustment direction and adjustment angle according to the instruction triggered by the user's operation on the control or slider in the option to adjust the angle; or, the distortion correction function menu includes an option to adjust the expression; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining a new expression template according to the instruction triggered by the user's operation on the control or slider in the option to adjust the expression; or, the distortion correction function menu includes an option to adjust the action; obtaining the transformation parameters input by the user on the distortion correction function menu includes: obtaining a new action template according to the instruction triggered by the user's operation on the control or slider in the option to adjust the action.

[0038] In a possible implementation, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the face of the target person than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the face of the target person than the relative positions of the facial features of the target person in the first image.

[0039] In a possible implementation, the distortion correction of the first image according to the transformation parameters includes: fitting the face of the target person in the first image with a standard face model according to the target distance to obtain the depth information of the face of the target person; performing perspective distortion correction on the first image according to the depth information and the transformation parameters to obtain the second image.

[0040] In a possible implementation, the performing perspective distortion correction on the first image according to the depth information and the transformation parameters to obtain the second image includes: establishing a first three-dimensional model of the face of the target person; transforming the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the face of the target person; obtaining a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtaining the second image according to the pixel displacement vector field of the face of the target person.

[0041] In a possible implementation, the obtaining a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model includes: performing perspective projection on the first three-dimensional model according to the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to a plurality of pixels in the first three-dimensional model; performing perspective projection on the second three-dimensional model according to the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to a plurality of pixels in the second three-dimensional model; calculating the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of the plurality of identical pixels included in the first three-dimensional model and the second three-dimensional model.

[0042] In a third aspect, the present application provides an image transformation method, which is applied to a first terminal. The method includes: obtaining a first image, where the first image includes the face of a target person; wherein, the face of the target person in the first image is distorted; displaying a distortion correction function menu on the screen, where the distortion correction function menu includes one or more sliders and / or one or more controls; receiving a distortion correction instruction, where the distortion correction instruction includes transformation parameters generated when the user performs a touch operation on the one or more sliders and / or the one or more controls, and the transformation parameters at least include an equivalent simulated shooting distance, and the equivalent simulated shooting distance is used to simulate the distance between the face of the target person and the camera when the shooting terminal shoots the face of the target person; performing a first process on the first image according to the transformation parameters to obtain a second image; the first process includes performing distortion correction on the first image; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0043] In a fourth aspect, the present application provides an image transformation device, including: an acquisition module, configured to obtain a first image for a target scene through a front camera, where the target scene includes the face of a target person; obtaining a target distance between the face of the target person and the front camera; a processing module, configured to, when the target distance is less than a preset threshold, perform a first process on the first image to obtain a second image; the first process includes performing distortion correction on the first image according to the target distance; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0044] In a possible implementation manner, the target distance includes the distance between the foremost part on the face of the target person and the front camera; or, the distance between a specified part on the face of the target person and the front camera; or, the distance between the central position on the face of the target person and the front camera.

[0045] In a possible implementation manner, the acquisition module is specifically configured to obtain the screen occupation ratio of the face of the target person in the first image; and obtain the target distance according to the screen occupation ratio and the field of view angle FOV of the front camera.

[0046] In a possible implementation manner, the acquisition module is specifically configured to obtain the target distance through a distance sensor, and the distance sensor includes a time-of-flight ranging method TOF sensor, a structured light sensor or a binocular sensor.

[0047] In a possible implementation, the preset threshold is less than 80 cm.

[0048] In a possible implementation, the second image includes a preview image or an image obtained after triggering the shutter.

[0049] In a possible implementation, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image, including: the relative proportion of the facial features of the target person in the second image is closer to the relative proportion of the facial features of the face of the target person than the relative proportion of the facial features of the target person in the first image; and / or, the relative position of the facial features of the target person in the second image is closer to the relative position of the facial features of the face of the target person than the relative position of the facial features of the target person in the first image.

[0050] In a possible implementation, the processing module is specifically configured to fit the face of the target person in the first image with a standard face model according to the target distance to obtain the depth information of the face of the target person; and correct the perspective distortion of the first image according to the depth information to obtain the second image.

[0051] In a possible implementation, the processing module is specifically configured to establish a first three-dimensional model of the face of the target person; transform the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the face of the target person; obtain a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the face of the target person.

[0052] In a possible implementation, the processing module is specifically configured to perform perspective projection on the first three-dimensional model according to the depth information to obtain a first coordinate set, the first coordinate set including coordinate values corresponding to a plurality of pixels in the first three-dimensional model; perform perspective projection on the second three-dimensional model according to the depth information to obtain a second coordinate set, the second coordinate set including coordinate values corresponding to a plurality of pixels in the second three-dimensional model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, the first coordinate value including the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value including the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel including any one of a plurality of identical pixels included in the first three-dimensional model and the second three-dimensional model.

[0053] Fifth aspect, the present application provides an image transformation device, including: an acquisition module, configured to acquire a first image, where the first image includes a face of a target person; wherein, the face of the target person in the first image is distorted; a display module, configured to display a distortion correction function menu; the acquisition module is further configured to acquire transformation parameters input by a user on the distortion correction function menu, where the transformation parameters at least include an equivalent simulated shooting distance, and the equivalent simulated shooting distance is used to simulate the distance between the face of the target person and the camera when a shooting terminal shoots the face of the target person; a processing module, configured to perform a first processing on the first image to obtain a second image; the first processing includes performing distortion correction on the first image according to the transformation parameters; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0054] In a possible implementation manner, the distortion of the face of the target person in the first image is caused by the fact that the target distance between the face of the target person and the second terminal is less than a first preset threshold when the second terminal shoots the first image; wherein, the target distance includes the distance between the foremost part on the face of the target person and the front camera; or, the distance between a specified part on the face of the target person and the front camera; or, the distance between the center position on the face of the target person and the front camera.

[0055] In a possible implementation manner, the target distance is obtained through the screen occupation ratio of the face of the target person in the first image and the FOV of the camera of the second terminal; or, the target distance is obtained through the equivalent focal length in the EXIF information of the first image in the exchangeable image file format.

[0056] In a possible implementation manner, the distortion correction function menu includes an option for adjusting the equivalent simulated shooting distance; the acquisition module is specifically configured to acquire the equivalent simulated shooting distance according to an instruction triggered by the user's operation on a control or a slider in the option for adjusting the equivalent simulated shooting distance.

[0057] In a possible implementation manner, when the distortion correction function menu is initially displayed, the value of the equivalent simulated shooting distance in the option for adjusting the equivalent simulated shooting distance includes a default value or a pre-calculated value.

[0058] In a possible implementation, the display module is further configured to display a pop-up window when the face of the target person is distorted, where the pop-up window is used to provide a selection control for whether to perform distortion correction; when the user clicks on the control for performing distortion correction on the pop-up window, an instruction generated in response to the user operation is triggered.

[0059] In a possible implementation, the display module is further configured to display a distortion correction control when the face of the target person is distorted, where the distortion correction control is used to open the distortion correction function menu; when the user clicks on the distortion correction control, an instruction generated in response to the user operation is triggered.

[0060] In a possible implementation, the distortion of the face of the target person in the first image is caused by the fact that when the second terminal captures the first image, the field of view angle FOV of the camera is greater than a second preset threshold, and the pixel distance between the face of the target person and the edge of the FOV is less than a third preset threshold; where the pixel distance includes the number of pixels between the foremost part on the face of the target person and the edge of the FOV; or, the number of pixels between a specified part on the face of the target person and the edge of the FOV; or, the number of pixels between the center position on the face of the target person and the edge of the FOV.

[0061] In a possible implementation, the FOV is obtained through the EXIF information of the first image.

[0062] In a possible implementation, the second preset threshold is 90°, and the third preset threshold is one-fourth of the length or width of the first image.

[0063] In a possible implementation, the distortion correction function menu includes an option for adjusting the displacement distance; the acquisition module is further configured to obtain the adjustment direction and the displacement distance according to an instruction triggered by the user's operation on the control or slider in the option for adjusting the displacement distance.

[0064] In a possible implementation, the distortion correction function menu includes an option for adjusting the relative position and / or relative proportion of the facial features; the acquisition module is further configured to obtain the adjustment direction, the displacement distance, and / or the facial feature size according to an instruction triggered by the user's operation on the control or slider in the option for adjusting the relative position and / or relative proportion of the facial features.

[0065] In a possible implementation, the distortion correction function menu includes an option for adjusting the angle; the obtaining module is further configured to obtain an adjustment direction and an adjustment angle according to an instruction triggered by a user's operation on a control or a slider in the option for adjusting the angle; or, the distortion correction function menu includes an option for adjusting an expression; the obtaining module is further configured to obtain a new expression template according to an instruction triggered by a user's operation on a control or a slider in the option for adjusting the expression; or, the distortion correction function menu includes an option for adjusting an action; the obtaining module is further configured to obtain a new action template according to an instruction triggered by a user's operation on a control or a slider in the option for adjusting the action.

[0066] In a possible implementation, that the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image includes: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the face of the target person than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the face of the target person than the relative positions of the facial features of the target person in the first image.

[0067] In a possible implementation, the processing module is specifically configured to fit the face of the target person in the first image with a standard face model according to the target distance to obtain depth information of the face of the target person; and perform perspective distortion correction on the first image according to the depth information and the transformation parameters to obtain the second image.

[0068] In a possible implementation, the processing module is specifically configured to establish a first three-dimensional model of the face of the target person; transform the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the face of the target person; obtain a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the face of the target person.

[0069] In a possible implementation, the processing module is specifically configured to perform perspective projection on the first 3D model according to the depth information to obtain a first coordinate set, where the first coordinate set includes coordinate values corresponding to a plurality of pixels in the first 3D model; perform perspective projection on the second 3D model according to the depth information to obtain a second coordinate set, where the second coordinate set includes coordinate values corresponding to a plurality of pixels in the second 3D model; calculate a coordinate difference between the first coordinate value and the second coordinate value to obtain a pixel displacement vector field of the target object, where the first coordinate value includes the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel includes any one of a plurality of identical pixels included in the first 3D model and the second 3D model.

[0070] In a possible implementation, it further includes: a recording module; the obtaining module is further configured to obtain a recording instruction according to a trigger operation of the user on the recording control; the recording module is configured to start recording the obtaining process of the second image according to the recording instruction until a stop recording instruction generated by a trigger operation of the user on the stop recording control is received.

[0071] In a sixth aspect, the present application provides an image transformation device, including: an obtaining module, configured to obtain a first image, where the first image includes a face of a target person; wherein, the face of the target person in the first image is distorted; a display module, configured to display a distortion correction function menu on a screen, where the distortion correction function menu includes one or more sliders and / or one or more controls; receive a distortion correction instruction, where the distortion correction instruction includes transformation parameters generated when the user performs a touch operation on the one or more sliders and / or the one or more controls, and the transformation parameters at least include an equivalent simulated shooting distance, and the equivalent simulated shooting distance is used to simulate the distance between the face of the target person and the camera when the shooting terminal shoots the face of the target person; a processing module, configured to perform a first process on the first image according to the transformation parameters to obtain a second image; the first process includes performing distortion correction on the first image; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0072] In a seventh aspect, the present application provides a device, including: one or more processors; a memory, configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of the first to third aspects described above.

[0073] In an eighth aspect, the present application provides a computer-readable storage medium including a computer program which, when executed on a computer, causes the computer to execute the method according to any one of the first to third aspects described above.

[0074] In a ninth aspect, the present application further provides a computer program product which includes computer program code that, when running on a computer, causes the computer to execute the method according to any one of the first to third aspects described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 FIG. shows an exemplary schematic diagram of an application architecture applicable to the image transformation method of the present application;

[0076] Figure 2 FIG. shows an exemplary structural diagram of a terminal 200;

[0077] Figure 3 is a flowchart of the first embodiment of the image transformation method of the present application;

[0078] Figure 4 FIG. shows an exemplary schematic diagram of a distance acquisition method;

[0079] Figure 5 FIG. shows an exemplary schematic diagram of a process for creating a three-dimensional face model;

[0080] Figure 6 FIG. shows an exemplary schematic diagram of an angular change in position movement;

[0081] Figure 7 FIG. shows an exemplary schematic diagram of perspective projection;

[0082] Figure 8a and Figure 8b exemplarily show the face projection effects at object distances of 30 cm and 55 cm respectively;

[0083] Figure 9 FIG. shows an exemplary schematic diagram of a pixel displacement vector expansion method;

[0084] Figure 10 is a flowchart of the second embodiment of the image transformation method of the present application;

[0085] Figure 11 FIG. shows an exemplary schematic diagram of a distortion correction function menu;

[0086] Figure 12a - Figure 12f exemplarily shows the process of a terminal performing distortion correction in a selfie scenario;

[0087] Figure 13a - Figure 13hExemplarily shows the process of distorting and correcting an image in an image library;

[0088] Figure 14 Exemplarily shows other examples of the distortion correction function menu;

[0089] Figure 15 This is a schematic structural diagram of an embodiment of the image transformation device of the present application. Detailed implementation manners

[0090] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0091] In the description of the embodiments of the present application, the terms "first", "second", etc. are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0092] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c may be single or multiple.

[0093] The present application proposes an image transformation method, so that the transformed image, especially the target object in the image, achieves the effect of real transformation in three-dimensional space.

[0094] Figure 1An exemplary schematic diagram showing an application architecture applicable to the image transformation method of the present application is shown as follows: Figure 1 As shown, the framework includes: an image acquisition module, an image processing module and a display module, wherein the image acquisition module is used to shoot or obtain the image to be processed, and the image acquisition module can be, for example, a camera, a video camera and other devices; the image processing module is used to transform the image to be processed to achieve distortion correction, and the image processing module can be any device with image processing capabilities, such as a terminal, a picture server, etc., and can also be any chip with image processing capabilities, such as a graphics processing unit (GPU) chip; the display module is used to display the image, and the display module can be, for example, a display, a terminal screen, a television, a projector, etc.

[0095] In the present application, the image acquisition module, the image processing module and the display module can all be integrated on the same device, in which case the processor of the device acts as a control module to control the image acquisition module, the image processing module and the display module to realize their respective functions. The image acquisition module, the image processing module and the display module can also be independent devices. For example, the image acquisition module uses devices such as cameras and video cameras, and the image processing module and the display module are integrated on one device, in which case the processor of the integrated device acts as a control module to control the image processing module and the display module to realize their respective functions, and the integrated device can also have wireless or wired transmission capabilities to receive the image to be processed from the image acquisition module, and the integrated device can also be provided with an input interface to obtain the image to be processed through the input interface. For another example, the image acquisition module uses devices such as cameras and video cameras, the image processing module uses devices with image processing capabilities, such as mobile phones, tablet computers, computers, etc., and the display module uses devices such as screens and televisions, and the three are connected by wireless or wired means to realize the transmission of image data. For another example, the image acquisition module and the image processing module are integrated on one device, and the integrated device has the capabilities of image acquisition and image processing, such as a mobile phone, a tablet computer, etc. The processor of the integrated device serves as a control module to control the image acquisition module and the image processing module to realize their respective functions. The device may also have wireless or wired transmission capabilities to transmit images to the display module. The integrated device may also be provided with an output interface to transmit images through the output interface.

[0096] It should be noted that the above application architecture may also be implemented in other hardware and / or software ways, and this application does not make any specific limitations on this.

[0097] The above image processing module is the core module of this application. Devices including this image processing module can be terminals (such as mobile phones, tablets, etc.), wearable devices with wireless communication functions (such as smart watches), computers with wireless transceiver functions, virtual reality (VR) devices, augmented reality (AR) devices, etc. This application does not make any limitations in this regard.

[0098] Figure 2 A schematic structural diagram of the terminal 200 is shown.

[0099] The terminal 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, a time of flight (TOF) sensor 280N, a structured light sensor 280O, a binocular sensor 280P, etc.

[0100] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the terminal 200. In other embodiments of this application, the terminal 200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0101] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0102] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0103] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 includes a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0104] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0105] The I2C interface is a two-way synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 210 may include multiple groups of I2C buses. The processor 210 may be respectively coupled to the touch sensor 280K, the charger, the flash, the camera 293, etc. through different I2C bus interfaces. For example, the processor 210 may be coupled to the touch sensor 280K through the I2C interface, enabling the processor 210 and the touch sensor 280K to communicate through the I2C bus interface, thereby implementing the touch function of the terminal 200.

[0106] The I2S interface can be used for audio communication. In some embodiments, the processor 210 may include multiple groups of I2S buses. The processor 210 may be coupled to the audio module 270 through the I2S bus to implement communication between the processor 210 and the audio module 270. In some embodiments, the audio module 270 may transmit audio signals to the wireless communication module 260 through the I2S interface, thereby implementing the function of answering a call through a Bluetooth headset.

[0107] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 270 and the wireless communication module 260 may be coupled through the PCM bus interface. In some embodiments, the audio module 270 may also transmit audio signals to the wireless communication module 260 through the PCM interface, thereby implementing the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0108] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus may include a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor 210 and the wireless communication module 260. For example, the processor 210 communicates with the Bluetooth module in the wireless communication module 260 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 270 may transmit audio signals to the wireless communication module 260 through the UART interface, thereby implementing the function of playing music through a Bluetooth headset.

[0109] The MIPI interface can be used to connect the processor 210 to peripheral devices such as the display screen 294 and the camera 293. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 210 and the camera 293 communicate through the CSI interface to implement the shooting function of the terminal 200. The processor 210 and the display screen 294 communicate through the DSI interface to implement the display function of the terminal 200.

[0110] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 210 to the camera 293, the display screen 294, the wireless communication module 260, the audio module 270, the sensor module 280, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0111] The USB interface 230 is an interface that complies with the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 230 can be used to connect a charger to charge the terminal 200, and can also be used to transfer data between the terminal 200 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other terminals, such as AR devices, etc.

[0112] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative and do not constitute a structural limitation on the terminal 200. In other embodiments of the present application, the terminal 200 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0113] The charging management module 240 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 240 can receive the charging input from the wired charger through the USB interface 230. In some embodiments of wireless charging, the charging management module 240 can receive the wireless charging input through the wireless charging coil of the terminal 200. While charging the battery 242, the charging management module 240 can also supply power to the terminal through the power management module 241.

[0114] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives the inputs from the battery 242 and / or the charging management module 240, and supplies power to the processor 210, the internal memory 221, the display screen 294, the camera 293, the wireless communication module 260, etc. The power management module 241 can also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be disposed in the processor 210. In some other embodiments, the power management module 241 and the charging management module 240 can also be disposed in the same device.

[0115] The wireless communication function of the terminal 200 can be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, and the baseband processor, etc.

[0116] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0117] The mobile communication module 250 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal 200. The mobile communication module 250 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves by the antenna 1, filter and amplify the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 250 can be disposed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 can be disposed in the same device.

[0118] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs an audio signal through an audio device (not limited to the speaker 270A, the receiver 270B, etc.), or displays an image or video through the display screen 294. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 210 and be provided in the same device as the mobile communication module 250 or other functional modules.

[0119] The wireless communication module 260 may provide solutions for wireless communications applied to the terminal 200, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 260 may be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signal, and transmits the processed signal to the processor 210. The wireless communication module 260 may also receive the signal to be transmitted from the processor 210, perform frequency modulation and amplification on it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.

[0120] In some embodiments, antenna 1 of terminal 200 is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, enabling terminal 200 to communicate with the network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0121] Terminal 200 implements the display function through the GPU, display screen 294, and application processor, etc. The GPU is a microprocessor for image processing, connected to display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or change display information.

[0122] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal 200 may include one or N display screens 294, where N is a positive integer greater than 1.

[0123] The terminal 200 can implement the shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.

[0124] The ISP is used to process the data fed back by the camera 293. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 293.

[0125] The camera 293 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the terminal 200 may include one or N cameras 293, where N is a positive integer greater than 1. Among them, one or more cameras 293 may be disposed on the front side of the terminal 200. For example, at the middle position on the top of the screen, which can be understood as the front camera of the terminal. Corresponding to the device of the binocular sensor, there may also be two front cameras. One or more cameras 293 may also be disposed on the back side of the terminal 200. For example, at the upper left corner of the back of the terminal, which can be understood as the rear camera of the terminal.

[0126] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal 200 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0127] The video codec is used to compress or decompress digital videos. The terminal 200 can support one or more video codecs. In this way, the terminal 200 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0128] The NPU is a neural-network (NN) computing processor. By drawing on the structure of the biological neural network, such as drawing on the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the terminal 200 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0129] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 200. The external memory card communicates with the processor 210 through the external memory interface 220 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0130] The internal memory 221 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the terminal 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the terminal 200 by running the instructions stored in the internal memory 221, and / or the instructions stored in the memory provided in the processor.

[0131] The terminal 200 can implement audio functions through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headphone jack 270D, and the application processor, etc. For example, music playback, recording, etc.

[0132] The audio module 270 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 can be provided in the processor 210, or some functional modules of the audio module 270 can be provided in the processor 210.

[0133] The speaker 270A, also called a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal 200 can listen to music or listen to a hands-free call through the speaker 270A.

[0134] The receiver 270B, also called a "handset", is used to convert an audio electrical signal into a sound signal. When the terminal 200 answers a call or a voice message, the voice can be listened to by bringing the receiver 270B close to the human ear.

[0135] The microphone 270C, also called a "microphone", "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 270C to input the sound signal into the microphone 270C. The terminal 200 can be provided with at least one microphone 270C. In some other embodiments, the terminal 200 can be provided with two microphones 270C. In addition to collecting sound signals, noise reduction functions can also be achieved. In some other embodiments, the terminal 200 can also be provided with three, four or more microphones 270C to achieve sound signal collection, noise reduction, and can also identify the sound source to achieve functions such as directional recording.

[0136] The headphone jack 270D is used to connect a wired headphone. The headphone jack 270D can be a USB interface 230, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0137] The pressure sensor 280A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 280A can be disposed on the display screen 294. There are many types of pressure sensors 280A, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 280A, the capacitance between the electrodes changes. The terminal 200 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 294, the terminal 200 detects the intensity of the touch operation according to the pressure sensor 280A. The terminal 200 can also calculate the position of the touch according to the detection signal of the pressure sensor 280A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.

[0138] The gyroscope sensor 280B can be used to determine the motion posture of the terminal 200. In some embodiments, the angular velocity of the terminal 200 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 280B. The gyroscope sensor 280B can be used for anti-shake during shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 280B detects the shaking angle of the terminal 200, calculates the distance that the lens module needs to compensate according to the angle, and enables the lens to cancel the shaking of the terminal 200 through reverse movement to achieve anti-shake. The gyroscope sensor 280B can also be used for navigation and somatosensory game scenarios.

[0139] The barometric pressure sensor 280C is used to measure the barometric pressure. In some embodiments, the terminal 200 calculates the altitude according to the barometric pressure value measured by the barometric pressure sensor 280C to assist in positioning and navigation.

[0140] The magnetic sensor 280D includes a Hall sensor. The terminal 200 can use the magnetic sensor 280D to detect the opening and closing of the flip case. In some embodiments, when the terminal 200 is a flip phone, the terminal 200 can detect the opening and closing of the flip based on the magnetic sensor 280D. Furthermore, according to the detected opening and closing state of the case or the flip, features such as automatic flip unlocking can be set.

[0141] The acceleration sensor 280E can detect the magnitude of the acceleration of the terminal 200 in various directions (generally three axes). When the terminal 200 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the attitude of the terminal and is applied to functions such as horizontal and vertical screen switching and pedometers.

[0142] The distance sensor 280F is used to measure distance. The terminal 200 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the terminal 200 can use the distance sensor 280F to measure distance to achieve rapid focusing.

[0143] The proximity light sensor 280G can include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The terminal 200 emits infrared light outward through the light-emitting diode. The terminal 200 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 200. When insufficient reflected light is detected, the terminal 200 can determine that there is no object near the terminal 200. The terminal 200 can use the proximity light sensor 280G to detect when the user holds the terminal 200 close to the ear during a call, so as to automatically turn off the screen to achieve power saving. The proximity light sensor 280G can also be used for automatic unlocking and locking in the case mode and pocket mode.

[0144] The ambient light sensor 280L is used to sense the ambient light brightness. The terminal 200 can adaptively adjust the brightness of the display screen 294 according to the sensed ambient light brightness. The ambient light sensor 280L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 280L can also cooperate with the proximity light sensor 280G to detect whether the terminal 200 is in the pocket to prevent accidental touch.

[0145] The fingerprint sensor 280H is used to collect fingerprints. The terminal 200 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering of incoming calls, etc.

[0146] The temperature sensor 280J is used to detect temperature. In some embodiments, the terminal 200 executes a temperature processing strategy using the temperature detected by the temperature sensor 280J. For example, when the temperature reported by the temperature sensor 280J exceeds a threshold, the terminal 200 reduces the performance of the processor near the temperature sensor 280J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the terminal 200 heats the battery 242 to prevent the terminal 200 from shutting down abnormally due to low temperature. In still other embodiments, when the temperature is lower than yet another threshold, the terminal 200 boosts the output voltage of the battery 242 to prevent abnormal shutdown caused by low temperature.

[0147] The touch sensor 280K, also referred to as a "touch control device". The touch sensor 280K can be disposed on the display screen 294, and together with the display screen 294 forms a touch screen, also referred to as a "touch control screen". The touch sensor 280K is used to detect touch operations applied thereto or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 294. In other embodiments, the touch sensor 280K can also be disposed on the surface of the terminal 200, at a different position from that of the display screen 294.

[0148] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 280M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 280M can also be disposed in the earphone to form a bone conduction earphone. The audio module 270 can analyze the voice signal based on the vibration signal of the vibrating bone mass of the vocal part acquired by the bone conduction sensor 280M to implement the voice function. The application processor can analyze the heart rate information based on the blood pressure pulsation signal acquired by the bone conduction sensor 280M to implement the heart rate detection function.

[0149] The keys 290 include a power-on key, volume keys, etc. The keys 290 can be mechanical keys. They can also be touch keys. The terminal 200 can receive key inputs and generate key signal inputs related to the user settings and function controls of the terminal 200.

[0150] The motor 291 can generate vibration prompts. The motor 291 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations for different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. For touch operations on different regions of the display screen 294, the motor 291 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving messages, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0151] The indicator 292 can be an indicator light and can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0152] The SIM card interface 295 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to achieve contact and separation from the terminal 200. The terminal 200 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 295 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 295 can also be compatible with different types of SIM cards. The SIM card interface 295 can also be compatible with external memory cards. The terminal 200 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the terminal 200 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the terminal 200 and cannot be separated from the terminal 200.

[0153] Figure 3 This is the flowchart of the first embodiment of the image transformation method of this application. As Figure 3 shown, the method of this embodiment can be applied to Figure 1 the application architecture shown, and its execution subject can be Figure 2 the terminal shown. The image transformation method can include:

[0154] Step 301: Obtain a first image of the target scene through the front camera.

[0155] The first image in this application is obtained by the front camera of the terminal in the scenario where the user takes a selfie. Usually, the FOV of the front camera is set to 70° - 110°, preferably 90°. The area directly opposite the front camera is the target scene, and this target scene includes the face of the target person (i.e., the user).

[0156] Optionally, the first image may be a preview image captured by the front camera of the terminal and displayed on the screen, at which time the shutter has not been triggered and the first image has not been formed on the image sensor; or, the first image may also be an image captured by the terminal but not displayed on the screen, at which time the shutter has not been triggered and the first image has not been formed on the sensor; or, the first image may further be an image captured by the terminal after the shutter is triggered and has been formed on the image sensor.

[0157] It should be noted that the first image may also be captured by the rear camera of the terminal, and no specific limitation is made thereto.

[0158] Step 302: Obtain the target distance between the face of the target person and the front camera.

[0159] The target distance between the face of the target person and the front camera may be the distance between the foremost part (such as the nose) on the face of the target person and the front camera; or, the target distance between the face of the target person and the front camera may also be the distance between a specified part (such as the eyes, mouth, or nose, etc.) on the face of the target person and the front camera; or, the target distance between the face of the target person and the front camera may further be the distance between the central position on the face of the target person (such as the nose in the frontal photo of the target person, or the zygomatic position in the profile photo of the target person, etc.) and the front camera. It should be noted that the definition of the above target distance may be determined according to the specific situation of the first image, and the present application does not make specific limitations thereto.

[0160] In a possible implementation manner, the terminal may obtain the above target distance by calculating the face screen occupation ratio, that is, first obtain the screen occupation ratio of the face of the target person in the first image (the ratio of the pixel area of the face to the pixel area of the first image), and then obtain the distance according to the screen occupation ratio and the field of view (FOV) of the front camera. Figure 4 An exemplary schematic diagram of the distance acquisition method is shown, as Figure 4 shown. Assuming that the length of the average face is 20 cm and the width is 15 cm, according to the screen occupation ratio P of the face, it can be estimated that at the target distance D between the face of the target person and the front camera, the true area S of the entire field of view of the first image is S = 20×15 / P cm 2 . According to the aspect ratio of the first image, the diagonal length L of the first image can be obtained. For example, when the aspect ratio of the first image is 1:1, the diagonal length is L = S 0.5 . According to the imaging relationship as shown in the above figure, the target distance D = L / (2×tan(0.5×FOV)).

[0161] In a possible implementation, the terminal can also obtain the above target distance through a distance sensor, that is, when the user takes a self-portrait, the distance between the front camera and the face of the target person in front can be measured through the distance sensor on the terminal. The distance sensor can include, for example, a time of flight (TOF) sensor, a structured light sensor, or a binocular sensor, etc.

[0162] Step 303: When the target distance is less than a preset threshold, perform a first processing on the first image to obtain a second image, where the first processing includes performing distortion correction on the first image according to the target distance.

[0163] Generally, when taking a self-portrait, the distance from the face to the front camera is relatively close, and there is a problem of perspective distortion of the face where "objects appear larger when closer and smaller when farther away". For example, when the distance between the face of the target person and the front camera is too small, due to the difference in the distances from different parts of the face to the camera, it may cause the nose in the image to appear larger and the face to be elongated, etc. When the distance between the face of the target person and the front camera is relatively large, the above problem may be weakened. Therefore, this application sets a threshold and believes that when the distance between the face of the target person and the front camera is less than this threshold, the first image containing the face of the target person is distorted and needs to be corrected for distortion. The value range of the preset threshold is within 80 centimeters. Optionally, the preset threshold can be set to 50 centimeters. It should be noted that the specific value of the preset threshold can be determined according to the performance of the front camera, the shooting light, etc. This application does not make specific limitations on this.

[0164] The terminal performs perspective distortion correction on the first image to obtain a second image. The face of the target person in the second image is closer to the true appearance of the face of the target person compared to the face of the target person in the first image, that is, the relative proportions and relative positions of the facial features of the target person in the second image are closer to the relative proportions and relative positions of the facial features of the face of the target person compared to the relative proportions and relative positions of the facial features of the target person in the first image.

[0165] As described above, when the first image is obtained, the distance between the face of the target person and the front camera is less than the preset threshold. It is very likely that due to the perspective distortion problem of the face of the front camera where "objects appear larger when closer and smaller when farther away", the sizes and stretches of the facial features of the face of the target person in the first image occur, and the relative proportions and relative positions of its facial features deviate from the relative proportions and relative positions of the true appearance of the face of the target person. The second image obtained by performing perspective distortion correction on the first image can eliminate the above situations where the sizes and stretches of the facial features of the face of the target person occur, so that the relative proportions and relative positions of the facial features of the face of the target person in the second image approach or even return to the relative proportions and relative positions of the true appearance of the face of the target person.

[0166] Optionally, the second image may be a preview image obtained by the front camera. That is, before the shutter is triggered, the front camera can obtain a first image facing the target scene area (the first image is displayed on the screen as a preview image, or the first image has never appeared on the screen). The terminal uses the above method to perform perspective distortion correction on the first image to obtain a second image, and displays the second image on the screen of the terminal. At this time, the second image seen by the user is a preview image that has been subjected to perspective distortion correction, the shutter has not been triggered, and the second image has not been imaged on the image sensor. Alternatively, the second image may also be an image obtained after the shutter is triggered. That is, after the shutter is triggered, the first image is an image imaged on the image sensor of the terminal. The terminal uses the above method to perform perspective distortion correction on the first image to obtain a second image, saves the second image, and displays it on the screen of the terminal. At this time, the second image seen by the user is an image in the picture library that has been subjected to perspective distortion correction and has been saved.

[0167] In this application, the process of obtaining the second image by performing perspective distortion correction on the first image may include: first, fitting the face of the target person in the first image with a standard face model according to the target distance to obtain the depth information of the face of the target person, and then performing perspective distortion correction on the first image according to the depth information to obtain the second image. The standard face model is a pre-created face model including facial features, and it is set with shape transformation coefficients, expression transformation coefficients, etc. By adjusting these coefficient values, the expression or shape of the face model can be changed. D represents the value of the target distance. The terminal assumes that the standard three-dimensional face model is placed directly in front of the camera, and the target distance between the standard three-dimensional face model and the camera is D. The specified points (such as the tip of the nose, the center point of the face model, etc.) on the standard three-dimensional face model correspond to the origin O of the three-dimensional coordinate system. Through perspective projection, a two-dimensional projection point set A of the feature points on the standard three-dimensional face model is obtained, and a point set B of the two-dimensional feature points of the face of the target person in the first image is obtained. Each point in the point set A has a corresponding unique matching point in the point set B, and the sum F of the two-dimensional coordinate distance differences of all the matching points is calculated. In order to obtain the true three-dimensional model of the face of the target person, the plane position (i.e., the above-mentioned specified point deviates from the origin O and moves up and down, left and right), shape, relative proportions of facial features, etc. of the standard three-dimensional face model can be adjusted multiple times to make the sum of the distance differences F reach the minimum or approach 0, that is, the point set A and the point set B approach complete one-to-one coincidence. Based on the above method, the true three-dimensional model of the face of the target person corresponding to the first image can be obtained. According to the coordinates (x, y, z) of any pixel point on the true three-dimensional model relative to the origin O, and then according to the target distance D between the origin O and the camera, the coordinates (x, y, z + D) of any pixel point on the true three-dimensional model of the face of the target person relative to the camera can be obtained. At this time, through perspective projection, the depth information of any pixel point of the face of the target person in the first image can be obtained. Optionally, a TOF sensor, a structured light sensor, etc. can also be used to directly obtain the depth information.

[0168] Based on the above fitting process, the terminal can establish a first three-dimensional model of the face of the target person. The above fitting process can adopt methods such as face feature point fitting, deep learning fitting, TOF camera depth fitting, etc. The first three-dimensional model can be presented in the form of a three-dimensional point cloud, which corresponds to the two-dimensional image of the face of the target person and presents the same facial expression, relative proportions and relative positions of facial features, etc.

[0169] Exemplarily, a three-dimensional model of the face of the target person is established by using the face feature point fitting method. Figure 5 An exemplary schematic diagram showing the creation process of the three-dimensional face model is shown, as Figure 5As shown, first, feature points of facial features such as facial features and contours are obtained from the input image. Then, with the help of a deformable basic 3D face model, according to the correspondence between the 3D key points of the face model and the above-mentioned obtained feature points, and according to the fitting parameters (target distance, FOV coefficient, etc.) and fitting optimization terms (facial rotation, translation, scaling parameters, face model deformation coefficients, etc.), the fitting of the real face shape is carried out. Since the corresponding positions of the contour feature points on the 3D face model change with the rotation angle of the face, the corresponding relationships at different rotation angles can be selected, and multiple iterative fittings can be used to ensure the fitting accuracy. Finally, the three-dimensional spatial coordinates of each point on the three-dimensional point cloud of the face that fits the target person are output. The distance from the camera to the nose, a 3D face model is established according to the distance, the standard model is placed at this distance, and projection is performed. The coordinate distance difference between the two (photo and two-dimensional projection) is minimized by continuous projection to obtain a 3D head, and the depth information is a vector. The fitting not only fits the shape but also fits the x-axis and y-axis.

[0170] The terminal transforms the pose and / or shape of the first 3D model of the face of the target person to obtain the second 3D model of the face of the target person.

[0171] Adjusting the pose and / or shape of the first 3D model needs to be adjusted according to the distance between the face of the target person in the first image and the front camera. Since the distance between the face of the target person in the first image and the front camera is less than the preset threshold, the first 3D model can be moved backward to obtain the second 3D model of the face of the target person. Based on the adjustment of the 3D model, the adjustment of the face of the target person in the real world can be simulated. Therefore, the adjusted second 3D model can present the result of the face of the target person moving backward compared to the front camera.

[0172] In a possible implementation, when the second 3D model is obtained by moving the position of the first 3D model, an angle compensation is performed on the second 3D model.

[0173] In the real world, when the face of the target person moves backward relative to the front camera, the face of the target person may change in angle relative to the camera. At this time, if it is not desired to retain the angle change caused by the backward movement, an angle compensation can be performed. Figure 6 An exemplary schematic diagram showing the angle change of the position movement is shown, such as Figure 6As shown, initially, the face of the target person is at a distance less than a preset threshold from the front camera, the angle relative to the front camera is α, the vertical distance from the front camera is tz1, and the horizontal distance from the front camera is tx. When the face of the target person moves backward, the angle relative to the front camera becomes β, the vertical distance from the front camera becomes tz2, and the horizontal distance from the front camera remains tx. Therefore, the change angle of the face of the target person relative to the front camera At this time, if only the first 3D model is moved backward, it will cause the angle change effect brought by the backward movement in the 2D image obtained based on the second 3D model. It is necessary to perform angle compensation on the second 3D model, that is, rotate the second 3D model by an angle of Δθ so that the finally obtained 2D image is the same as the corresponding 2D image before the backward movement (i.e., the face of the target person in the original image).

[0174] The terminal obtains the pixel displacement vector field of the target object according to the depth information, the first 3D model, and the second 3D model.

[0175] After obtaining the two 3D models before and after the transformation (the above-mentioned first 3D model and second 3D model), the present application performs perspective projection on the first 3D model according to the depth information to obtain a first coordinate set, and the first coordinate set includes two-dimensional coordinate values corresponding to the projection of multiple sampling points in the first 3D model. Perspective projection is performed on the second 3D model according to the depth information to obtain a second coordinate set, and the second coordinate set includes two-dimensional coordinate values corresponding to the projection of multiple sampling points in the second 3D model.

[0176] Projection is a method of transforming three-dimensional coordinates into two-dimensional coordinates. Common projection methods include orthogonal projection and perspective projection, etc. Taking perspective projection as an example, the basic perspective projection model consists of two parts: the viewpoint E and the view plane P, and the viewpoint E is not on the view plane P. The viewpoint E can be considered as the position of the camera. The view plane P is the two-dimensional plane for rendering the perspective view of the three-dimensional target object. Figure 7 An exemplary schematic diagram of perspective projection is shown, as Figure 7 As shown, for any point X in the real world, a ray starting from the viewpoint E and passing through the point X is constructed, and the intersection point Xp of this ray and the view plane P is the perspective projection of the point X. The objects in the three-dimensional world can be regarded as composed of a set of points {Xi}. In this way, rays Ri starting from the viewpoint E and passing through the points Xi are respectively constructed, and the set of intersection points of these rays Ri and the view plane P is the two-dimensional projection diagram of the objects in the three-dimensional world at the viewpoint E.

[0177] Based on the above principle, each sampling point of the three-dimensional model is projected respectively to obtain the corresponding pixel points on the two-dimensional plane. These pixel points on the two-dimensional plane can be represented by a coordinate value within the two-dimensional plane, and thus a coordinate set corresponding to the sampling point of the three-dimensional model can be obtained. In this application, a first coordinate set corresponding to the first three-dimensional model and a second coordinate set corresponding to the second three-dimensional model can be obtained.

[0178] Calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the face of the target person. The first coordinate value is the coordinate value corresponding to the first sampling point in the first coordinate set, the second coordinate value is the coordinate value corresponding to the first sampling point in the second coordinate set, and the first sampling point is any one of the multiple identical point clouds included in the first three-dimensional model and the second three-dimensional model.

[0179] The second three-dimensional model is obtained by performing a pose and / or shape transformation on the first three-dimensional model. Therefore, the two contain a large number of identical sampling points, and even the sampling points contained in the two are exactly the same. In this way, there will be many groups of coordinate values in the first coordinate set and the second coordinate set that correspond to the same sampling point, that is, a certain sampling point corresponds to a coordinate value in the first coordinate set and also corresponds to a coordinate value in the second coordinate set. Calculate the coordinate difference between the above first coordinate value and the above second coordinate value, that is, calculate the coordinate difference between the first coordinate value and the second coordinate value on the x-axis and the coordinate difference on the y-axis respectively to obtain the coordinate difference of the first pixel. By calculating the coordinate differences of all the identical sampling points included in the first three-dimensional model and the second three-dimensional model, the pixel displacement vector field of the face of the target person can be obtained. The pixel displacement vector field is composed of the coordinate differences of each sampling point.

[0180] In a possible implementation manner, when the coordinate difference between the coordinate value corresponding to the pixel at the edge position of the face of the target person in the image and the coordinate values of the pixels in the surrounding area is greater than a preset threshold, the coordinate values in the second coordinate set are adjusted by translation or scaling. The surrounding area is adjacent to the face of the target person.

[0181] In order to keep the size and position of the face of the target person consistent, when the displacement of the edge position of the face of the target person is too large (which can be measured by a preset threshold) resulting in background distortion, appropriate alignment points and scaling scales can be selected to adjust the coordinate values in the second coordinate set by translation or scaling. The principle of translation or scaling adjustment can be to make the displacement of the edge of the face of the target person relative to the surrounding area as small as possible. The surrounding area includes the background area or the edge area of the field of view. For example, when the lowest point of the first coordinate set is located at the boundary of the edge area of the field of view and the lowest point of the second coordinate set deviates too much from the boundary, translate the second coordinate set so that the lowest point coincides with the boundary of the edge area of the field of view. Figure 8a and Figure 8bExemplarily show the face projection effects at object distances of 30 cm and 55 cm respectively.

[0182] The terminal obtains the transformed image according to the pixel displacement vector field of the face of the target person.

[0183] In a possible implementation manner, algorithmic constraint correction is performed on the face of the target person, the field-of-view edge region, and the background region according to the pixel displacement vector field of the face of the target person to obtain the transformed image. The field-of-view edge region is a strip region located at the image edge, and the background region is other regions in the image except for the face of the target person and the field-of-view edge region.

[0184] The image is divided into three regions. One is the region occupied by the face of the target person, another is the edge region of the image (i.e., the field-of-view edge region), and the third is the background region (i.e., the background part outside the face of the target person, which does not include the edge region of the image).

[0185] This application can determine the initial image matrices corresponding to the face of the target person, the field-of-view edge region, and the background region respectively according to the pixel displacement vector field of the face of the target person; construct the constraint terms corresponding to the face of the target person, the field-of-view edge region, and the background region respectively, and construct the regularization constraint term for the image; according to the constraint terms and the regularization constraint term corresponding to the face of the target person, the field-of-view edge region, and the background region respectively, and the weight coefficients corresponding to each constraint term, obtain the pixel displacement matrices corresponding to the face of the target person, the field-of-view edge region, and the background region respectively; according to the initial image matrices corresponding to the face of the target person, the field-of-view edge region, and the background region respectively and the pixel displacement matrices corresponding to the face of the target person, the field-of-view edge region, and the background region respectively, obtain the transformed image through color mapping.

[0186] In a possible implementation manner, the pixel displacement vector field of the face of the target person is expanded through an interpolation algorithm to obtain the pixel displacement vector field of the mask region, and the mask region includes the face of the target person; algorithmic constraint correction is performed on the mask region, the field-of-view edge region, and the background region according to the pixel displacement vector field of the mask region to obtain the transformed image.

[0187] The interpolation algorithm may include assigning the pixel displacement vector of the first sampling point to the second sampling point as the pixel displacement vector of the second sampling point. The second sampling point is any sampling point located outside the face region of the target person and within the mask region, and the first sampling point is the pixel point on the boundary contour of the face of the target person that is closest to the second pixel point.

[0188] Figure 9 Show an exemplary schematic diagram of the pixel displacement vector expansion method, as Figure 9As shown, the target area is the human face, and the mask area is the area of the human head, which includes the human face. The pixel displacement vector field of the human face is expanded to the entire mask area using the above interpolation algorithm to obtain the pixel displacement vector field of the mask area.

[0189] The image is divided into four areas. One is the area occupied by the human face, and the other is a partial area related to the human face. This partial area can change accordingly with the pose and / or shape transformation of the target object. For example, when the human face rotates, the corresponding human head will definitely rotate as well. The human face and the above partial area form the mask area. The third is the edge area of the image (i.e., the edge area of the field of view), and the fourth is the background area (i.e., the background part outside the target object, which does not include the edge area of the image).

[0190] This application can determine the initial image matrices corresponding to the mask area, the field of view edge area, and the background area respectively according to the pixel displacement vector field of the mask area; construct the constraint terms corresponding to the mask area, the field of view edge area, and the background area respectively, and construct the regularization constraint term for the image; according to the constraint terms and the regularization constraint term corresponding to the mask area, the field of view edge area, and the background area respectively, and the weight coefficients corresponding to each constraint term, obtain the pixel displacement matrices corresponding to the mask area, the field of view edge area, and the background area respectively; according to the initial image matrices corresponding to the mask area, the field of view edge area, and the background area respectively and the pixel displacement matrices corresponding to the mask area, the field of view edge area, and the background area respectively, obtain the transformed image through color mapping.

[0191] The constraint terms for the mask area, the background area, and the field of view edge area respectively, as well as the regularization constraint term for the global image, are described below.

[0192] (1) The constraint term corresponding to the mask area is used to constrain the target image matrix corresponding to the mask area in the image to approximate the image matrix after geometric transformation using the pixel displacement vector field of the mask area in the previous step, so as to correct the distortion of the mask area. This geometric transformation represents a spatial mapping, that is, a pixel displacement vector field is mapped to another image matrix through transformation. The geometric transformation in this application can be at least one of image translation transformation (Translation), image scaling transformation (Scale), and image rotation transformation (Rotation).

[0193] For the convenience of description, the constraint term corresponding to the mask area can be simply referred to as the mask constraint term. When there are multiple target objects in the image, different target objects can correspond to different mask constraint terms.

[0194] The mask constraint term can be denoted as Term1, and the arithmetic expression of Term1 is as follows:

[0195] Term1(i,j) = SUM (i,j)∈HeadRegionk ||M0(i,j) + Dt(i,j) - Func1 k [M1(i,j)]||

[0196] Among them, for the image matrix M0(i,j) of the pixel points located in the head region (i.e., (i,j) ∈ HeadRegionk), the coordinate values after conformal transformation of the target object are M1(i,j) = [u1(i,j), v1(i,j)] T , Dt(i,j) represents the displacement matrix corresponding to M0(i,j), k represents the k-th mask region of the image, and Func1 k represents the geometric transformation function corresponding to the k-th mask region, and ||...|| represents the vector 2-norm.

[0197] The mask constraint term Term1(i,j) needs to ensure that under the action of the displacement matrix Dt(i,j), the image matrix M0(i,j) tends to be an appropriate geometric transformation of M1(i,j), including at least one transformation operation among image rotation, image translation, and image scaling.

[0198] The geometric transformation function Func1 corresponding to the k-th mask region k means that all points within the k-th mask region share the same geometric transformation function Func1 k , and different mask regions correspond to different geometric transformation functions. The geometric transformation function Func1 k can be specifically expressed as:

[0199]

[0200] Among them, ρ 1k represents the scaling coefficient of the k-th mask region, θ 1k represents the rotation angle of the k-th mask region, TX 1k and TY 1k respectively represent the horizontal displacement and vertical displacement of the k-th mask region.

[0201] Term1(i,j) can be specifically expressed as:

[0202]

[0203] Among them, du(i,j) and dv(i,j) are unknowns to be solved, and this term needs to be minimized as much as possible when solving the constraint equation later.

[0204] (2) The constraint term corresponding to the field of view edge region is used to constrain the pixel points in the initial image matrix corresponding to the field of view edge region in the image to displace along the edge of the image or towards the outside of the image, so as to maintain or expand the field of view edge region.

[0205] For ease of description, the constraint term corresponding to the field of view edge region can be abbreviated as the field of view edge constraint term. The field of view edge constraint term can be denoted as Term3, and the arithmetic expression of Term3 is as follows:

[0206] Term3(i,j) = SUM (i,j)∈EdgeRegion ||M0(i,j)+Dt(i,j)-Func3 (i,j) [M0(i,j)]||

[0207] Where, M0(i,j) represents the image coordinates of the pixel points located in the field of view edge region (i.e., (i,j) ∈ EdgeRegion), Dt(i,j) represents the displacement matrix corresponding to this M0(i,j), Func3 (i,j) represents the displacement function of M0(i,j), and ||...|| represents the vector 2-norm.

[0208] The field of view edge constraint term needs to ensure that under the action of the displacement matrix Dt(i,j), the image matrix M0(i,j) tends to have an appropriate displacement for the coordinate value M0(i,j). The displacement rule is to move only along the edge region or appropriately towards the outside of the edge region, avoiding moving towards the inside of the edge region. The advantage of doing this is that it can minimize the loss of image information caused by subsequent rectangular cropping, and even can gain and expand the image content of the field of view edge region.

[0209] Suppose a pixel point A located in the field of view edge region of the image, whose image coordinates are [u0, v0] T , the tangential vector of this point A along the field of view boundary is denoted as y(u0, v0), and the normal vector towards the outside of the image is denoted as x(u0, v0). When the boundary region is known, x(u0, v0) and y(u0, v0) are also known. Then Func3 (i,j) can be specifically expressed as:

[0210]

[0211] Where, α(u0(i,j), v0(i,j)) needs to be restricted to be not less than 0 to ensure that this point will not displace towards the inside of the field of view edge. The positive or negative of β(u0(i,j), v0(i,j)) does not need to be restricted. α(u0(i,j), v0(i,j)) and β(u0(i,j), v0(i,j)) are intermediate unknowns that do not need to be explicitly solved.

[0212] Term3(i,j) can be specifically expressed as:

[0213]

[0214] Among them, du(i,j) and dv(i,j) are unknowns to be solved, and this term needs to be minimized as much as possible when solving the constraint equations later.

[0215] (3) The constraint term corresponding to the background region is used to constrain the displacement of the pixel points in the image matrix corresponding to the background region in the image, and the first vector corresponding to the pixel point before displacement and the second vector corresponding to the pixel point after displacement are kept parallel as much as possible, so that the image content in the background region is smooth and continuous and the image content passing through the portrait in the background region is continuous and consistent in human visual perception; among them, the first vector represents the vector between the pixel point before displacement and the neighboring pixel points corresponding to the pixel point before displacement; the second vector represents the vector between the pixel point after displacement and the neighboring pixel points corresponding to the pixel point after displacement.

[0216] For the sake of convenience of description, the constraint term corresponding to the background region can be abbreviated as the background constraint term, and the background constraint term can be denoted as Term4. The arithmetic expression of Term4 is as follows:

[0217] Term4(i,j) = SUM (i,j)∈Bkg Re gion {Func4 (i,j) (M0(i,j), M0(i,j) + Dt(i,j))}

[0218] Among them, M0(i,j) represents the image coordinates of the pixel point located in the background region (i.e., (i,j) ∈ Bkg Region), Dt(i,j) represents the displacement matrix corresponding to this M0(i,j), Func4 (i,j) represents the displacement function of M0(i,j), and ||...|| represents the vector 2-norm.

[0219] The background constraint term needs to ensure that the coordinate value M0(i,j) under the action of the displacement matrix Dt(i,j) tends to be an appropriate displacement of the coordinate value M0(i,j). In this application, each pixel point in the background area can be divided into different control domains, and this application does not limit the size, shape, and number of the control domains. In particular, for the background pixel points located at the boundary between the target object and the background, their control domains need to extend across the target object to the other end of the target object. Suppose there is a certain background pixel point A and its set of control domain pixel points {Bi}. The control domain is the neighborhood of point A, and the control domain of A extends across the intermediate mask area to the other end of the mask area. Bi represents the neighborhood pixel points of A. After they are displaced, they are respectively moved to A′ and {B′i}. The displacement rule is that the background constraint term will limit the vectors ABi and A′B′i to keep their directions parallel as much as possible. The advantage of doing this is that it can ensure a smooth transition between the target object and the background area, and the image content passing through the human figure in the background area can be continuous and consistent in human vision, avoiding phenomena such as distortion or holes and wire drawing in the background image. Func4 (i,j) It can be specifically expressed as:

[0220]

[0221] Where:

[0222]

[0223]

[0224]

[0225] Among them, angle[] represents the included angle between two vectors, and vec1 represents the background point [i,j] before correction T and the vector formed by a certain point in its control domain, and vec2 represents the background point [i,j] after correction T and the vector formed by a certain point in its control domain after correction, SUM (i+di,j+dj)∈CtrlRegion represents the sum of the included angles of all vectors in the control domain.

[0226] Term4(i,j) can be specifically expressed as:

[0227] Term4(i,j) = SUM (i,j)∈BkgRegion {SUM (i+di,j+dj)∈CtrlRegion {angle[vec1,vec2]}}

[0228] Among them, du(i,j) and dv(i,j) are unknowns to be solved, and this term needs to be minimized as much as possible when solving the constraint equation later.

[0229] (4) The regularization constraint term is used to constrain that the difference between the displacement matrices of any two adjacent pixel points in the displacement matrices corresponding to the background region, the mask region, and the field-of-view edge region in the image is less than a preset threshold, so that the global image content of the image is smooth and continuous.

[0230] The regularization constraint term can be denoted as Term5, and the arithmetic expression of Term5 is as follows:

[0231] Term5(i,j) = SUM (i,j)∈AllRegion {Func5 (i,j) (Dt(i,j))}

[0232] For the entire image range (i.e., for the pixel point M0(i,j) of ( i ,j) ∈AllRegion ), the regularization constraint term needs to ensure that the displacement matrix Dt(i,j) of adjacent pixel points is smooth and continuous to avoid local excessive jumps. The limiting principle is that the difference between the displacement at the point [i,j] T and the displacement of its neighborhood point (i + di, j + dj) should be as small as possible (i.e., less than a certain threshold). Func5 (i,j) can be specifically expressed as:

[0233]

[0234] Term5(i,j) can be specifically expressed as:

[0235]

[0236] Among them, du(i,j) and dv(i,j) are unknowns to be solved, and this term needs to be ensured to be as small as possible when solving the constraint equation later.

[0237] According to each constraint term and the weight coefficient corresponding to each constraint term, the displacement matrix corresponding to each region is obtained.

[0238] Specifically, weight coefficients can be set for the constraint terms and the regularization constraint term of the mask region, the field-of-view edge region, and the background region, and a constraint equation is established according to each constraint term and the corresponding weight coefficient, and by solving this constraint equation, the offset of each position point in each region can be obtained.

[0239] Assume that the coordinate matrix of the image after algorithm constraint correction (also called the target image matrix) is Mt(i,j), Mt(i,j) = [ut(i,j), vt(i,j)] T , and its displacement matrix compared with the image matrix M0(i,j) is Dt(i,j), Dt(i,j) = [du(i,j), dv(i,j)] T , that is to say:

[0240] Mt(i,j) = M0(i,j) + Dt(i,j)

[0241] ut(i,j) = u0(i,j) + du(i,j)

[0242] vt(i,j) = v0(i,j) + dv(i,j)

[0243] Weighing coefficients are assigned to each constraint term, and the following constraint equation is constructed:

[0244] Dt(i,j) = (du(i,j), dv(i,j))

[0245] = arg min(α1(i,j) × Term1(i,j) + α2(i,j) × Term2(i,j) + α3(i,j) × Term3(i,j) + α4(i,j) × Term4(i,j) + α5(i,j) × Term5(i,j))

[0246] where α1(i,j) to α5(i,j) are the weighing coefficients (weighing matrix) corresponding to Term1 to Term5 respectively.

[0247] By using the least squares method or the gradient descent method or various improved algorithms to solve this constraint equation, the displacement matrix Dt(i,j) of each pixel point of the image is finally obtained. Based on this displacement matrix Dt(i,j), the transformed image can be obtained.

[0248] This application realizes the three-dimensional transformation effect of the image by means of a three-dimensional model, corrects the perspective distortion of the face of the target person in the image taken at close range, so that the relative proportions and relative positions of the facial features of the face of the target person after correction are closer to the relative proportions and relative positions of the facial features of the face of the target person, and can significantly improve the imaging effect in the selfie scenario.

[0249] In a possible implementation manner, for the scenario of recording a video, the terminal can adopt Figure 3 the method of the embodiment shown to perform distortion correction processing on multiple image frames in the recorded video respectively to obtain a video after distortion correction. The terminal can directly play the video after distortion correction on the screen, or the terminal can also display the video before distortion correction in a part of the area on the screen and the video after distortion correction in another part of the area in a split-screen manner, as shown in Figure 12f .

[0250] Figure 10 is the flowchart of the second embodiment of the image transformation method of this application. As shown in Figure 10 shown, the method of this embodiment can be applied to Figure 1 the application architecture shown, and its execution subject can beFigure 2 The terminal shown. The image transformation method may include:

[0251] Step 1001, obtain a first image.

[0252] In this application, the first image has been stored in the picture library of the second terminal. The first image may be a photo taken by the second terminal or a frame of a video taken by the second terminal. The application does not specifically limit the acquisition method of the first image.

[0253] It should be noted that the first terminal and the second terminal that currently perform distortion correction processing on the first image may be the same device or different devices.

[0254] In a possible implementation, the second terminal, as the device for acquiring the first image, may be any device with a shooting function, such as a camera, a video camera, etc. The second terminal stores the captured image locally or in the cloud. The first terminal, as the processing device for performing distortion correction on the first image, may be any device with image processing functions, such as a mobile phone, a computer, a tablet computer, etc. The first terminal may receive the first image from the second terminal or the cloud through wired or wireless communication, or the first terminal may obtain the first image captured by the second terminal through a storage medium (such as a USB flash drive).

[0255] In a possible implementation, the first terminal has both a shooting function and an image processing function, such as a mobile phone, a tablet computer, etc. The first terminal obtains the first image from the local picture library, or the first terminal captures and obtains the first image based on an instruction triggered by the shutter being pressed.

[0256] In a possible implementation, the first image includes the face of the target person. The distortion of the face of the target person in the first image is caused by the fact that the target distance between the face of the target person and the second terminal is less than the first preset threshold when the second terminal captures the first image. That is, when the second terminal captures the first image, the target distance between the face of the target person and the camera is small. Usually, when taking a self-portrait with a front camera, the distance between the face and the camera is small, and there is a problem of perspective distortion of the face where "the closer the object, the larger it appears; the farther the object, the smaller it appears". For example, when the target distance between the face of the target person and the front camera is too small, due to the difference in the distances from different parts of the face to the camera, it may cause the nose in the image to be larger and the face to be elongated. When the distance between the face of the target person and the front camera is large, the above problems may be weakened. Therefore, this application sets a threshold, believing that when the distance between the face of the target person and the front camera is less than this threshold, the first image containing the face of the target person is distorted and needs to be corrected for distortion. The value range of the preset threshold is within 80 centimeters. Optionally, the preset threshold can be set to 50 centimeters. It should be noted that the specific value of the preset threshold can be determined according to the performance of the front camera, the shooting light, etc., and this application does not make specific limitations on this.

[0257] The target distance between the face of the target person and the camera can be the distance between the foremost part (such as the nose) on the face of the target person and the camera; or, the target distance between the face of the target person and the camera can also be the distance between a specified part (such as the eyes, mouth, or nose, etc.) on the face of the target person and the camera; or, the target distance between the face of the target person and the camera can further be the distance between the central position on the face of the target person (such as the nose in the front view of the target person, or the cheekbone position in the side view of the target person, etc.) and the camera. It should be noted that the definition of the above target distance can be determined according to the specific situation of the first image, and this application does not make specific limitations on this.

[0258] The terminal can obtain the target distance based on the screen occupation ratio of the face of the target person in the first image and the FOV of the camera of the second terminal. It should be noted that when the second terminal includes multiple cameras, after shooting the first image, the information of the camera that shoots the first image will be recorded in the exchangeable image file format (EXIF) information. Therefore, the FOV of the second terminal mentioned above refers to the FOV of the camera recorded in the EXIF information. The principle can refer to step 302 above and will not be elaborated here. The FOV can be obtained from the FOV in the EXIF information of the first image, or calculated based on the equivalent focal length in the EXIF information. For example, fov = 2.0 × atan(43.27 / 2f), where 43.27 is the diagonal length of a 135mm film, and f represents the equivalent focal length. The terminal can also obtain the target distance based on the target shooting distance saved in the EXIF information of the first image. EXIF is specifically set for digital camera photos and can record the attribute information and shooting data of digital photos. The terminal can directly read data such as the target shooting distance, FOV, or equivalent focal length when the first image is taken from the EXIF information, and then obtain the above target distance. The principle can also refer to step 302 above and will not be elaborated here.

[0259] In a possible implementation, the first image includes the face of the target person. The distortion of the face of the target person in the first image is caused by the fact that when the second terminal shoots the first image, the FOV of the camera is greater than the second preset threshold, and the pixel distance between the face of the target person and the edge of the FOV is less than the third preset threshold. If the camera of the terminal is a wide-angle camera, then when the face of the target person is at the edge position of the FOV of the camera, it will also cause distortion, while when the face of the target person is at the middle position of the FVO of the camera, the distortion will be reduced or even disappear. Therefore, this application sets two thresholds and believes that when the FOV of the camera of the terminal is greater than the corresponding threshold and the pixel distance between the face of the target person and the edge of the FOV is less than the corresponding threshold, the first image containing the face of the target person has distortion and needs to be corrected for distortion. The threshold corresponding to the FOV is 90°, and the threshold corresponding to the pixel distance is one-fourth of the length or width of the first image. It should be noted that the specific value of the threshold can be determined according to the performance of the camera, shooting light, etc. This application does not make specific limitations on this.

[0260] The pixel distance can be the number of pixels between the foremost part of the face of the target person and the boundary of the first image; or, the pixel distance is the number of pixels between the specified part of the face of the target person and the boundary of the first image; or, the pixel distance is the number of pixels between the central position of the face of the target person and the boundary of the first image.

[0261] Step 1002: Display the distortion correction function menu.

[0262] When there is distortion in the face of the target person, the terminal displays a pop-up window, which is used to provide a selection control for whether to perform distortion correction. For example, Figure 13c as shown; or the terminal displays a distortion correction control, which is used to open the distortion correction function menu. For example, Figure 12d as shown. When the user clicks the "Yes" control or clicks the distortion correction control, in response to the instruction generated by the user operation, the distortion correction function menu is displayed.

[0263] This application provides a distortion correction function menu, which includes options for changing transformation parameters. For example, options for adjusting the equivalent simulated shooting distance, options for adjusting the displacement distance, options for adjusting the relative position and / or relative proportion of facial features, etc. These options can adjust the corresponding transformation parameters in the form of sliders or controls. That is, the transformation parameters can be changed by adjusting the position of a slider, or the selected value can be determined by triggering a control. When the distortion correction function menu is initially displayed, the values of the transformation parameters in each option on the menu can be default values or pre-calculated values. For example, the slider corresponding to the equivalent simulated shooting distance can initially be located at the position with a value of 0, as Figure 11 shown, or the terminal obtains an adjustment amount of the equivalent simulated shooting distance according to the image transformation algorithm, and displays the slider at the position corresponding to the value of this adjustment amount, as Figure 13e shown.

[0264] Figure 11 Shows an exemplary schematic diagram of the distortion correction function menu, as Figure 11As shown in the figure, the function menu includes four areas: the left part of the function menu is used to display the first image, which is placed under a coordinate axis including the xyz axes; the upper right part of the function menu includes two controls, one for saving pictures and the other for starting the transformation process of recording images. The corresponding operations can be enabled by triggering the controls; the lower right part of the function menu includes an expression control and an action control. Among them, the expression templates available for selection under the expression control include six expressions: unchanged, happy, sad, disgusted, surprised, and angry. When the user clicks on the control corresponding to the expression, it means that the expression template is selected. The terminal can transform the relative proportions and relative positions of the facial features of the face in the left part of the image according to the selected expression template. For example, for happy, usually the eyes on the human face will become smaller and the corners of the mouth will turn upwards; for surprised, usually both the eyes and the mouth will open wide; for angry, usually the eyebrows will frown and the corners of the eyes and mouth will turn downwards; the action templates available for selection under the action control include six actions: no action, nod, shake head left and right, shake head in Xinjiang dance style, blink, and laugh. When the user clicks on the control corresponding to the action, it means that the action template is selected. The terminal can transform the facial features of the face in the left part of the image multiple times in the order of the action according to the selected action template. For example, for nod, usually nodding includes two actions: lowering the head and raising the head; for shaking the head left and right, usually shaking the head left and right includes two actions: shaking the head to the left and shaking the head to the right; for blink, usually blinking includes two actions: closing the eyes and opening the eyes; the lower part of the function menu includes four sliders, corresponding to distance, shake head left and right, shake head up and down, and turn head clockwise respectively. Among them, by moving the position of the distance slider left and right, the distance between the face and the camera can be adjusted; by moving the position of the shake head left and right slider left and right, the direction (left or right) and angle of the face shaking the head can be adjusted; by moving the position of the shake head up and down slider left and right, the direction (up or down) and angle of the face shaking the head can be adjusted; by moving the position of the turn head clockwise slider left and right, the angle of the face turning clockwise or counterclockwise can be adjusted. The user can adjust the corresponding transformation parameter values by controlling one or more of the above sliders or controls. After detecting the user's operation, the terminal obtains the corresponding transformation parameter values.

[0265] Step 1003: Obtain the transformation parameters input by the user on the distortion correction function menu.

[0266] Obtain the transformation parameters according to the user's operations on the options included in the distortion correction function menu (for example, according to the user's dragging operation on the slider associated with the transformation parameter; and / or, according to the user's triggering operation on the control associated with the transformation parameter).

[0267] In a possible implementation, the distortion correction function menu includes an option to adjust the displacement distance; obtaining transformation parameters input by the user on the distortion correction function menu, including: obtaining the adjustment direction and displacement distance according to an instruction triggered by the user's operation on a control or slider in the option to adjust the displacement distance.

[0268] In a possible implementation, the distortion correction function menu includes an option to adjust the relative position and / or relative proportion of facial features; obtaining transformation parameters input by the user on the distortion correction function menu, including: obtaining the adjustment direction, displacement distance, and / or facial feature size according to an instruction triggered by the user's operation on a control or slider in the option to adjust the relative position and / or relative proportion of facial features.

[0269] In a possible implementation, the distortion correction function menu includes an option to adjust the angle; obtaining transformation parameters input by the user on the distortion correction function menu, including: obtaining the adjustment direction and adjustment angle according to an instruction triggered by the user's operation on a control or slider in the option to adjust the angle.

[0270] In a possible implementation, the distortion correction function menu includes an option to adjust the expression; obtaining transformation parameters input by the user on the distortion correction function menu, including: obtaining a new expression template according to an instruction triggered by the user's operation on a control or slider in the option to adjust the expression.

[0271] In a possible implementation, the distortion correction function menu includes an option to adjust the action; obtaining transformation parameters input by the user on the distortion correction function menu, including: obtaining a new action template according to an instruction triggered by the user's operation on a control or slider in the option to adjust the action.

[0272] Step 1004: Perform a first processing on the first image to obtain a second image, where the first processing includes performing distortion correction on the first image according to the transformation parameters.

[0273] The terminal performs perspective distortion correction on the first image according to the transformation parameters to obtain a second image. The face of the target person in the second image is closer to the true appearance of the face of the target person compared to the face of the target person in the first image, that is, the relative proportion and relative position of the facial features of the target person in the second image are closer to the relative proportion and relative position of the facial features of the target person's face compared to the relative proportion and relative position of the facial features of the target person in the first image.

[0274] By performing perspective distortion correction on the first image to obtain the second image, situations such as size changes and stretching of the facial features of the face of the above-mentioned target person can be eliminated. Therefore, the relative proportion and relative position of the facial features of the face of the target person in the second image approach, or even return to, the relative proportion and relative position of the true appearance of the face of the target person.

[0275] In a possible implementation, the terminal can also obtain a recording instruction according to the user's trigger operation on the start recording control, and then start recording the process of obtaining the second image. At this time, the start recording control becomes a stop recording control. When the user triggers the stop recording control, the terminal can receive a stop recording instruction and then stop recording.

[0276] The recording process in this application can include at least the following two scenarios:

[0277] One scenario is that for the first image, before the start of distortion correction, the user clicks the start recording control, and the terminal responds to the instruction generated by this operation to start the screen recording function, that is, record the screen of the terminal. At this time, the user clicks the control on the distortion correction function menu or drags the slider on the distortion correction function menu to set the transformation parameters. As the user operates, the terminal performs distortion correction processing or other transformation processing on the first image based on the transformation parameters, and then displays the processed second image on the screen. The process from the operation of the control or slider on the distortion correction function menu to the transformation process from the first image to the second image is recorded by the screen recording function of the terminal. When the user clicks the stop recording control, the terminal closes the screen recording function, and at this time, the terminal obtains the video of the above process.

[0278] Another scenario is that for a video, before the start of distortion correction, the user clicks the start recording control, and the terminal responds to the instruction generated by this operation to start the screen recording function, that is, record the screen of the terminal. At this time, the user clicks the control on the distortion correction function menu or drags the slider on the distortion correction function menu to set the transformation parameters. As the user operates, the terminal performs distortion correction processing or other transformation processing on multiple image frames in the video based on the transformation parameters, and then plays the processed video on the screen. The process from the operation of the control or slider on the distortion correction function menu to the process of playing the processed video is recorded by the screen recording function of the terminal. When the user clicks the stop recording control, the terminal closes the screen recording function, and at this time, the terminal obtains the video of the above process.

[0279] In a possible implementation, the terminal can also obtain a storage instruction according to the user's trigger operation on the save picture control, and then store the currently obtained image in the picture library.

[0280] In this application, the process of obtaining the second image by performing perspective distortion correction on the first image can refer to the description in step 303, with the difference that: in this application, in addition to being able to perform image transformation for the backward movement of the face of the target person, it can also perform image transformation for other transformations of the face of the target person, such as shaking the head, changing the expression, changing the action, etc. On this basis, in addition to being able to perform distortion correction on the first image, any transformation can also be performed on the first image. Even if there is no distortion in the first image, the face of the target person can be transformed according to the transformation parameters input by the user in the distortion correction function menu, so that the face of the target person in the transformed image is closer to the appearance under the selected transformation parameters.

[0281] In a possible implementation manner, in step 1001, in addition to obtaining the first image including the face of the target person, the annotation data of the first image can also be obtained, such as image segmentation mask, annotation box position, feature point position, etc. In step 1003, after obtaining the transformed second image, the annotation data can also be transformed according to the same transformation method as in step 1003 to obtain a new annotation file.

[0282] This can augment the picture library with annotated data, that is, generate various transformed images from one image. The transformed images do not need to be manually annotated again, and the effects of the transformed images are natural and real, with a large difference from the original image, which is beneficial for deep learning training.

[0283] In a possible implementation manner, in step 1003, the transformation method for the first three-dimensional model is specified, such as specifying to move the first three-dimensional model about 6 cm to the left or right (the distance between the left and right eyes of the face). After obtaining the transformed second image, the first image and the second image are respectively used as the inputs for the left and right eyes of the VR device, and the three-dimensional display effect of the target subject can be achieved.

[0284] This can convert ordinary two-dimensional images or videos into VR input sources and achieve the three-dimensional display effect of the face of the target person. Since the depth of the three-dimensional model itself is considered, the three-dimensional effect of the entire target object can be achieved.

[0285] This application uses a three-dimensional model to achieve the three-dimensional transformation effect of the image, performs perspective distortion correction on the face of the target person in the images taken at close range in the picture library, so that the relative proportions and relative positions of the facial features of the face of the target person after correction are closer to the relative proportions and relative positions of the facial features of the face of the target person, and the imaging effect can be changed. Further, corresponding transformations can be performed on the face of the target person according to the transformation parameters input by the user, so as to achieve diverse transformations of the face and realize the virtual transformation of the image.

[0286] It should be noted that in the above embodiments, a three-dimensional model is established for the face of the target person in the image, and then the distortion correction of the face is realized. The image transformation method provided by the present application can also be applied to the distortion correction of any other target object. The difference is that the object for establishing the three-dimensional model changes from the face of the target person to the target object. Therefore, the present application does not make specific limitations on the object of image transformation.

[0287] The following describes the image transformation method of the present application with two specific embodiments.

[0288] Embodiment 1 Figure 12a - Figure 12f Exemplarily shows the process of the terminal performing distortion correction in a self-timer scenario.

[0289] As Figure 12a shown, the user clicks on the camera icon on the desktop of the terminal to turn on the camera.

[0290] As Figure 12b shown, the camera by default turns on the rear camera, and the image of the target scene obtained by the rear camera is displayed on the screen of the terminal. The user clicks on the camera switch control on the camera function menu to switch the camera to the front camera.

[0291] As Figure 12c shown, the first image of the target scene obtained by the front camera is displayed on the screen of the terminal, and the target scene includes the face of the user.

[0292] When the distance between the face of the target person in the first image and the front camera is less than a preset threshold,

[0293] In one case, the terminal uses the image transformation method in the above Figure 3 shown embodiment to perform distortion correction on the first image to obtain a second image. Since the shutter is not triggered, the second image is displayed as a preview image on the screen of the terminal. As Figure 12d shown, at this time, the terminal displays the words "distortion correction" on the screen and also displays a close control. If the user does not need to display the corrected preview image, they can click on this close control. After the terminal receives the corresponding instruction, the first image is then displayed as the preview image on the screen.

[0294] In another case, as Figure 12e shown, the user clicks on the shutter control on the camera function menu, and the terminal uses the method in the above Figure 3 shown embodiment to perform distortion correction on the captured image (the first image) to obtain a second image. At this time, the second image is saved into the picture library.

[0295] In the third case, as Figure 12f shown, the terminal uses the above Figure 3In the method of the illustrated embodiment, the first image is subjected to distortion correction to obtain a second image. Since the shutter is not triggered, the terminal displays both the first image and the second image as preview images on the screen of the terminal, and the user can intuitively see the difference between the first image and the second image.

[0296] Embodiment 2 Figure 13a - Figure 13h Exemplarily shows the process of performing distortion correction on the images in the picture library.

[0297] As Figure 13a shown, the user clicks on the picture library icon on the desktop of the terminal to open the picture library.

[0298] As Figure 13b shown, the user selects the first image to be subjected to distortion correction in the picture library. When the first image was taken, the distance between the face of the target person therein and the camera was less than a preset threshold. The foregoing distance information can be obtained by the distance acquisition method in the method embodiment shown above, and will not be elaborated here. Figure 10

[0299] Figure 13c As shown, when the terminal detects that the face of the target person in the currently displayed image is distorted, it can pop up a window on the screen. The window displays the words "This image is distorted. Do you want to perform distortion correction?" and below this line, controls for "Yes" and "No" are displayed. When the user clicks "Yes", a distortion correction function menu is displayed. It should be noted that Figure 13c provides an example of the window for allowing the user to select whether to perform distortion correction, but this does not limit the interface or display content of the window. For example, the content of the words displayed on the window, the size of the words, the font of the words, the content on the two controls corresponding to "Yes" and "No", etc. can all be implemented in other ways, and the present application does not make specific limitations on the implementation manner of the window.

[0300] Figure 12d Figure 12d When the terminal detects that the face of the target person in the currently displayed image is distorted, it can also, as shown, display a control on the screen, and this control is used to open or close the distortion correction function menu. It should be noted that

[0301] Figure 11 Figure 14

[0302] or Figure 11 shown.

[0302] When the distortion correction function menu is initially displayed, the transformation parameter values in each option on the menu can be default values or pre-calculated values. For example, for the slider corresponding to distance, initially it can be located at the position with a value of 0, as Figure 11 shown, or the terminal obtains an adjustment amount for the distance according to the image transformation algorithm and displays the slider at the position corresponding to the value of this adjustment amount, as Figure 13d shown.

[0303] As Figure 13e shown, the user can adjust the distance by dragging the slider corresponding to the distance. As Figure 13f shown, the user can also trigger the control to select the expression to be transformed (happy). This process can refer to the method in the above Figure 10 shown embodiment and will not be elaborated here.

[0304] As Figure 13g shown, when the user is satisfied with the transformation result, the user can trigger the save picture control. When the terminal receives the instruction to save the picture, it stores the currently obtained image in the picture library.

[0305] As Figure 13h shown, before selecting the transformation parameters, the user can trigger the start recording control. After the terminal receives the start recording instruction, it starts screen recording and records the process of obtaining the second image. During this process, the start recording control becomes the stop recording control. When the user triggers the stop recording control, the terminal stops screen recording.

[0306] Embodiment 3 Figure 14 Exemplarily shows other examples of the distortion correction function menu.

[0307] As Figure 14 shown, the distortion correction function menu includes controls for selecting facial features and adjusting their positions. The user first selects the nose (the oval control in front of the nose is black), and then clicks the control representing enlargement. Each time the user clicks, the terminal will, according to the set step size, use the method in the above Figure 10 shown embodiment to enlarge the nose of the target person.

[0308] It should be noted that the Figure 11 , Figure 12a - Figure 12f , Figure 13a - Figure 13h , Figure 14 shown embodiments are all examples, but none of them constitute a limitation on the desktop of the terminal, the distortion correction function menu, the camera interface, etc. The present application does not make specific limitations on these.

[0309] Figure 15 is the structural schematic diagram of the embodiment of the image transformation device of the present application. As Figure 15 shown, the device in this embodiment can be applied to Figure 2The terminal shown. The image transformation module includes: an acquisition module 1501, a processing module 1502, a recording module 1503, and a display module 1504. Among them,

[0310] In a selfie scenario, the acquisition module 1501 is configured to obtain a first image for a target scenario through a front camera, where the target scenario includes the face of a target person; obtain a target distance between the face of the target person and the front camera; the processing module 1502 is configured to, when the target distance is less than a preset threshold, perform a first processing on the first image to obtain a second image; the first processing includes performing distortion correction on the first image according to the target distance; where the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0311] In a possible implementation, the target distance includes the distance between the most front-end part on the face of the target person and the front camera; or, the distance between a specified part on the face of the target person and the front camera; or, the distance between the central position on the face of the target person and the front camera.

[0312] In a possible implementation, the acquisition module 1501 is specifically configured to obtain the screen occupancy ratio of the face of the target person in the first image; obtain the target distance according to the screen occupancy ratio and the field of view angle FOV of the front camera.

[0313] In a possible implementation, the acquisition module 1501 is specifically configured to obtain the target distance through a distance sensor, and the distance sensor includes a time-of-flight ranging TOF sensor, a structured light sensor, or a binocular sensor.

[0314] In a possible implementation, the preset threshold is less than 80 cm.

[0315] In a possible implementation, the second image includes a preview image or an image obtained after triggering the shutter.

[0316] In a possible implementation, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image, including: the relative proportion of the facial features of the target person in the second image is closer to the relative proportion of the facial features of the face of the target person than the relative proportion of the facial features of the target person in the first image; and / or, the relative position of the facial features of the target person in the second image is closer to the relative position of the facial features of the face of the target person than the relative position of the facial features of the target person in the first image.

[0317] In a possible implementation, the processing module 1502 is specifically configured to fit the face of the target person in the first image with a standard face model according to the target distance to obtain the depth information of the face of the target person; and perform perspective distortion correction on the first image according to the depth information to obtain the second image.

[0318] In a possible implementation, the processing module 1502 is specifically configured to establish a first three-dimensional model of the face of the target person; perform transformation on the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the face of the target person; obtain a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the face of the target person.

[0319] In a possible implementation, the processing module 1502 is specifically configured to perform perspective projection on the first three-dimensional model according to the depth information to obtain a first coordinate set, where the first coordinate set includes coordinate values corresponding to a plurality of pixels in the first three-dimensional model; perform perspective projection on the second three-dimensional model according to the depth information to obtain a second coordinate set, where the second coordinate set includes coordinate values corresponding to a plurality of pixels in the second three-dimensional model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, where the first coordinate value includes the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel includes any one of a plurality of identical pixels included in the first three-dimensional model and the second three-dimensional model.

[0320] In a scenario of performing distortion correction on an image based on transformation parameters input by a user, the acquisition module 1501 is configured to acquire a first image, where the first image includes the face of a target person; wherein, the face of the target person in the first image has distortion; the display module 1504 is configured to display a distortion correction function menu; the acquisition module 1501 is further configured to acquire transformation parameters input by the user on the distortion correction function menu, where the transformation parameters at least include an equivalent simulated shooting distance, and the equivalent simulated shooting distance is used to simulate the distance between the face of the target person and the camera when the shooting terminal shoots the face of the target person; the processing module 1502 is configured to perform a first process on the first image to obtain a second image; the first process includes performing distortion correction on the first image according to the transformation parameters; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image.

[0321] In a possible implementation, the distortion of the face of the target person in the first image is caused by the target distance between the face of the target person and the second terminal being less than a first preset threshold when the second terminal captures the first image; wherein, the target distance includes the distance between the foremost part on the face of the target person and the front camera; or, the distance between a specified part on the face of the target person and the front camera; or, the distance between the center position on the face of the target person and the front camera.

[0322] In a possible implementation, the target distance is obtained based on the screen occupation ratio of the face of the target person in the first image and the FOV of the camera of the second terminal; or, the target distance is obtained based on the equivalent focal length in the EXIF information of the first image in the Exchangeable Image File Format.

[0323] In a possible implementation, the distortion correction function menu includes an option to adjust the equivalent simulated shooting distance; the obtaining module 1501 is specifically configured to obtain the equivalent simulated shooting distance according to an instruction triggered by a user's operation on a control or slider in the option to adjust the equivalent simulated shooting distance.

[0324] In a possible implementation, when the distortion correction function menu is initially displayed, the value of the equivalent simulated shooting distance in the option to adjust the equivalent simulated shooting distance includes a default value or a pre-calculated value.

[0325] In a possible implementation, the display module 1504 is further configured to, when the face of the target person is distorted, display a pop-up window, and the pop-up window is used to provide a selection control for whether to perform distortion correction; when the user clicks the control for performing distortion correction on the pop-up window, an instruction generated by responding to the user's operation is triggered.

[0326] In a possible implementation, the display module 1504 is further configured to, when the face of the target person is distorted, display a distortion correction control, and the distortion correction control is used to open the distortion correction function menu; when the user clicks the distortion correction control, an instruction generated by responding to the user's operation is triggered.

[0327] In a possible implementation, the distortion of the face of the target person in the first image is caused by the fact that when the second terminal captures the first image, the field of view (FOV) of the camera is greater than a second preset threshold, and the pixel distance between the face of the target person and the edge of the FOV is less than a third preset threshold; wherein the pixel distance includes the number of pixels between the foremost part of the face of the target person and the edge of the FOV; or, the number of pixels between a specified part of the face of the target person and the edge of the FOV; or, the number of pixels between the center position of the face of the target person and the edge of the FOV.

[0328] In a possible implementation, the FOV is obtained from the EXIF information of the first image.

[0329] In a possible implementation, the second preset threshold is 90°, and the third preset threshold is one quarter of the length or width of the first image.

[0330] In a possible implementation, the distortion correction function menu includes an option for adjusting the displacement distance; the obtaining module 1501 is further configured to obtain an adjustment direction and a displacement distance according to an instruction triggered by an operation of a control or a slider in the option for adjusting the displacement distance.

[0331] In a possible implementation, the distortion correction function menu includes an option for adjusting the relative position and / or relative proportion of facial features; the obtaining module 1501 is further configured to obtain an adjustment direction, a displacement distance, and / or facial feature sizes according to an instruction triggered by an operation of a control or a slider in the option for adjusting the relative position and / or relative proportion of facial features.

[0332] In a possible implementation, the distortion correction function menu includes an option for adjusting the angle; the obtaining module 1501 is further configured to obtain an adjustment direction and an adjustment angle according to an instruction triggered by an operation of a control or a slider in the option for adjusting the angle; or, the distortion correction function menu includes an option for adjusting the expression; the obtaining module 1501 is further configured to obtain a new expression template according to an instruction triggered by an operation of a control or a slider in the option for adjusting the expression; or, the distortion correction function menu includes an option for adjusting the action; the obtaining module 1501 is further configured to obtain a new action template according to an instruction triggered by an operation of a control or a slider in the option for adjusting the action.

[0333] In a possible implementation, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image, including: the relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the face of the target person than the relative proportions of the facial features of the target person in the first image; and / or, the relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the face of the target person than the relative positions of the facial features of the target person in the first image.

[0334] In a possible implementation, the processing module 1502 is specifically configured to fit the face of the target person in the first image with a standard face model according to the target distance to obtain the depth information of the face of the target person; and correct the perspective distortion of the first image according to the depth information and the transformation parameters to obtain the second image.

[0335] In a possible implementation, the processing module 1502 is specifically configured to establish a first three-dimensional model of the face of the target person; transform the pose and / or shape of the first three-dimensional model according to the transformation parameters to obtain a second three-dimensional model of the face of the target person; obtain a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the face of the target person.

[0336] In a possible implementation, the processing module 1502 is specifically configured to perform perspective projection on the first three-dimensional model according to the depth information to obtain a first coordinate set, where the first coordinate set includes coordinate values corresponding to a plurality of pixels in the first three-dimensional model; perform perspective projection on the second three-dimensional model according to the depth information to obtain a second coordinate set, where the second coordinate set includes coordinate values corresponding to a plurality of pixels in the second three-dimensional model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the target object, where the first coordinate value includes the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel includes any one of the plurality of same pixels included in the first three-dimensional model and the second three-dimensional model.

[0337] In a possible implementation, it further includes: a recording module 1503; the obtaining module is further configured to obtain a recording instruction according to a trigger operation of the user on a recording control; the recording module 1503 is configured to start recording the process of obtaining the second image according to the recording instruction until a stop recording instruction generated by a trigger operation of the user on a stop recording control is received.

[0338] In a possible implementation, the display module 1504 is configured to display a distortion correction function menu on the screen, the distortion correction function menu including one or more sliders and / or one or more controls; receive a distortion correction instruction, the distortion correction instruction including transformation parameters generated when the user performs a touch operation on the one or more sliders and / or the one or more controls, the transformation parameters at least including an equivalent simulated shooting distance, the equivalent simulated shooting distance being used to simulate the distance between the face of the target person and the camera when the shooting terminal shoots the face of the target person; the processing module 1502 is configured to perform a first processing on the first image according to the transformation parameters to obtain a second image; the first processing includes performing distortion correction on the first image; wherein, the face of the target person in the second image is closer to the real appearance of the face of the target person than the face of the target person in the first image.

[0339] The device in this embodiment can be used to execute Figure 3 or Figure 10 the technical solutions of the method embodiment shown, the implementation principles and technical effects of which are similar and will not be elaborated here.

[0340] In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware-encoded processor, or executed and completed by the combination of the hardware and software modules in the encoded processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0341] The memories mentioned in the above embodiments may be volatile memories or non-volatile memories, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memories of the systems and methods described herein are intended to include but not be limited to these and any other suitable types of memories.

[0342] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0343] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0344] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0345] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0346] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0347] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of the present application, or the part that contributes to the prior art, or this part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0348] As described above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image transformation method, characterized in that Including: Obtaining a first image of a target scene through a front camera, where the target scene includes the face of a target person; Obtaining a target distance between the face of the target person and the front camera; When the target distance is less than a preset threshold, performing a first processing on the first image to obtain a second image; the first processing includes performing distortion correction on the first image according to the target distance; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person compared to the face of the target person in the first image; Among them, the performing distortion correction on the first image according to the target distance includes: Fitting the face of the target person in the first image with a standard face model according to different rotation angles of the face of the target person and the target distance to obtain depth information of the face of the target person; Performing perspective distortion correction on the first image according to the depth information to obtain the second image.

2. The method according to claim 1, wherein The target distance includes the distance between the foremost part on the face of the target person and the front camera; or, the distance between a specified part on the face of the target person and the front camera; or, the distance between the center position on the face of the target person and the front camera.

3. The method according to claim 1 or 2, characterized in that The obtaining the target distance between the face of the target person and the front camera includes: Obtaining the screen occupancy ratio of the face of the target person in the first image; Obtaining the target distance according to the screen occupancy ratio and the field of view angle FOV of the front camera.

4. The method according to claim 1 or 2, characterized in that, The obtaining the target distance between the face of the target person and the front camera includes: Obtaining the target distance through a distance sensor, and the distance sensor includes a time-of-flight ranging TOF sensor, a structured light sensor or a binocular sensor.

5. The method according to any one of claims 1-4, characterized in that The preset threshold is less than 80 centimeters.

6. The method according to any one of claims 1-5, characterized in that, The second image includes a preview image or an image obtained after triggering the shutter.

7. The method according to any one of claims 1-6, characterized in that, The face of the target person in the second image is closer to the true appearance of the face of the target person compared to the face of the target person in the first image, including: The relative proportion of the facial features of the target person in the second image is closer to the relative proportion of the facial features of the face of the target person compared to the relative proportion of the facial features of the target person in the first image; and / or, The relative position of the facial features of the target person in the second image is closer to the relative position of the facial features of the face of the target person compared to the relative position of the facial features of the target person in the first image.

8. The method according to any one of claims 1-7, characterized in that, The performing perspective distortion correction on the first image according to the depth information to obtain the second image includes: Establishing a first three-dimensional model of the face of the target person; Performing transformation on the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the face of the target person; Obtaining a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model and the second three-dimensional model; Obtaining the second image according to the pixel displacement vector field of the face of the target person.

9. The method according to claim 8, characterized in that, Obtaining the pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model includes: Performing perspective projection on the first three-dimensional model according to the depth information to obtain a first coordinate set, where the first coordinate set includes coordinate values corresponding to multiple pixels in the first three-dimensional model; Performing perspective projection on the second three-dimensional model according to the depth information to obtain a second coordinate set, where the second coordinate set includes coordinate values corresponding to multiple pixels in the second three-dimensional model; Calculating the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the face of the target person, where the first coordinate value includes the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel includes any one of the multiple identical pixels included in the first three-dimensional model and the second three-dimensional model.

10. An image transformation device, characterized in that, Includes: An acquisition module for acquiring a first image of a target scene including the face of a target person through a front camera; Obtaining the target distance between the face of the target person and the front camera; A processing module for, when the target distance is less than a preset threshold, performing a first processing on the first image to obtain a second image; the first processing includes performing distortion correction on the first image according to the target distance; wherein, the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image; Specifically, the processing module is configured to fit the face of the target person in the first image with a standard face model according to different rotation angles of the face of the target person and the target distance to obtain the depth information of the face of the target person; and performing perspective distortion correction on the first image according to the depth information to obtain the second image.

11. The device according to claim 10, characterized in that, The target distance includes the distance between the frontmost part of the face of the target person and the front camera; or, the distance between a specified part of the face of the target person and the front camera; or, the distance between the center position of the face of the target person and the front camera.

12. The device according to claim 10 or 11, characterized in that, Specifically, the acquisition module is configured to obtain the screen occupancy ratio of the face of the target person in the first image; and obtain the target distance according to the screen occupancy ratio and the field of view angle FOV of the front camera.

13. The device according to claim 10 or 11, characterized in that, Specifically, the acquisition module is configured to obtain the target distance through a distance sensor, and the distance sensor includes a time-of-flight ranging TOF sensor, a structured light sensor, or a binocular sensor.

14. The device according to any one of claims 10-13, wherein The preset threshold is less than 80 centimeters.

15. The device according to any one of claims 10-14, characterized in that, The second image includes a preview image or an image obtained after triggering the shutter.

16. The device according to any one of claims 10-15, characterized in that, That the face of the target person in the second image is closer to the true appearance of the face of the target person than the face of the target person in the first image includes: The relative proportions of the facial features of the target person in the second image are closer to the relative proportions of the facial features of the target person's face than the relative proportions of the facial features of the target person in the first image; and / or, The relative positions of the facial features of the target person in the second image are closer to the relative positions of the facial features of the target person's face than the relative positions of the facial features of the target person in the first image.

17. The device according to any one of claims 10 - 16, characterized in that, The processing module is specifically configured to establish a first three-dimensional model of the face of the target person; transform the pose and / or shape of the first three-dimensional model to obtain a second three-dimensional model of the face of the target person; obtain a pixel displacement vector field of the face of the target person according to the depth information, the first three-dimensional model, and the second three-dimensional model; and obtain the second image according to the pixel displacement vector field of the face of the target person.

18. The device according to claim 17, characterized in that, The processing module is specifically configured to perform perspective projection on the first three-dimensional model according to the depth information to obtain a first coordinate set, where the first coordinate set includes coordinate values corresponding to a plurality of pixels in the first three-dimensional model; perform perspective projection on the second three-dimensional model according to the depth information to obtain a second coordinate set, where the second coordinate set includes coordinate values corresponding to a plurality of pixels in the second three-dimensional model; calculate the coordinate difference between the first coordinate value and the second coordinate value to obtain the pixel displacement vector field of the face of the target person, where the first coordinate value includes the coordinate value corresponding to a first pixel in the first coordinate set, the second coordinate value includes the coordinate value corresponding to the first pixel in the second coordinate set, and the first pixel includes any one of the plurality of identical pixels included in the first three-dimensional model and the second three-dimensional model.

19. A device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.

20. A computer-readable storage medium, characterized in that, Comprising a computer program, when the computer program is executed on a computer, the computer executes the method according to any one of claims 1-9.

21. A computer program product, characterized in that, The computer program product includes computer program code, when the computer program code runs on a computer or a processor, the computer or the processor executes the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Identification photo camera capable of performing human image posture photography prompting and human image posture detection method

    CN105046246A

  • Face distortion correction method and device, terminal equipment and storage medium

    CN111080545A