Image processing method and apparatus
By generating correction parameters that adapt to the shooting parameters to correct the facial image, the problem of facial distortion during selfies is solved, and the image accuracy and user experience are improved.
Patent Information
- Application Number
- PCT/CN2025/084448
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-24
- Publication Date
- 2025-10-02
AI Technical Summary
When users use terminal devices to take selfies at close range, facial images are prone to distortion, affecting the user experience.
The image processing device generates correction parameters based on the camera's shooting parameters and the face image, and directly corrects the face image, avoiding errors caused by local correction and improving correction accuracy.
Effectively eliminate the distortion of facial images, generate images that are closer to the user's actual appearance, and improve user experience.
Smart Images

Figure CN2025084448_02102025_PF_FP_ABST
Abstract
Description
Image processing method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on March 28, 2024, with application number 202410367295.X and application name “Image Processing Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of image processing, and in particular to an image processing method and apparatus. Background Art
[0003] With the development of terminal technology, the application scenarios of terminal devices are becoming more and more extensive. For example, users can use mobile phones to take selfies. However, due to the imaging characteristics of the camera, when users use terminal devices to take portraits at close range, the selfie images may have facial distortion issues, such as enlarged noses and elongated faces, which affects the user experience. Summary of the Invention
[0004] The present application provides an image processing method and apparatus, wherein the image processing apparatus corrects a face image in an image to overcome lens distortion.
[0005] In a first aspect, the present application provides an image processing method. The method includes: obtaining at least one first correction parameter based on a first image and a first shooting parameter, wherein the first image is captured by a camera based on the first shooting parameter, the first image includes a facial image, the at least one first correction parameter corresponds one-to-one with at least one pixel in the first image, and the at least one first correction parameter is associated with the shooting distance between the camera and the face. Based on the at least one first correction parameter, a distorted portion of the facial image in the first image is corrected to obtain a first corrected image.
[0006] Thus, the image processing device of the present application obtains correction parameters for correcting the pixels of the facial image based on the first image containing the facial image and the shooting parameters at the time of shooting. By introducing the shooting parameters, the generated correction parameters can be adapted to the shooting scene corresponding to the current shooting parameters, thereby improving the accuracy of the correction of the image obtained under the current shooting parameters. Furthermore, the correction parameters are generated corresponding to the first image and correspond to the pixels in the first image, and can correct the entire first image, avoiding correction errors caused by local correction, further improving the correction accuracy, and effectively eliminating the distortion problem of the facial image, so that the facial image is closer to the user's actual appearance and improving the user experience.
[0007] Illustratively, the number of correction parameters is less than or equal to the number of pixels in the image.
[0008] In one possible implementation, based on a first image and a first shooting parameter, obtaining at least one first correction parameter includes: inputting the first image and the first shooting parameter into an AI network, and obtaining a first correction grid output by the AI network; wherein at least one grid point in the first correction grid corresponds one-to-one to at least one pixel point in the first image, and each grid point in the first correction grid corresponds to a first correction parameter. In this way, the electronic device or apparatus in the embodiment of the present application may pre-store an AI network, which may take an image and shooting parameters as input and output a correction grid adapted to the image and shooting parameters. In actual applications, the AI network may output a correction grid and corresponding correction parameters based on different images and shooting parameters, thereby improving the accuracy of image correction of the image. In addition, the AI network in this application takes the entire image as input and generates a correction grid corresponding to the entire image. The existing technology requires detecting facial feature points and using facial feature points and / or face detection frames as input to the AI network. This application directly uses facial images or downsampled images of facial images as input to the AI network. There is no need to input facial feature points or face detection frames, and there is no need to perform operations such as face detection. This can improve the accuracy of the generated correction parameters while reducing processing complexity, effectively improving the efficiency of image correction.
[0009] In one possible implementation, based on at least one first correction parameter, a distorted portion of a facial image in a first image is corrected to obtain a first corrected image. The method includes: obtaining a first correction degree parameter; correcting the distorted portion of the facial image in the first image based on the at least one first correction parameter to obtain the first corrected image, including: updating the first correction parameter corresponding to each grid point in a first correction grid based on the first correction degree parameter; and correcting the distorted portion of the facial image in the first image based on the updated first correction grid to obtain the first corrected image. In this way, in embodiments of the present application, the degree of correction of the image by the correction grid can be adjusted by setting the correction degree parameter, so that by setting different correction degree parameters, dynamic adjustment of the degree of correction of the facial image by the correction parameters can be achieved.
[0010] In one possible implementation, obtaining the first correction parameter includes: obtaining the first correction parameter in response to a received user operation; or obtaining the first correction parameter based on the deflection angle of the facial image in the first image. In this way, the present application can provide a user interface that allows users to manually adjust the correction level based on their personal needs and preferences, thereby meeting their personalized correction requirements. Furthermore, the electronic device can automatically adjust the correction level based on the facial deflection angle to accommodate the correction requirements of different facial capture scenarios.
[0011] In one possible implementation, obtaining a first correction level parameter in response to a received user operation includes: displaying a correction level parameter adjustment option in a user interface; determining the first correction level parameter in response to a received adjustment operation on the correction level parameter adjustment option; and displaying the first correction level parameter in the user interface. The method further includes displaying a first corrected image in the user interface. Thus, by providing a user interface, the present application enables users to adjust the correction intensity based on their personal needs and preferences. Furthermore, by providing a user interaction interface, interactivity is improved, enabling personalized customization of correction scenarios to better adapt to different scenarios and usage habits. Furthermore, in embodiments of the present application, an adjustment option is provided, allowing users to adjust the correction level with simple operations, providing a convenient operation method. Furthermore, by displaying the correction level parameter and the corrected image in the user interface, the present application enables users to intuitively understand the correlation between the current correction level of the correction level parameter and the corrected image (also understood as the correction effect). Based on the adjusted correction effect, users can more accurately adjust the correction level parameter to achieve a satisfactory visual effect and experience.
[0012] In one possible implementation, the AI network is trained based on multiple training images containing facial images and corresponding shooting parameters, where the multiple training images containing facial images correspond to different shooting distances and / or different shooting parameters. Thus, by training the AI network based on different shooting distances and shooting parameters as training datasets during the training process, the AI network can generate a suitable correction grid based on images obtained at different shooting distances and corresponding shooting parameters during actual application, thereby improving the accuracy of the correction grid in image correction.
[0013] Exemplarily, the shooting parameter may be a FOV parameter. Different FOV parameters may correspond to different aspect ratios of the image. At the same shooting distance, the FOV parameters of the image are different, and the output correction grid and its corresponding correction parameters are also different. For example, at the same shooting distance, the larger the image FOV parameter, the more content in the picture, and the smaller the proportion of the image occupied by the face. The larger the image FOV parameter, the less content in the picture, and the larger the proportion of the image occupied by the face. In an embodiment of the present application, the shooting distance is associated with the degree of distortion, that is, when the shooting distance is the same, the degree of distortion of the face is the same. However, when the FOV is different, that is, the proportion of the face image in the image is different, the position of the distortion may change, and accordingly, the pixel positions corresponding to the correction parameters of the same value will also be different.
[0014] In one possible implementation, correcting distorted portions of a facial image in a first image based on at least one first correction parameter to obtain a first corrected image includes: obtaining facial motion information based on the facial image; updating the correction parameters corresponding to the grid points in the first correction grid based on the facial motion information and the correction parameters corresponding to the grid points in a second correction grid; wherein the second correction grid is obtained based on the shooting parameters corresponding to the previous frame and the previous frame of the first image; and correcting the distorted portions of the facial image in the first image based on the updated first correction grid to obtain the first corrected image. In this manner, the electronic device smoothes the correction parameters of the current correction grid using the previous correction grid based on the sequence of video image frames and the temporal relationship between image frames, thereby improving the consistency and stability of the correction grids through temporal smoothing techniques. For example, after obtaining the correction grids, the electronic device may perform a weighted average of the correction parameters of each corresponding grid point in the two correction grids based on the previous correction grid to update the correction parameters of each grid point in the first correction grid. The electronic device can correct the first image based on the updated first correction grid. The electronic device can obtain motion information such as the speed or displacement of the face, and determine the corresponding smoothing amplitude based on the speed or displacement of the face, and the smoothing amplitude can optionally be a weight value in the weighted average. In an embodiment of the present application, the faster the face moves, that is, the greater the displacement of the face image between the two image frames, the smaller the smoothing amplitude, that is, the smaller the weight value of the correction grid corresponding to the previous image (for example, 0.1), and the larger the weight value of the correction grid corresponding to the current image (for example, 0.9). Conversely, if the face moves slower, that is, the smaller the displacement of the face image between the two image frames, the smoothing amplitude is larger, that is, the larger the weight value of the correction grid corresponding to the previous image (for example, 0.9), and the smaller the weight value of the correction grid corresponding to the current image (for example, 0.1). Through this time domain smoothing technology, the coherence and stability between the correction grids can be effectively improved, thereby improving the coherence and stability between the corrected images.
[0015] In one possible implementation, based on at least one first correction parameter, the distorted portion of the facial image in the first image is corrected. Before obtaining the first corrected image, the method further includes: obtaining the depth value of each pixel in the first image; based on the depth value, determining whether the shooting scene corresponding to the facial image is a planar image; based on the judgment result, determining whether to correct the distorted portion of the facial image in the first image based on at least one first correction parameter. In this way, the present application can further determine whether the facial image is distorted by determining whether the shooting scene corresponding to the facial image corresponds to a planar image. In the case where the shooting scene corresponding to the facial image is a planar image, the image processing device can determine that the facial image is not distorted, and then determine that there is no need to correct the facial image, thereby avoiding the overhead caused by unnecessary correction.
[0016] In one possible implementation, the first image corresponds to a first shooting distance; the method further includes: obtaining at least one second correction parameter based on a second image and a first shooting parameter, wherein the second image is captured by a camera based on the first shooting parameter, the second image includes a facial image, the second image corresponds to a second shooting distance, the second shooting distance is less than the first shooting distance, the at least one second correction parameter corresponds one-to-one with at least one pixel in the second image, the degree of distortion of the facial image in the second image is greater than the degree of distortion of the facial image in the first image, and the second correction parameter corresponding to the distorted portion of the facial image in the second image is greater than the first correction parameter corresponding to the distorted portion of the facial image in the first image; and correcting the distorted portion of the facial image in the second image based on the at least one second correction parameter to obtain a second corrected image. In this way, in the embodiment of the present application, corresponding correction parameters can be generated for facial images with different degrees of distortion, and the distorted portion of the facial image is corrected using the correction parameters to obtain an image of the facial image that is close to the user's actual appearance, effectively overcoming the facial distortion problem caused by close-range shooting, improving visual effects, and enhancing user experience.
[0017] In one possible implementation, the first correction parameter indicates the corrective displacement of the corresponding pixel in the first image. Thus, the electronic device can use the correction parameter to correct the pixels that cause facial distortion, deforming the pixels to an appropriate position based on the corrective displacement, thereby eliminating facial distortion in the image and improving the visual quality.
[0018] In one possible implementation, the first shooting parameter includes a field of view (FOV) parameter. Thus, by referencing the FOV parameter, embodiments of the present application can be adapted to correct facial distortion in different shooting scenarios. That is, the degree of distortion of a facial image may vary under the influence of different FOV parameters. The correction parameters generated in this application based on the FOV parameter can effectively improve the accuracy of facial image correction.
[0019] In a possible implementation, the first image is a preview image captured in real time by a camera. In this way, the image processing device can correct each preview image in real time while the user is previewing the image, thereby improving the user experience.
[0020] In one possible implementation, the method further includes: correcting a thumbnail corresponding to the preview image based on at least one first correction parameter to obtain a corrected thumbnail; and / or correcting an image stored in an album corresponding to the preview image based on the at least one first correction parameter to obtain a corrected saved image. In this way, the image processing device corrects the thumbnail and saved image based on the correction parameters corresponding to the preview image, avoiding repeated generation of corrected images and improving image correction processing efficiency.
[0021] Exemplarily, the image processing device can also store the correspondence between the correction grid and the timestamp. The timestamp represents the acquisition time of the first image, which can also be understood as the generation time of the correction grid. In this way, the image processing device can find the corresponding correction grid based on the timestamp of the thumbnail and / or saved image to achieve time alignment, thus avoiding errors in the correction result caused by inconsistencies between the correction grid and the image to be corrected (including the thumbnail and / or saved image).
[0022] In one possible implementation, obtaining at least one first correction parameter based on the first image and the first shooting parameters includes: downsampling the first image; and obtaining the at least one first correction parameter based on the downsampled first image and the first shooting parameters. In this way, the image processing device downsamples the image captured by the camera before inputting it into the AI network, reducing the processing complexity of the AI network and improving the efficiency of the AI network in generating the correction grid.
[0023] Exemplarily, the resolution of the downsampled image is lower than the resolution of the first image captured by the camera, and accordingly, the generated correction grid corresponds to the downsampled image. The number of dots in the correction grid is the same as the number of pixels in the downsampled image, both of which are lower than the number of pixels in the first image. The electronic device may interpolate the correction grid to obtain a correction grid with the same number of dots as the first image, thereby correcting each pixel in the first image based on the correction parameters corresponding to each dot.
[0024] In second aspect, the present application provides an image processing device, including: a first acquisition module, used to acquire at least one first correction parameter based on a first image and a first shooting parameter, wherein the first image is captured by a camera based on the first shooting parameter, the first image includes a face image, at least one first correction parameter corresponds one-to-one to at least one pixel point in the first image, and the first correction parameter is associated with the shooting distance between the camera and the face; a correction module, used to correct the distorted part of the face image in the first image based on at least one first correction parameter to obtain a first corrected image.
[0025] In one possible implementation, the acquisition module is specifically configured to: input the first image and the first shooting parameter into the AI network, and obtain a first correction grid output by the AI network; wherein at least one grid point in the first correction grid corresponds one-to-one to at least one pixel point in the first image, and each grid point in the first correction grid corresponds to a first correction parameter.
[0026] In one possible implementation, the device also includes: a second acquisition module, used to obtain a first correction degree parameter; a correction module, specifically used to: update the first correction parameter corresponding to each grid point in the first correction grid based on the first correction degree parameter; based on the updated first correction grid, correct the distorted part of the facial image in the first image to obtain a first corrected image.
[0027] In a possible implementation, the second acquisition module is specifically configured to: acquire the first correction degree parameter in response to a received user operation; or acquire the first correction degree parameter based on a deflection angle of the face image in the first image.
[0028] In one possible implementation, the second acquisition module is specifically used to: determine the first correction degree parameter in response to a received adjustment operation on the correction degree parameter adjustment option in the user interface; display the first correction degree parameter in the user interface; the device also includes: a display module for displaying the first corrected image in the user interface.
[0029] In one possible implementation, the AI network is trained based on multiple training images containing facial images and corresponding shooting parameters, wherein the multiple training images containing facial images have different shooting distances and corresponding shooting parameters.
[0030] In one possible implementation, the correction module is specifically configured to: obtain facial motion information based on a facial image; update first correction parameters of the points in a first correction grid based on the facial motion information and the correction parameters corresponding to the points in a second correction grid; wherein the second correction grid is obtained based on a previous frame image adjacent to the first image and shooting parameters corresponding to the previous frame image; and correct the distorted portion of the facial image in the first image based on the updated first correction grid to obtain a first corrected image.
[0031] In one possible implementation, the correction module is further used to: obtain a depth value of each pixel in the first image; based on the depth value, determine whether the shooting scene corresponding to the facial image is a planar image; based on the judgment result, determine whether to correct the distorted part of the facial image in the first image based on at least one first correction parameter.
[0032] In one possible implementation, the shooting distance corresponding to the first image is the first shooting distance; the first acquisition module is further used to obtain at least one second correction parameter based on the second image and the first shooting parameter, wherein the second image is captured by the camera based on the first shooting parameter, the second image includes a face image, the shooting distance corresponding to the second image is the second shooting distance, the second shooting distance is smaller than the first shooting distance, at least one second correction parameter corresponds one-to-one to at least one pixel point in the second image, the degree of distortion of the face image of the second image is greater than the degree of distortion of the face image of the first image, and the second correction parameter corresponding to the distorted part of the face image of the second image is greater than the first correction parameter corresponding to the distorted part of the face image of the first image; the correction module is further used to correct the distorted part of the face image in the second image based on the at least one second correction parameter to obtain a second corrected image.
[0033] In a possible implementation, the first correction parameter is used to indicate a correction displacement of a corresponding pixel point in the first image.
[0034] In a possible implementation, the first shooting parameter includes a field of view FOV parameter.
[0035] In a possible implementation, the first image is a preview image captured in real time by a camera.
[0036] In one possible implementation, the correction module is further used to: correct the thumbnail corresponding to the preview image based on at least one first correction parameter to obtain a corrected thumbnail; and / or correct the image saved in the album corresponding to the preview image based on at least one first correction parameter to obtain a corrected saved image.
[0037] In a possible implementation, the acquisition module is further configured to: perform downsampling processing on the first image; and acquire the at least one first correction parameter based on the downsampled first image and the first shooting parameter.
[0038] In a third aspect, the present application provides a graphical user interface on a computer device, characterized in that the computer device has a display screen, a camera, a memory, and one or more processors that execute one or more instructions stored in the memory, wherein: in response to a first operation received, an image preview interface is displayed on the display screen; a first corrected image is displayed in the image preview interface; wherein the first corrected image is obtained by correcting the distorted part of the facial image in the first image based on at least one first correction parameter; the first image is captured by the camera based on the first shooting parameter, the first image includes a facial image, the at least one first correction parameter corresponds one-to-one to at least one pixel point in the first image, the first correction parameter is obtained based on the first image and the first shooting parameter, and the first correction parameter is associated with the shooting distance between the camera and the face.
[0039] In one possible implementation, the image preview interface also includes a correction degree parameter adjustment option, wherein: in response to receiving a second operation on the correction degree parameter adjustment option, a first correction degree parameter is displayed in the image preview interface, and the first correction degree parameter is used to indicate the correction degree of at least one correction parameter on the distorted part of the facial image.
[0040] In a fourth aspect, the present application provides an image processing apparatus. The apparatus comprises: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory and, when executed by the one or more processors, cause the apparatus to perform instructions of the method of the first aspect or any possible implementation of the first aspect.
[0041] In a fifth aspect, an embodiment of the present application provides a computer-readable medium for storing a computer program, wherein the computer program includes instructions for executing the method in the first aspect or any possible implementation of the first aspect.
[0042] In a sixth aspect, an embodiment of the present application provides a computer program comprising instructions for executing the method in the first aspect or any possible implementation of the first aspect.
[0043] In a seventh aspect, embodiments of the present application provide a chip comprising a processing circuit and transceiver pins. The transceiver pins and the processing circuit communicate with each other via an internal connection path, and the processing circuit executes the method of the first aspect or any possible implementation of the first aspect to control the receive pin to receive a signal and to control the transmit pin to send a signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIG1 is a schematic diagram showing a hardware structure of an electronic device;
[0045] FIG2 is a schematic structural diagram of an exemplary camera module;
[0046] FIG3 is a schematic diagram illustrating a software structure of an electronic device;
[0047] FIG4 is a schematic diagram showing an exemplary software structure of a camera;
[0048] FIG5 is a schematic diagram of an exemplary user interface;
[0049] FIG6 is a schematic diagram of an exemplary photographic image;
[0050] FIG7 is a flow chart showing an exemplary image processing method;
[0051] FIG8 is a schematic diagram of FOV shown as an example;
[0052] FIG9 is a schematic diagram illustrating an exemplary image correction process;
[0053] FIG10 a is a schematic diagram of an exemplary correction grid;
[0054] 10b to 10e are schematic diagrams of exemplary user interfaces;
[0055] FIG11 is a schematic diagram illustrating an exemplary image correction;
[0056] FIG12 is a schematic diagram of an exemplary user interface;
[0057] FIG13 is a schematic diagram showing an exemplary display of a saved image;
[0058] FIG14 is a schematic diagram showing the structure of an exemplary image processing device;
[0059] FIG15 is a schematic diagram showing the structure of an exemplary image processing device. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0061] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0062] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.
[0063] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0064] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. For example, "multiple processing units" means two or more processing units; "multiple systems" means two or more systems.
[0065] FIG1 shows a schematic diagram of the structure of an electronic device 100. It should be understood that the electronic device 100 shown in FIG1 is only an example of an electronic device, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in FIG1 may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits. It should be noted that in the embodiment of the present application, only the electronic device 100 is described as a mobile phone. In other embodiments, the electronic device 100 may also be a tablet, wearable device, smart home device (such as a smart TV, smart door lock), vehicle-mounted device, monitoring device, or other device with a camera function, which is not limited in the present application.
[0066] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0067] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0068] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0069] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0070] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0071] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.
[0072] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0073] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0074] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0075] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0076] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0077] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0078] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0079] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0080] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0081] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU (Graphics Processing Unit), NPU (Neural Network Processing Unit), display screen 194 and application processor.
[0082] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0083] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0084] Exemplarily, the camera 193 may be located in the edge area of the electronic device, may be an under-screen camera, or may be a liftable camera. The camera 193 may include a rear camera or a front camera. The embodiment of the present application does not limit the specific position and form of the camera 193. The electronic device 100 may include cameras with one or more focal lengths. For example, cameras with different focal lengths may include a telephoto camera, a wide-angle camera, an ultra-wide-angle camera, or a panoramic camera.
[0085] FIG2 is a schematic diagram of an exemplary camera module (also referred to as a camera module; in the embodiment of the present application, the camera and the camera module may be interchangeable and will not be described again herein). Referring to FIG2 , FIG2 (a) and (b) schematically illustrate the front 102 and back 103 of the electronic device 100, respectively. The front 102 of the electronic device 100 may be understood as the side facing the user when the user is using the electronic device 100, and the back 103 of the electronic device 100 may be understood as the side facing away from the user when the user is using the electronic device 100.
[0086] The camera module 104 is used to capture still images or videos. The camera module 104 can be set on the front and / or back of the electronic device 100. As shown in Figure 2 (a), when the camera module 104 is set on the front of the electronic device 100, the front camera 104-1 can be used to shoot the scene on the front side of the electronic device 100, such as for selfies, and in some embodiments it can be called a front camera. It should be noted that the layout of the cameras shown in Figure 2 (a) (such as horizontal arrangement and spacing) is only an illustrative example and is not limited in this application.
[0087] When the camera module 104 is set on the back of the electronic device 100, the rear camera 104-2 can be used to shoot the scene on the back side of the electronic device 100, and in some embodiments, it can be called a rear camera. As shown in Figure 2 (b), the rear camera of the mobile phone includes 4 cameras, which can be regarded as a rear camera module or as 4 separate cameras. Among them, the 4 cameras can include but are not limited to: wide-angle camera, ultra-wide-angle camera, panoramic camera, etc., which are not limited in this application. When shooting, the user can select the corresponding camera module according to the shooting needs.
[0088] It should be noted that the embodiment of the present application does not limit the number of camera modules 104 provided, and can be one, two, four, or even more. For example, one or more camera modules 104 can be provided on the front of the electronic device 100, and / or one or more camera modules 104 can be provided on the back of the electronic device 100. When multiple camera modules 104 are provided, the multiple camera modules 104 can be completely identical or different, for example, the multiple camera modules 104 have different lens optical parameters, different lens installation positions, different lens shapes, etc. The embodiment of the present application also does not impose any restrictions on the relative positions of the multiple camera modules when they are provided.
[0089] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0090] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0091] The NPU is a neural network computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, voice recognition, and text comprehension.
[0092] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0093] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0094] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0095] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0096] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0097] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0098] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.
[0099] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0100] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100.
[0101] FIG3 is a software structure block diagram of the electronic device 100 according to an embodiment of the present application.
[0102] The layered architecture of electronic device 100 divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other via software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the framework layer, the Android runtime and system libraries, and the kernel layer.
[0103] The application layer can include a series of application packages.
[0104] As shown in FIG3 , the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video call, and short message.
[0105] The framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0106] As shown in FIG3 , the framework layer may include but is not limited to: a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, an image processing module, and the like.
[0107] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0108] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0109] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0110] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0111] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0112] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0113] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0114] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0115] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0116] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0117] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0118] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0119] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0120] A 2D graphics engine is a drawing engine for 2D drawings.
[0121] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, audio drivers, sensor drivers, Wi-Fi drivers, etc., which are not limited in this application.
[0122] It is understood that the components included in the framework layer, system library, and runtime layer shown in FIG3 do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or combine or split some components, or arrange the components differently.
[0123] Figure 4 is a schematic diagram of the software and hardware structure of the camera (also referred to as a camera device) provided in an embodiment of the present application. As shown in Figure 4, the embodiment of the present application uses a Linux system with a camera as a layered architecture as an example to illustrate the structure of the camera. The layered architecture of the camera divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. The architecture of the camera includes, from top to bottom: an application layer and a kernel layer.
[0124] The application layer may include, but is not limited to, applications such as image processing applications. Optionally, image processing applications may include, but are not limited to, sub-functions (or sub-applications) such as video input (VI), video processing sub-system (VPSS), video encoder (VENC), and video graphic system (VGS). Image processing applications are used to process images; for example, one or more sub-functions in the image processing application may be used to perform image processing processes such as noise reduction and color correction on the image; the processed image is then output to the electronic device. The application programs included in the application layer shown in FIG4 are merely illustrative and are not limited in this application. It is understood that the applications included in the application layer do not constitute a specific limitation on the camera. In other embodiments of the present application, the camera may include more or fewer applications than those included in the application layer shown in FIG4. Optionally, each application in the application layer may be pre-installed on the camera or the electronic device including the camera before leaving the factory, or may be installed when the camera or the electronic device including the camera is upgraded.
[0125] The kernel layer is the middle layer between the hardware and the software (i.e., the application layer). It passes application requests to the hardware and acts as a low-level driver, addressing various devices and components in the system. The kernel layer manages hardware devices and provides access to applications. The kernel layer includes one or more components. For example, it may include the system call interface, process management, memory management, calibration module, network stack, vision processing module, ISP driver, sensor driver, etc.
[0126] In an embodiment of the present application, the correction module can store and call a pre-trained AI network, which can also be referred to as an AI correction network or a correction network, which is not limited in this application. Specifically, the correction module can obtain the image and shooting parameters (such as FOV) captured by the camera, and output a correction grid based on the image and shooting parameters through the AI network.
[0127] Exemplarily, the visual processing module can run an image algorithm to process the image. For example, image algorithms may include, but are not limited to, image warping algorithms and face recognition algorithms. The visual processing module can invoke the image warping algorithm and, based on the correction grid, correct (also understood as "update") the image captured by the camera, outputting the corrected image. Furthermore, the visual processing module can invoke a face recognition algorithm to identify faces in the image.
[0128] It should be noted that the embodiments of the present application are all described by taking the correction module and visual processing module in the camera to execute corresponding steps to achieve image correction as an example. In other embodiments, the steps executed by the correction module and / or visual processing module can also be executed by a module in the electronic device (for example, a module in the kernel layer), and this application does not limit this.
[0129] Exemplarily, the camera hardware includes a CPU, an ISP, and sensors. Optionally, the hardware may also include devices such as memory. The ISP is used to process images and video streams and output the processed video streams and images in two ways. The CPU is merely an illustrative example; various microcontrollers such as a microcontroller unit (MCU) or devices that function as processors or microcontrollers may be alternatives to the CPU.
[0130] The ISP chip houses processing circuitry and transceiver pins. The processing circuitry controls the pins to send and receive data or signals, enabling communication with the camera's CPU and electronic devices (such as projectors). The processing circuitry also supports applications in the application layer, such as image processing applications.
[0131] The CPU chip houses processing circuitry and transceiver pins. Similarly, the processing circuitry controls the transceiver pins to send or receive data or signals, communicating with the electronic device and the camera's ISP, respectively. Furthermore, the processing circuitry supports applications or functions within applications, such as facial recognition applications.
[0132] It should be noted that, in the embodiments of the present application, each application (or module) is described as the main body for implementing each function. In fact, the function of each application is implemented by the processing circuit in the ISP or CPU, and will not be repeated below.
[0133] The sensor is the photosensitive element of the camera, which is used to collect light signals and convert the collected light signals into electrical signals. The electrical signals are then passed to the ISP for processing and converted into images or video streams.
[0134] Optionally, the CPU and ISP can be integrated on the chip or on different chips and connected via a bus. The CPU can respond to the request of the electronic device and output a control signal to the ISP through the control channel between the CPU and the ISP to trigger the corresponding processing circuit (also understood as a module) in the ISP. The CPU can also output data to the ISP through the data channel between the CPU and the ISP, such as the palm coordinates described in the embodiments of the present application. The ISP can output data to the CPU through the data channel between the ISP and the CPU. It should be noted that the control channel and data channel described above can refer to the same physical circuit or different physical circuits, and this application does not limit this.
[0135] FIG5 is a schematic diagram of an exemplary user interface. Referring to FIG5 , after the user clicks the camera application, the electronic device responds to the received user operation and displays a photo taking interface 500. The photo taking interface 500 includes but is not limited to: a preview interface 501 and other controls (or options).
[0136] Exemplarily, the image preview interface 501 displays the image captured by the camera in real time, which can also be called a preview image. Optionally, the image preview interface 501 can also include a zoom option 503 (also called a focal length setting option, which is not limited in this application), and the zoom option 505 is used to adjust the shooting focal length so that the preview interface displays an image within the zoom range corresponding to the zoom factor. Optionally, in an embodiment of the present application, the default initial zoom factor is 1x (also expressed as 1x). In an embodiment of the present application, the electronic device can provide different zoom factors, such as 3x (3x), 0.8x (0.8x), etc., which is not limited in this application.
[0137] Still referring to FIG. 5 , exemplary controls within camera interface 500 include, but are not limited to, a capture button 502 and mode options 504. If a user taps capture button 502, the electronic device responds to the received user action by taking a photo or recording a video. Exemplarily, mode options 504 include, but are not limited to, aperture, portrait, photo, and video. The mode options in FIG. 5 are for illustrative purposes only and are not intended to be limiting.
[0138] Exemplarily, the photo taking interface 500 may further include more options or controls, such as a settings option 505. If the user clicks on the settings option 505, the electronic device may display a settings interface box in response to the received user operation. The settings interface box may include, but is not limited to, photo ratio options, video resolution options, and the like. Exemplarily, the photo ratio option provides a variety of different photo ratios in photo taking scenarios, such as 4:3, 1:1, and full screen. The video resolution option provides different resolutions and corresponding recording image ratios in video recording scenarios, and the recording ratios include, but are not limited to, [15:9], [full screen], [21:9], or [15:9]. Among them, the various ratios involved in the embodiments of the present application may optionally be the ratio of the length (horizontal direction) to the width (vertical direction) of the image. The various numerical values are only illustrative examples and can be set according to actual needs. This application does not limit them.
[0139] For example, let's take a user selfie scenario as an example. The front camera captures an image in real time, including the user image and a background image (the background image is not shown). The preview image captured by the camera in real time is displayed in the preview interface 501. Optionally, the electronic device may perform image processing on the image captured by the camera to generate a preview image. For example, the resolution of the preview image may be lower than the image actually captured by the camera, but this is not limited in this application. Still referring to Figure 5, if the user clicks the capture option 502, the electronic device, in response to the received user operation, saves the image (or image frame) corresponding to the currently displayed preview image as a captured image, for example, in an album. Optionally, as described above, the electronic device may perform image processing on the image captured by the camera to generate and display the preview image. When saving the image, the electronic device may optionally perform image processing such as noise reduction on the image captured by the camera to generate an image (which may be referred to as a saved image in this embodiment of the application) and store it in a local album (which may also be another folder, but this is not limited in this application). Optionally, the resolution (or size) of the saved image is generally greater than the resolution of the preview image. It can be understood that to ensure the smoothness of real-time preview, the electronic device can reduce the resolution of the preview image to reduce the computational complexity during image processing. When saving the image, to preserve a clear image, the electronic device can perform image processing such as noise reduction, temperature adjustment, and color adjustment. In other words, the image captured by the camera can be output to different processes for processing, one of which can generate a preview image based on the image captured by the camera, and another process can generate a saved image based on the image captured by the camera. Optionally, in the following embodiments, another process can generate a thumbnail image based on the image captured by the camera.
[0140] Optionally, in response to the user clicking on the shooting option 502, the electronic device acquires the image and saves the image in the album, while the preview interface 501 can display the photographed image (also referred to as a selfie image, a photographed image, etc., which is not limited in this application), that is, the camera stops acquiring images. Among them, the photographed image displayed in the preview interface 501 can be an image that has been processed (such as noise reduction, etc.), that is, the resolution of the image is optionally greater than the original preview image. Optionally, in response to the user clicking on the shooting option 502, the electronic device acquires the image and saves the image in the album, and the preview interface 501 can continue to display the preview image captured by the camera in real time. It can be set according to actual needs and is not limited in this application.
[0141] FIG6 is a schematic diagram of an exemplary photographed image. Referring to FIG6 , image 601 is an image taken by a user, and image 601 includes a user image and a background image. The user image includes the user's facial image and images of other body parts. In this example, when a user uses a front-facing camera to take a portrait at close range, due to the difference in distances between different parts of the face and the camera, as shown in FIG6 , this usually results in facial image distortion problems such as a larger nose in the image, an elongated face (i.e., a longer middle part), and a reduced distance between the eyes. This results in a difference between the photographed photo and the actual facial appearance (or what the user sees when looking in the mirror).
[0142] The present application provides an image processing method. As shown in FIG6 , image 602 is an image corrected based on the image processing method in an embodiment of the present application, and may be referred to as corrected image 602. Based on the image processing method in an embodiment of the present application, image 601 is corrected to obtain corrected image 602. After the pupil distance, nose, and / or atrium of the facial image in corrected image 602 are corrected, facial distortion can be effectively eliminated, achieving what you see is what you get. That is, the captured photo is closer to the visual effect of the user's face when looking in the mirror, effectively improving image quality and enhancing the user experience.
[0143] FIG7 is a flowchart of an exemplary image processing method. Referring to FIG7 , the method specifically includes but is not limited to the following steps:
[0144] S701: Acquire a first image and first shooting parameters.
[0145] Exemplarily, the electronic device captures a first image, which may also be referred to as an image to be corrected. In one example, the image to be corrected may be a preview image displayed on a preview interface. That is, while the camera is capturing preview images in real time, the electronic device may perform the correction process in FIG. 7 on each preview image, and the preview interface of the electronic device may display each corrected preview image.
[0146] In another example, the image to be corrected may also be an image in an album. For example, the camera captures a preview image in real time, and the preview interface of the electronic device displays an uncorrected preview image. The user clicks on the shooting option, and the electronic device saves the current preview image, which can also be understood as the image corresponding to the preview image currently captured by the camera, to the album in response to the received user operation. Optionally, as described above, the image captured by the camera can be output to different threads, and one thread is used to process the image to generate a preview image with a lower resolution. Another thread processes the image (such as noise reduction, etc.) to generate an image with a resolution greater than the preview image, and saves it to the album. Subsequently, the electronic device performs correction processing on the image saved in the album (i.e., the image to be corrected).
[0147] Of course, similar to the photo scene, the video recording scene can also adopt real-time correction and delayed correction methods. That is, the electronic device can correct each image in the video stream captured by the camera in real time during the shooting process. The corrected video stream is displayed in the interface of the electronic device. Similarly, the electronic device can also correct each image in the video stream after the video stream is saved to the album. In the embodiment of the present application, the real-time correction method allows the user to observe the correction effect in real time and adjust the degree of correction. The specific adjustment method can be referred to below. The delayed processing method, that is, the method of correcting the image or video stream that has been saved (for example, in the album), can reduce the CPU occupancy rate. For example, the electronic device can execute the correction process at any time before the user views the image or video. Since the CPU occupancy rate is lower when not in the shooting mode, the calculation speed during the correction can be improved. The specific correction timing can be set according to actual needs and is not limited in this application. Optionally, the electronic device can also correct images such as thumbnails in the album. That is, the image to be corrected can also be an image of the thumbnail type, which is not limited in this application.
[0148] It should be noted that the embodiments of this application only describe the correction of images or video streams obtained from a camera shot using an electronic device. The image processing methods in the embodiments of this application can also be applied to scenarios such as facial recognition in surveillance cameras and smart TVs, making the corrected facial image closer to the actual appearance of the face, thereby improving the accuracy of facial recognition and other functions of the device.
[0149] Exemplarily, the electronic device acquires a first image captured by a camera and a first shooting parameter corresponding to the first image captured. In an embodiment of the present application, the shooting parameter may include a field of view (FOV) parameter, etc. It should be noted that the embodiment of the present application only uses the FOV parameter as an example for illustration. In other embodiments, the shooting parameter may also include other parameters, such as a brightness parameter, an exposure parameter, etc., which are not limited in this application.
[0150] Field of view, also known as the field angle of view, refers to the angular range in which a camera can receive an image in an imaging scene, and is also often referred to as the field of view. The size of the field of view determines the field of view of an optical instrument. The larger the field of view angle, the larger the field of view and the smaller the optical magnification. In layman's terms, at a certain fixed distance, if the size of the target object is too large and exceeds the field of view range, it will not be completely captured by the lens. As shown in Figure 8, FOV can be quantified from three directions, namely H FOV (horizontal field of view angle), V FOV (vertical field of view angle), and D FOV (diagonal field of view angle). In an embodiment of the present application, the shooting parameters may include at least two of H FOV, V FOV, and D FOV.
[0151] Exemplarily, different combinations of camera shooting ratios (also understood as the image aspect ratio, such as the photo ratio described above) and focal length parameters (also referred to as zoom parameters) correspond to a set of FOV information. For example, as shown in FIG5 , in a photo shooting scenario, the photo ratio (i.e., resolution) may include, but is not limited to, 4:3, 1:1, and full screen. Focal length parameters may include, but are not limited to, 1x, 0.8x, and wide angle. Any combination of resolution and focal length may correspond to a set of FOV parameters. It can be understood that different resolutions and focal lengths correspond to different FOV parameters. For example, in a photo shooting scenario, during the shooting process, the camera obtains the focal length parameter and shooting ratio in response to a received user operation. Based on the focal length parameter and shooting ratio, the corresponding FOV parameter can be calculated. The specific calculation method can be referenced in existing technologies and is not limited in this application. The camera can capture an image within the corresponding range based on the obtained FOV parameters. When executing step S701, after the electronic device (specifically, the camera) captures the image, it can further obtain the FOV parameters corresponding to the image capture.
[0152] In one example, as described above, the image to be corrected may be an image captured in real time. Accordingly, the electronic device may obtain the currently captured image and read the FOV parameters corresponding to the capture of the current image.
[0153] In another example, as described above, the image to be corrected can also be an image stored in an album (or other folder). In this example, the camera can mark the acquisition time of the image and save the corresponding timestamp. In addition, the camera obtains the FOV parameters when the image is acquired, and also records them corresponding to the timestamp. In this way, when executing subsequent processes, the electronic device can obtain the image acquired at that time point and the corresponding FOV parameters based on the timestamp, and align the timestamps to accurately match the image and FOV parameters to avoid calculation errors.
[0154] In one possible implementation, the electronic device may be provided with a correction switch. For example, the correction switch may be included in the photo taking interface. If the user clicks the correction switch, the electronic device starts the correction process, i.e., executes the process in FIG7 .
[0155] In another possible implementation, the electronic device may set a self-starting process for the correction process. Exemplarily, when the correction switch is turned on, the electronic device may obtain the shooting distance and, based on the shooting distance, determine whether to start the correction process. In one example, if the shooting distance is greater than a preset threshold (for example, 80 cm, which can be set according to actual needs and is not limited in this application), the correction process is not started, that is, the subsequent steps are not continued. It can be understood that when the shooting distance is far, the problem of facial image distortion usually does not occur. In another example, if the shooting distance is less than or less than or equal to the preset threshold, the correction process is started, that is, the subsequent correction process is continued. Optionally, the step of determining whether to start the process can be performed before executing S701, or at any time after executing S701 and before executing S703, and is not limited in this application. Optionally, the judgment process can also be applied to the step of obtaining the degree of correction (the concept of the degree of correction can be referred to below and will not be repeated here) in the following embodiments. For example, if the shooting distance is greater than a preset threshold, the degree of correction can be determined to be 0. If the shooting distance is less than the preset threshold, the degree of correction can be determined to be 10 (or 1).
[0156] Alternatively, the electronic device can obtain the shooting distance based on a variety of methods. The shooting distance can optionally be the distance between the camera and the subject (e.g., the user). Alternatively, in an embodiment of the present application, the facial image in the image is corrected. Accordingly, the shooting distance can optionally be the distance between the camera and the face, without estimating the distance between the camera and other background objects.
[0157] In one example, the electronic device can estimate the shooting distance by the size of the face image (for example, the proportion of the face image in the image). If the proportion of the face is smaller, the shooting distance is larger, and vice versa. In another example, the shooting distance can be further estimated based on the size and FOV of the face image to improve the accuracy of the shooting distance estimation. In another example, the electronic device can also estimate the shooting distance by the number of faces. For example, the more faces in the image, the longer the shooting distance, and vice versa. The shooting distance is closer. In other embodiments, the electronic device can also estimate the shooting distance by parameters measured by sensors such as TOF (Time of Flight) and laser. The specific acquisition method can be set according to actual needs and is not limited in this application.
[0158] In another possible implementation, the electronic device may also perform facial image recognition on the image (the specific recognition method can refer to existing technologies and will not be described in detail in this application). If the image to be corrected does not include a facial image, the steps subsequent to S701 may not be performed. That is, if it is determined that the image to be corrected does not include a face, the image does not need to be corrected according to the embodiment of this application. Of course, in this example, the electronic device can continue to perform other corrections on the image, such as correcting the verticality of lines in the image, which is not limited in this application. Optionally, this judgment step can also be performed after S702, that is, the electronic device obtains the correction grid and, after determining that the image does not include a facial image, no subsequent correction operations are performed. The specific timing can be set according to actual needs and is not limited in this application. Optionally, this judgment process can also be applied to the step of obtaining the correction degree (the concept of the correction degree can be referred to below and will not be described in detail here) in the following embodiments. For example, if the image does not include a face, the correction degree can be determined to be 0, and if the image includes at least one face, the correction degree can be determined to be 10 (or 1).
[0159] In another possible implementation, the electronic device can also perform plane image recognition (also known as 2D scene recognition) on the image to be corrected. If the image to be corrected is generated based on a 2D image (i.e., a plane image) (it can also be understood that the shooting scene or the shooting object is a plane image), there is no need to perform the subsequent correction process. It can be understood that image distortion is caused by the projection of 3D objects in the real world onto a 2D plane, resulting in the problem of "near big and far small" in the image. When taking a static 2D photo without depth, there will be no image distortion problem, that is, there is no need to perform the correction process. Specifically, the electronic device obtains the depth value of each pixel of the image to be corrected, determines whether the shooting scene includes a 2D scene based on the depth value, and determines whether to continue to perform the subsequent correction process based on the recognition result. For example, the electronic device recognizes the captured image to obtain the depth value of each pixel in the image (the specific acquisition method can refer to the existing technology and will not be repeated here). Alternatively, the electronic device can also recognize the shooting scene based on TOF technology to determine the depth value of the pixel in the captured image. Among them, if there are some pixels in the image (a threshold value can be set, for example, greater than or equal to 50% of the pixels, which can be set according to actual needs and is not limited in this application) whose depth difference (or the derivative of the depth value) is 0 (or other relatively small threshold values, which can be set according to actual needs and are not limited in this application), that is, some pixels are on a plane, then it can be determined that the scene captured by the camera is a 2D scene, or the scene corresponding to the captured image includes a 2D scene or a 2D image. In other words, the image captured by the camera is generated based on a 2D scene (or a scene containing a 2D scene). Optionally, the electronic device can also perform face recognition on the image to obtain a face image, and perform depth recognition on part of the face image to determine whether the depth difference of the pixels of the face image part is 0. If it is 0, it can be determined that the face image part is generated based on a 2D image. For example, the preview image may be an image obtained by capturing a photo on a desktop that includes a face image with a camera. In this image, the portion of the photo corresponding to the face image is generated based on a 2D image, while other backgrounds in the image, such as the desktop or other objects on the desktop, have a certain depth difference, that is, a 2D image generated from a 3D scene. For another example, if the background portion of the shooting scene is a 2D image, that is, the face is in front of the 2D image, then in the image captured by the camera, the face image is obtained based on a 3D scene (that is, a three-dimensional face), and there is a difference between the depth values of the corresponding pixels, while the difference between the depth values of the pixels corresponding to other background portions may be 0. Optionally, this judgment step can be performed after S701, that is, if it is detected that the image to be corrected is generated from a 2D image, there is no need to continue to perform operations such as generating a correction grid.Optionally, the judgment step can also be performed after S702, that is, after the correction grid is generated, it is judged that the image to be corrected is generated as a 2D image, and the subsequent correction operation is no longer performed. The specific timing can be set according to actual needs, and this application does not limit it. Optionally, the judgment process can also be applied to the step of obtaining the degree of correction (the concept of the degree of correction can be referred to below and will not be repeated here) in the following embodiments. For example, if the shooting scene includes a 2D image, the degree of correction can be determined to be 0. If the shooting scene does not include a 2D image or the face image is a non-2D image, the degree of correction can be determined to be 10 (or 1). In some instances, if there is a certain tilt angle between the plane image including the face and the camera, there is a difference between the depth values of the pixels of the face image acquired by the electronic device. In this example, the electronic device can determine whether the shooting scene corresponding to the face image is a plane image based on the derivative value of the depth value.
[0160] That is to say, the method of determining whether to continue the correction process in the embodiment of the present application can be understood as depending on the execution order of the modules. For example, the judgment process is executed by the visual processing module, and the generation of the correction grid (i.e., S702) is executed by the correction module. Accordingly, the correction module can continue to execute S702 to generate the correction grid, and at the same time, the visual processing module can execute the above judgment process in parallel. In this way, after the correction module generates the correction grid, it can be combined with the judgment result of the visual processing module to determine whether to further execute the subsequent process to improve processing efficiency. That is, if the visual processing module determines that the subsequent process should be continued, the electronic device can continue the subsequent process based on the obtained correction grid, reducing the time taken to wait for the judgment result. Of course, the method of first judging and then executing the subsequent process can effectively reduce the computing pressure of the electronic device. That is, when there is a demand (i.e., the judgment result indicates to continue execution), the subsequent correction grid generation process is continued, which can reduce the overhead of redundant calculations.
[0161] S702: Obtain at least one first correction parameter based on the first image and the first shooting parameter.
[0162] Exemplarily, after the electronic device acquires the first image (also referred to as the image to be corrected) and the corresponding first shooting parameters, it can obtain n correction parameters corresponding to n pixels in the first image based on the first image and the first shooting parameters. Specifically, the electronic device can generate a first correction mesh based on the first image and the first shooting parameters. The first correction mesh includes n dots, and the n dots correspond one-to-one to the n pixels in the first image. Each dot includes a correction parameter, and the correction parameter is used to indicate the correction offset (also referred to as correction displacement, which is not limited in this application) of the corresponding pixel in the first image in the first image. In the embodiment of the present application, only the grid is used as an example for explanation. In other embodiments, the correction parameters can also be in an array or other form, which is not limited in this application.
[0163] Exemplarily, after the electronic device starts the camera application, it can initialize the pre-stored AI network (also called the AI correction network). Initializing the AI network can also be understood as running the computer program corresponding to the AI network. The electronic device uses the first image and the first shooting parameter as the input of the AI network. The AI network outputs a first correction grid based on the first image and the first shooting parameter. Figure 9 is an exemplary diagram of the image correction process. Please refer to Figure 9. In this example, the scenario of correcting the image captured by the camera in real time is used as an example. The sensor driver obtains the first image captured by the camera. The correction module obtains the first image output by the sensor driver, and the correction module obtains the first FOV parameter corresponding to the first image captured by the camera. The correction module calls the AI network and uses the first image and the first FOV parameter as the input of the AI network. The AI network outputs the first correction grid based on the first image and the first FOV parameter. The correction module outputs the first correction grid to the visual processing module.
[0164] Exemplarily, the AI network is pre-trained by a server. The server trains the AI network based on a large number of images containing distorted facial images and different FOV parameters (which can also be understood as training the weights of the AI network), so that the AI network can output a correction grid that meets expectations. A correction grid that meets expectations can be understood as one that, after correcting the preset distorted facial image using the correction grid, can produce a facial image that is close to the actual appearance of the face, thereby achieving a good correction effect.
[0165] For example, after the server trains the AI network, it can push the AI network or the weight value of the AI network to each terminal. Optionally, the terminal may have pre-stored the AI network before leaving the factory. Optionally, in an application scenario, after the server updates the AI network, it can also send the weight value corresponding to the updated AI network to the electronic device, so that the electronic device can update the AI network based on the updated weight value.
[0166] Optionally, the AI network in the embodiment of the present application may be a convolutional neural network (CNN) or other types of networks, which is not limited in this application.
[0167] In an embodiment of the present application, the degree of distortion of the facial image is associated with the shooting distance, wherein the smaller the shooting distance, that is, the closer the distance between the face and the camera, the greater the degree of distortion of the facial image, and conversely, the farther the distance between the face and the camera, the smaller the degree of distortion of the facial image. During the training process, the training data set of the AI network includes a plurality of images containing facial images and corresponding shooting parameters, the shooting distances between the images are the same or different, and / or the FOV parameters corresponding to the images are the same or different. That is to say, during the training stage, the AI network is trained through a large number of images with different degrees of distortion and different FOVs, so that in actual application, the AI network can generate corresponding correction grids for images with different degrees of distortion and different FOVs to cope with different shooting environments.
[0168] Optionally, under the same shooting environment (i.e., the same camera and the same subject) and the same shooting distance, that is, when the degree of distortion of the facial image is the same, the correction grids and correction parameters generated by different FOVs are different. Specifically, different FOV parameters can correspond to different aspect ratios of the image. At the same shooting distance, the FOV parameters of the image are different, and the output correction grids and their corresponding correction parameters are also different. For example, at the same shooting distance, the larger the image FOV parameter, the more content in the picture, and the smaller the proportion of the image occupied by the face. The larger the image FOV parameter, the less content in the picture, and the larger the proportion of the image occupied by the face. In the embodiment of the present application, the shooting distance is associated with the degree of distortion, that is, when the shooting distance is the same, the degree of distortion of the face is the same. However, when the FOV is different, that is, the proportion of the facial image in the image is different, the position of the distortion may change, and accordingly, the pixel positions corresponding to the correction parameters of the same value will also be different.
[0169] In the embodiments of this application, the AI network is trained based on the entire image. Accordingly, in practical applications, the AI network can generate a corresponding correction grid based on the entire image. Compared to existing techniques that require identifying facial features and determining correction parameters based on these features, this application eliminates the need for facial recognition processing, effectively reducing the processing complexity of the AI network and improving its processing efficiency.
[0170] In one possible implementation, the number of dots (or grids) in the correction grid generated by the AI network in step S702 (to distinguish it from the first correction grid, the correction grid is called the initial correction grid) may also be smaller than the number of pixels in the image. For example, the resolution of the s-th image is 1400*1080, that is, it includes 1400*1080 pixels. The initial correction grid output by the AI network may include 64*64 grids. Among them, the size of the initial correction grid is still the same as the size of the first image, that is, the initial correction grid completely covers the first image, and some of the pixels do not correspond to dots. In this example, the electronic device (such as a correction module) can perform interpolation processing (also called upsampling processing) on the initial correction to generate a first correction grid with the same number of dots as the number of pixels. Optionally, during the interpolation process, the electronic device may perform a weighted average of the correction parameters of at least one grid point surrounding the position of the grid point to be inserted (e.g., the nine grid points surrounding the grid point to be inserted and closest to it, or other locations or numbers of grid points, which can be set according to actual needs and are not limited by this application) to obtain the correction parameters corresponding to the grid point to be inserted, and then insert the grid point into the initial correction grid. After performing a subtraction on all pixels that do not have a corresponding grid point, the electronic device can obtain the first correction grid.
[0171] For example, Figure 10a is a schematic diagram of an exemplary correction grid. Referring to Figure 10a, the first image includes n pixels (for example, 25 pixels, which is only an illustrative example and is not limited in this application). The electronic device executes S702, and the AI network outputs an initial correction grid. The initial correction grid includes m grid points (i.e., the intersection points of the grid), for example, 6 grid points, which is only an illustrative example and is not limited in this application. The electronic device performs a difference on the initial correction grid. For example, the electronic device can perform a weighted average based on the correction parameters of grid point 1, grid point 2, grid point 3, and grid point 4 to obtain the correction parameters corresponding to the grid point to be inserted. The electronic device inserts the grid point to be inserted into the initial correction grid, and the corresponding correction parameters are the parameters calculated after the above weighted average. The grid point to be inserted can also be called an inserted grid point in the first correction grid. The inserted grid point corresponds to pixel point 1 in the first image. It can be understood that the correction parameters corresponding to the inserted grid point are used to indicate the displacement of pixel point 1 in the first image.
[0172] Optionally, in an embodiment of the present application, the correction module may downsample the first image to obtain a downsampled image. The resolution of the downsampled image is smaller than the resolution of the first image (e.g., a preview image), which can also be understood as the number of pixels of the downsampled image is smaller than the number of pixels of the first image. Exemplarily, the correction module obtains a first correction grid based on the downsampled image and the first FOV parameter, thereby reducing the difficulty and complexity of the operation and improving the efficiency of the AI network in generating the correction grid. Optionally, in this example, the number of points of the generated first correction grid may be less than or equal to the number of pixels of the downsampled image. The AI network generates a correction grid with a number of points smaller than the number of pixels of the downsampled image, which can further reduce the processing complexity of the AI network. In the above example, the number of dots in the correction grid generated by the AI network is smaller than the number of pixels in the first image. The electronic device (which can be a correction module or a visual processing module) can perform interpolation processing on the correction grid to obtain a number of dots that is consistent with the number of pixels in the first image (i.e., the image captured by the camera, such as a preview image), so that the dots of the first correction grid correspond one-to-one to the pixels of the first image.
[0173] S703: Correct the distorted portion of the facial image of the first image based on at least one first correction parameter to obtain a first corrected image.
[0174] For example, after obtaining at least one first correction parameter, the electronic device may correct the facial image portion of the first image based on the at least one first correction parameter to correct the lens distortion portion of the facial image. Specifically, the electronic device corrects the first image based on the correction grid output by the AI network to obtain a first corrected image. The facial image in the first corrected image closely resembles the actual facial appearance, achieving a "mirror" effect.
[0175] Specifically, still referring to FIG9 , the correction module outputs the first correction grid to the visual processing module. The visual processing module may correct the first image output by the sensor based on the first correction image and output the first correction image. Optionally, the visual processing module may run a related graphics algorithm, such as an image warping algorithm. The visual processing module may correct at least one pixel in the first image using the image warping algorithm based on the correction parameters corresponding to each point in the correction grid.
[0176] In an embodiment of the present application, in an embodiment of the present application, the number of grids (or the number of dots) in the correction grid corresponds one-to-one to the pixels of the image, that is, the correction grid in the present application covers all the pixels in the image. Each dot includes a correction parameter, which is used to indicate the correction offset of the corresponding pixel in the image. It can also be understood as being used to indicate the twisting direction (also known as deformation direction, movement direction, etc., which is not limited in this application) and twisting amplitude (also known as deformation amplitude, movement amplitude, etc., which is not limited in this application) when the visual processing module performs the image warping algorithm on the image. It can also be understood as being used to indicate the position of the corresponding pixel in the corrected image. Due to facial distortion, the visual processing module can correct at least one pixel of the facial image part based on the correction parameters in the correction grid, correct the positions of some pixels that cause facial distortion, and obtain a corrected image. It can be understood that, among the correction parameters corresponding to the n points of the correction grid, the correction parameters for pixels in the distorted face image portion are greater than 0, while the correction parameters corresponding to other pixels (e.g., pixels corresponding to the undistorted portion and / or background portion) can optionally be 0, indicating that no correction is required. The correction grid in the embodiments of the present application is a global vector generated based on the entire image (i.e., each correction parameter can also be considered a vector). The corrected image generated based on this correction grid is more natural and smooth.
[0177] For example, as shown in Figure 6, due to the problem of facial distortion, some pixels in the face image in image 601 (for example, at least one pixel corresponding to the image of the nose part) are offset, so that the width of the nose part in the image is greater than the actual width of the user's nose. Take pixel 6011 in the nose image part as an example. The coordinate position of pixel 6011 in image 601 is (x1, y1). Among them, the coordinate position is established with the upper left corner of image 601 as the coordinate origin. The coordinate system can be set according to actual needs and is not limited in this application. The AI network outputs a correction grid based on image 601 and the corresponding FOV parameters. The correction grid includes a grid point A corresponding to pixel point 6011, and grid point A includes correction parameter A. Correction parameter A is used to indicate the offset of pixel point 6011 in image 601, which can also be understood as indicating the torsion direction and torsion amplitude. Accordingly, in step S703, the electronic device (e.g., the visual processing module in the camera) can correct the corresponding pixel 6011 based on the correction parameter A of the dot A (e.g., twisting or shifting the pixel position), so that it moves to the location of the pixel 6011 in the image 602, for example, the coordinate position (x2, y2) (wherein the coordinate system of the image 602 is based on the upper left corner of the image 602 as the origin, and the image 602 and the image 601 have the same size). The offset between the coordinate position (x2, y2) and the coordinate position (x1, y1) in the horizontal direction (e.g., the x-axis direction) and the vertical direction (e.g., the y-axis direction) is the correction displacement (or correction offset) described above. The visual processing module performs the same correction processing on the other pixels of the nose and the pixels corresponding to other parts (e.g., the pixels of the eyes), so as to shorten the width of the nose, shorten the middle part of the face, and / or shorten the distance between the eyes, thereby obtaining a corrected image 602, which is close to the actual appearance of the user.
[0178] In one possible implementation, as described above, the electronic device can perform the aforementioned correction processing on each preview image captured in real time by the camera, similar to the processing performed on the first image. After acquiring the second image, the electronic device can generate a second correction grid based on the FOV parameters corresponding to the second image and the second image. In one example, if the distance between the user and the camera changes, the degree of distortion or the location of distortion in the second image relative to the first image may also change accordingly (for example, the nose is distorted in the first image, while the chin is distorted in the second image). Accordingly, given the same FOV parameters, the correction grids output by the AI network based on different images will also differ. For example, in the first correction grid, the correction parameter A corresponding to the nose may be greater than 0, while the correction parameter corresponding to the pixels in the chin may be 0 (or a smaller value). In contrast, in the second correction grid, the correction parameter A corresponding to the nose may be 0 (or a smaller value), while the correction parameter B corresponding to the pixels in the chin may be greater than 0. The electronic device can correct the second image based on the second correction grid to obtain a second corrected image.
[0179] Optionally, in an embodiment of the present application, for the correction process of two adjacent images, the electronic device can update the correction parameters of the correction grid obtained this time based on the correction parameters of the previous correction grid (i.e., the correction grid corresponding to the adjacent previous image) to avoid image mutations, which lead to correction mutations caused by large differences in correction parameters. For example, still taking the second correction grid as an example, after the electronic device obtains the second correction grid, it can perform weighted averaging of the correction parameters of each corresponding grid point in the two correction grids based on the previous correction grid (e.g., the first correction grid) to update the correction parameters of each grid point in the second correction grid. The electronic device can correct the second image based on the updated second correction grid. In one example, the electronic device can obtain motion information such as the speed or displacement of the face, and determine the corresponding smoothing amplitude based on the speed or displacement of the face. The smoothing amplitude can optionally be the weight value in the weighted average. In an embodiment of the present application, the faster the face moves, that is, the greater the displacement of the face image between the two image frames, the smaller the smoothing amplitude, that is, the smaller the weight value of the correction grid corresponding to the previous image (for example, the first correction grid) is (for example, 0.1), and the larger the weight value of the correction grid corresponding to the current image (for example, the second correction grid) is (for example, 0.9). On the contrary, if the face moves slower, that is, the smaller the displacement of the face image between the two image frames, the larger the smoothing amplitude, that is, the larger the weight value of the correction grid corresponding to the previous image (for example, 0.9), and the smaller the weight value of the correction grid corresponding to the current image (for example, 0.1). Optionally, the electronic device can set different smoothing amplitude gears (i.e., weight value ratios) corresponding to different face movement speeds or displacements. The specific values can be set according to actual needs and are not limited by this application. For example, taking displacement as an example, the electronic device pre-sets the displacement threshold of the facial image between two image frames to 5 pixels (or 5% of the image width, etc., which can be set according to actual needs and is not limited by this application). If the displacement of the facial image between the two images is less than or equal to 5 pixels, the weight value of the correction grid corresponding to the previous image is 0.9, and the weight value of the correction grid corresponding to the current image is 0.1. If the displacement of the facial image between the two images is greater than 5 pixels and the excess value is 10% to 20% (for example, 10% of 5 pixels), the weight value of the correction grid corresponding to the previous image is 0.8, and the weight value of the correction grid corresponding to the current image is 0.2. If the displacement of the facial image between the two images is greater than 5 pixels and the excess value is 20% to 30%, the weight value of the correction grid corresponding to the previous image is 0.8, and the weight value of the correction grid corresponding to the current image is 0.2... The above values are only illustrative examples and can be set according to actual needs and are not limited by this application.Optionally, the electronic device can obtain the facial motion speed or displacement by performing facial recognition on the previous and current image frames, and determining the motion speed and / or displacement based on the position of the facial frame in the two image frames. This temporal smoothing technology can effectively improve the consistency and stability between correction grids, and thus the consistency and stability between corrected images.
[0180] In another possible implementation, the user can also adjust the photo ratio (or video ratio) and / or focal length to adjust the FOV parameters when the camera captures the image. Accordingly, as the FOV parameters are modified, the degree of distortion and / or position of the facial image will also change. For example, still taking the real-time preview scene as an example, the electronic device determines the updated photo ratio and / or focal length in response to the received user operation. The camera can recalculate the FOV parameters based on the updated photo ratio and / or focal length to obtain a third FOV parameter. The electronic device obtains a third image captured by the camera based on the third FOV parameter, and obtains the third FOV parameter. The electronic device can obtain a third correction grid based on the third image and the third FOV parameter through the AI network. And correct the third image based on the third correction grid. That is, in the embodiment of the present application, by introducing the FOV parameter and generating the correction grid based on the FOV parameter and the image, it is possible to achieve dynamic updating of the correction parameter to adapt to different FOV parameter scenarios. It can be understood that different FOV parameters may cause different degrees of distortion and / or distortion positions of facial images. This application generates a correction grid through FOV parameters to correct the image, thereby meeting the scene requirements of different FOV parameters, so that the corrected image is more in line with the user's actual appearance and improves the accuracy of the correction.
[0181] In another possible implementation, the user can also select a corresponding correction degree based on the observed corrected image, wherein the correction degree is used to indicate the degree of influence of the correction parameter on the image. For example, Figures 10b and 10c are schematic diagrams of user interfaces shown exemplarily. Referring to Figure 10b, the electronic device displays a photo preview interface in response to the user clicking the camera application. The interface description can be referred to above and will not be repeated here. Exemplarily, the electronic device obtains a first image and a first FOV parameter. Based on the judgment conditions described above, the electronic device determines that the first image includes a face image, and the shooting distance is less than a preset threshold (e.g., 80cm), and the acquired image is generated based on a 3D real scene image, which meets the correction process triggering conditions. The electronic device starts to perform image correction on the first image. Optionally, a correction prompt box 1001 is displayed in the preview interface 501, such as "Image correction is turned on" to prompt that the currently displayed image is an image after image correction. Correction prompt box 1001 may include a cancel option. If the user clicks the cancel option, the electronic device, in response to the received user action, disables the image correction function for this time, i.e., no further image correction is performed until the camera application is closed. Optionally, the image correction function can be automatically enabled again after the camera application is next launched. The user can manually disable the image correction function in the camera settings. After this function is disabled, the electronic device will no longer automatically initiate the image correction process.
[0182] Still referring to Figure 10b, in this example, the electronic device uses a default correction level of 10 as an example. It should be noted that the numerical values in the embodiments of this application are merely illustrative examples and are not intended to be limiting. For example, the electronic device obtains a current correction level (referred to as the first correction level) of 10, which can also be understood as 100%. Based on the first correction level, the electronic device updates the correction parameters of each dot in the first correction grid. For example, the correction parameters of each dot are multiplied by the coefficient corresponding to the correction level. For example, if the correction level is 10, the corresponding coefficient is 1. The electronic device corrects the first image based on the updated first correction grid to obtain a first corrected image. The electronic device displays the first corrected image in the preview interface 501. The electronic device performs the same processing on subsequent images. During the preview process, if the user clicks the correction prompt box 1001 (see Figure 10c), the electronic device displays the correction level setting box 1002 in response to the received user operation. The correction level setting box 1002 displays the current correction level (e.g., 10) and also includes adjustment options. The user can click on the adjustment option to adjust the degree of correction. For example, the user clicks on the "-" option to reduce the degree of correction to 5. Accordingly, the electronic device obtains the second image captured by the camera (there may be multiple images between the second image and the first image) and the second FOV parameter corresponding to the capture of the second image. The electronic device generates a second correction grid based on the second image and the second FOV parameter. The electronic device obtains the correction coefficient corresponding to the current correction degree (i.e., the correction degree is 5), for example, 0.5. Based on the correction degree, the electronic device updates the correction parameters of each point in the second correction grid, for example, multiplying each correction parameter by 0.5. The electronic device corrects the second image based on the updated second correction grid to obtain a second corrected image. The electronic device displays the second corrected image in the preview interface 501. The correction effects vary under different correction levels. For example, the correction level for nose width in the first image is a first correction level, and the correction level for nose width in the second image is a second correction level. Assuming that the first correction level is greater than the second correction level, the nose width in the first correction image is reduced from the first width in the first image to the second width, and in the second correction image, the nose width is reduced from the first width in the second image to a third width. The third width may be greater than the second width. In this way, the present application can provide a user interaction interface so that users can adjust the correction level as needed to improve the user experience.
[0183] The user interface in Figure 10c is only an illustrative example. In other embodiments, the correction degree adjustment options can also be other graphics. In one example, Figure 10d is an illustrative user interface diagram. Please refer to Figure 10d. After the image correction is turned on, the user can click on the correction prompt box 1001. The electronic device responds to the received user operation and displays the correction degree setting gear 1003. Among them, it includes 0 to 10 gears, and each vertical bar represents a gear. It should be noted that the operation of triggering the display of the correction degree setting gear and the correction degree setting box in other embodiments is only an illustrative example. The user can also long press the correction prompt box 1001, or the corresponding degree adjustment option will be automatically displayed after the correction is turned on. This is not limited in this application. Still referring to Figure 10d, the interface also includes a pointer 1004. The user can move the pointer 1004 to select the corresponding correction degree gear.
[0184] In another example, Figure 10e is a schematic diagram of an exemplary user interface. Please refer to Figure 10e. After the image correction is turned on, the user can click the correction prompt box 1001. The electronic device responds to the received user operation and displays the correction degree setting control 1005. The correction degree setting control 1005 includes a sliding option control 1006 (also called a sliding bar, not limited in this application). The user can adjust the correction degree by sliding the control left and right. The electronic device can determine the corresponding correction degree based on the user's operation on the sliding option control 1006. The current corresponding correction degree can be displayed in the interface, for example, a correction degree of 7.3. In addition, the electronic device can update the correction parameters of the correction grid based on the correction degree set by the user.
[0185] In another possible implementation, in addition to user settings and / or default values, the correction degree can also be set based on the deflection angle of the facial image in the image. Specifically, the electronic device can identify the facial image in the first image to obtain the deflection angle of the facial image in the image, which can also be understood as the turning angle of the face. The electronic device can pre-set correction levels corresponding to different deflection angle ranges, where the smaller the deflection angle, the greater the correction degree, and the larger the deflection angle, the smaller the correction degree. The specific values can be set according to actual needs and are not limited by this application. For example, if the electronic device obtains a deflection angle of 0 degrees for the face, i.e., the face is facing the camera directly, the electronic device determines the corresponding correction degree to be 10 (i.e., 100%). For another example, if the electronic device obtains a deflection angle of greater than or equal to 90 degrees for the face, i.e., the face is facing the camera sideways, the electronic device determines the corresponding correction degree to be 0. Optionally, the electronic device can also set only two correction levels: 0 and 10 (or 1). For example, if the turning angle of the face is less than 90 degrees, the correction degree is 10. If the face's turning angle is greater than or equal to 90 degrees, the degree of correction is 0. Of course, in some instances, this judgment condition can also be applied to S702, that is, before generating the correction grid. For example, if the face's turning angle is greater than or equal to 90 degrees, S702 is not executed; if it is less than 90 degrees, S702 is continued. Alternatively, this judgment condition can also be before S703. For example, if the face's turning angle is greater than or equal to 90 degrees, S703 is not executed; if it is less than 90 degrees, S703 is continued.
[0186] In another possible implementation, as shown in FIG11 , the electronic device (e.g., a visual processing module) may correct the thumbnail image based on the correction grid corresponding to the preview image to obtain a corrected thumbnail image, and / or correct the saved image saved to the album based on the correction grid corresponding to the preview image to obtain a corrected saved image.
[0187] For example, FIG12 is a schematic diagram of an exemplary user interface. Please refer to FIG12 (1). In the real-time correction scenario, the camera captures an image, and the camera can perform image processing on the captured image (the relevant processing can refer to existing technologies, such as downsampling, etc., to reduce the resolution of the image, which is not limited in this application) to generate a preview image. The correction module can generate a corresponding correction grid based on the preview image, and correct the preview image based on the correction grid to obtain a corrected preview image. The electronic device displays the corrected preview image in the preview interface. The electronic device corrects each preview image and displays the corrected image (for example, called a corrected preview image).
[0188] For example, the user can click on the shooting option 502. As shown in (2) of Figure 12, in response to the received user operation, the electronic device obtains the image captured by the camera, the correction grid and the timestamp (used to indicate the image capture time). The electronic device displays a thumbnail generated based on the image captured by the camera in the thumbnail box 500. Optionally, the image in the thumbnail is an uncorrected image, for example, as shown in image 601. Optionally, as described above, the image captured by the camera can be processed by different threads, one thread is used to generate a preview image from the image captured by the camera, another thread can generate a saved image based on the image captured by the camera, and another thread can generate a thumbnail image based on the image captured by the camera. That is, based on the image captured by the camera, the electronic device can obtain three different types of images, and the resolutions of the three images are different to meet different user needs. For example, the resolution of the preview image is low to meet the smoothness of the preview image. The resolution of the saved image is high to meet the clarity requirement. The resolution of the thumbnail is the lowest to meet the need for fast image output, that is, after the user clicks on the shooting option, the electronic device can display the corresponding thumbnail in the thumbnail box 500.
[0189] In one example, the user clicks on the thumbnail box 500. As shown in (3) of Figure 12, the electronic device responds to the received user operation and enlarges the thumbnail in the thumbnail box 500 in the interface, for example, image 601. This image is an uncorrected thumbnail. The electronic device (for example, a visual processing module) obtains the thumbnail (the thumbnail can also be the first image described in the embodiment of the present application), and the electronic device queries the correction grid corresponding to the timestamp of the thumbnail (used to indicate the acquisition time of the image) and queries the correction grid saved corresponding to the timestamp. The correction grid is the correction grid generated based on the preview image and the corresponding FOV parameters as described above (for example, it can be the first correction grid described in the embodiment of the present application). As shown in Figure 11, the electronic device corrects the thumbnail based on the correction grid corresponding to the preview image to obtain a corrected thumbnail. As shown in (4) of Figure 12, the electronic device displays the corrected thumbnail, that is, the corrected thumbnail, for example, image 602. The description of image 601 and image 602 can be referred to Figure 6 and will not be repeated here. Optionally, the electronic device may obtain a corresponding saved image (which may also be referred to as a finished image or a large image, which is not limited in this application) based on the image captured by the camera, and the electronic device saves the saved image to an album.
[0190] For example, referring to FIG11 , an electronic device (e.g., a visual processing module) acquires a saved image (also understood as the first image described in the embodiments of this application). Based on the timestamp of the saved image (indicating the time when the image was captured), the electronic device queries the correction grid stored corresponding to the timestamp. This correction grid is the correction grid generated based on the preview image and the corresponding FOV parameters described above.
[0191] Exemplarily, the electronic device corrects the stored image based on the correction grid to obtain a corrected stored image. The correction method can be referred to above and will not be described in detail here. Optionally, the electronic device can perform some image processing on the stored image (such as adjusting the temperature, adjusting the color, etc.). The electronic device can also process the current image based on the previous frame image. The specific processing process can refer to the existing technology and is not limited in this application. The electronic device replaces the thumbnail displayed in (4) of Figure 12 with the stored image. The stored image is the image after correction and image processing. Exemplarily, the stored images in the electronic device correspond to different display modes. Figure 13 is an exemplary display diagram of the stored image. Please refer to Figure 13. The electronic device can display a small image of the corrected stored image in the album in response to the received user operation.
[0192] In another example, after the electronic device obtains the thumbnail, it can correct the thumbnail and display the corrected thumbnail in the thumbnail frame 500. Accordingly, after the user clicks the thumbnail frame 500, the image displayed on the user interface is the corrected thumbnail.
[0193] In another example, after the electronic device obtains the thumbnail and displays the thumbnail in the thumbnail frame, the electronic device corrects the thumbnail. Accordingly, after the user clicks the thumbnail frame 500, the image displayed on the user interface is the corrected thumbnail.
[0194] That is to say, in an embodiment of the present application, the electronic device may correct the thumbnail and / or saved image corresponding to the first image based on the correction grid corresponding to the first image to obtain the corresponding corrected image. In an embodiment of the present application, a thumbnail, a saved image and / or a preview image may also be used as the first image. For example, the electronic device may also generate a correction grid based on the thumbnail and correct the thumbnail. Then, the saved image may be corrected based on the correction grid of the thumbnail. And / or, the electronic device may generate a correction grid based on the saved image and correct the saved image. This application does not limit this.
[0195] Thus, in this embodiment of the present application, the electronic device can achieve timestamp alignment by recording the correction grid and the corresponding timestamp, ensuring consistent correction effects across preview images, saved images, and thumbnails. Furthermore, by using the correction grid for the preview image to correct the saved image and thumbnail corresponding to that preview image, duplicate correction grid generation is avoided, thereby improving image correction efficiency.
[0196] For example, in the above embodiments, the correction of the preview image captured by the camera in real time by the electronic device is used as an example. In the embodiments of the present application, the electronic device can also adopt a delayed correction method. For example, still taking the scene in Figure 10b as an example, the user can click the shooting option 502, and the electronic device responds to the received user operation, and the electronic device obtains the currently displayed preview image and the corresponding first FOV parameter. The electronic device saves the preview image to the album. Optionally, the electronic device can perform image processing on the preview image and save it to the album. That is to say, the parameters such as resolution of the first image saved in the album and the preview image may be different, and this application does not limit this. In this example, the electronic device saves the first timestamp (used to indicate the acquisition time of the first image), the first image and the first FOV parameter accordingly. During the correction process, the electronic device can search for the first FOV parameter corresponding to the first timestamp based on the first image and the first timestamp. The electronic device can perform a subsequent correction process based on the first image and the first FOV parameter to ensure that the first FOV parameter matches the first image to generate an accurate correction grid, that is, the generated correction grid is in line with the shooting scene (that is, corresponding to the FOV parameter) requirements when capturing the preview image corresponding to the first image.
[0197] Optionally, the electronic device may also perform some post-image processing on the corrected image, such as beautification, which is not limited in this application. Optionally, the electronic device generates a correction grid based on the first image. Before correcting the first image, the electronic device may perform image processing such as noise reduction on the first image, and then correct the processed first image based on the correction grid.
[0198] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the interaction between various network elements. It can be understood that in order to realize the above functions, the wireless signal coverage detection device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0199] The embodiment of the present application can divide the functional modules of the image processing device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0200] In the case of dividing the functional modules according to the corresponding functions, Figure 14 shows a possible structural diagram of the image processing device 1400 involved in the above embodiment. As shown in Figure 14, the image processing device may include: a first acquisition module 1401, used to obtain at least one first correction parameter based on the first image and the first shooting parameter, wherein the first image is captured by the camera based on the first shooting parameter, the first image includes a face image, the at least one first correction parameter corresponds one-to-one to at least one pixel point in the first image, and the first correction parameter is associated with the shooting distance between the camera and the face; a correction module 1402, used to correct the distorted part of the face image in the first image based on the at least one first correction parameter to obtain a first corrected image.
[0201] In one possible implementation, the first acquisition module is specifically configured to input the first image and the first shooting parameter into the AI network, and obtain a first correction grid output by the AI network; wherein at least one grid point in the first correction grid corresponds one-to-one to at least one pixel point in the first image, and the grid point in the first correction grid is used to indicate the first correction parameter.
[0202] In one possible implementation, the device also includes: a second acquisition module 1403, used to obtain a first correction degree parameter; a correction module, specifically used to: update the first correction parameter corresponding to each grid point in the first correction grid based on the first correction degree parameter; based on the updated first correction grid, correct the distorted part of the facial image in the first image to obtain a first corrected image.
[0203] In a possible implementation, the second acquisition module 1403 is specifically configured to: acquire the first correction degree parameter in response to a received user operation; or acquire the first correction degree parameter based on a deflection angle of the face image in the first image.
[0204] In one possible implementation, the second acquisition module 1403 is specifically used to: determine the first correction degree parameter in response to a received adjustment operation of the correction degree parameter adjustment option in the user interface; display the first correction degree parameter in the user interface; the device also includes: a display module for displaying the first corrected image in the user interface.
[0205] In one possible implementation, the AI network is trained based on multiple training images containing facial images and corresponding shooting parameters, wherein the multiple training images containing facial images have different shooting distances and corresponding shooting parameters.
[0206] In one possible implementation, the correction module 1402 is specifically configured to: obtain facial motion information based on a facial image; update first correction parameters of the points in a first correction grid based on the facial motion information and correction parameters corresponding to the points in a second correction grid, wherein the second correction grid is obtained based on a previous frame image adjacent to the first image and corresponding shooting parameters; and correct the distorted portion of the facial image in the first image based on the updated first correction grid to obtain a first corrected image.
[0207] In one possible implementation, the correction module 1402 is further used to: obtain a depth value of each pixel in the first image; based on the depth value, determine whether the shooting scene corresponding to the facial image is a planar image; and based on the judgment result, determine whether to correct the distorted part of the facial image in the first image based on at least one first correction parameter.
[0208] In one possible implementation, the shooting distance corresponding to the first image is the first shooting distance; the first acquisition module 1401 is further used to obtain at least one second correction parameter based on the second image and the first shooting parameter, wherein the second image is captured by the camera based on the first shooting parameter, the second image includes a face image, the shooting distance corresponding to the second image is the second shooting distance, the second shooting distance is smaller than the first shooting distance, at least one second correction parameter corresponds one-to-one to at least one pixel point in the second image, the degree of distortion of the face image of the second image is greater than the degree of distortion of the face image of the first image, and the second correction parameter corresponding to the distorted part of the face image of the second image is greater than the first correction parameter corresponding to the distorted part of the face image of the first image; the correction module 1402 is further used to correct the distorted part of the face image in the second image based on the at least one second correction parameter to obtain a second corrected image.
[0209] In a possible implementation, the first correction parameter is used to indicate a correction displacement of a corresponding pixel point in the first image.
[0210] In a possible implementation, the first shooting parameter includes a field of view FOV parameter.
[0211] In a possible implementation, the first image is a preview image captured in real time by a camera.
[0212] In one possible implementation, the correction module is further used to: correct the thumbnail corresponding to the preview image based on at least one first correction parameter to obtain a corrected thumbnail; and / or correct the image saved in the album corresponding to the preview image based on at least one first correction parameter to obtain a corrected saved image.
[0213] In a possible implementation, the first acquisition module 1401 is further configured to: perform downsampling processing on the first image; and acquire the at least one first correction parameter based on the downsampled first image and the first shooting parameter.
[0214] In another example, Figure 15 shows a schematic block diagram of an image processing device 1500 according to an embodiment of the present application. The image processing device may include a processor 1501 and a transceiver / transceiver pin 1502, and optionally, a memory 1503. The processor 1501 may be configured to execute the steps performed by the image processing device in each method of the aforementioned embodiments, and control the receive pin to receive signals and the transmit pin to send signals.
[0215] The various components of image processing apparatus 1500 are coupled together via bus 1504. Bus system 1504 includes not only a data bus but also a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 1504 in the figure.
[0216] Optionally, the memory 1503 may be used to store instructions in the aforementioned method embodiment.
[0217] It should be understood that the image processing device 1500 according to an embodiment of the present application may correspond to the electronic device in each method of the aforementioned embodiment, and the above-mentioned and other management operations and / or functions of each element in the image processing device 1500 are respectively for implementing the corresponding steps of each of the aforementioned methods. For the sake of brevity, they will not be repeated here.
[0218] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0219] Based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes at least one section of code, and the at least one section of code can be executed by an electronic device to control the electronic device to implement the above method embodiment.
[0220] Based on the same technical concept, an embodiment of the present application also provides a computer program, which, when executed by an electronic device, is used to implement the above method embodiment.
[0221] The program may be stored in whole or in part on a storage medium packaged with the processor, or may be stored in whole or in part on a memory not packaged with the processor.
[0222] Based on the same technical concept, the embodiment of the present application further provides a processor, which is used to implement the above method embodiment. The above processor can be a chip.
[0223] The steps of the method or algorithm described in conjunction with the disclosure of the embodiments of the present application can be implemented in hardware or by executing software instructions by a processor. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a network device. Of course, the processor and storage medium can also be present in a network device as discrete components.
[0224] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0225] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. An image processing method, characterized in that: include: Obtaining at least one first correction parameter based on a first image and a first shooting parameter, wherein the first image is captured by a camera based on the first shooting parameter, the first image includes a face image, the at least one first correction parameter corresponds one-to-one to at least one pixel in the first image, and the at least one first correction parameter is associated with a shooting distance between the camera and the face; Based on the at least one first correction parameter, the distorted portion of the facial image in the first image is corrected to obtain a first corrected image.
2. The method according to claim 1, characterized in that The acquiring at least one first correction parameter based on the first image and the first shooting parameter includes: The first image and the first shooting parameter are input into an AI network, and a first correction grid is obtained as output by the AI network; wherein at least one grid point in the first correction grid corresponds one-to-one to at least one pixel point in the first image, and each grid point in the first correction grid corresponds to one of the first correction parameters.
3. The method according to claim 2, characterized in that Before correcting the distorted portion of the facial image in the first image based on the at least one first correction parameter to obtain the first corrected image, the method includes: Obtaining a first correction degree parameter; The correcting the distorted portion of the facial image in the first image based on the at least one first correction parameter to obtain a first corrected image includes: Based on the first correction degree parameter, updating a first correction parameter corresponding to each grid point in the first correction grid; Based on the updated first correction grid, the distorted portion of the facial image in the first image is corrected to obtain the first corrected image.
4. The method according to claim 3, characterized in that The obtaining of the first correction degree parameter includes: In response to the received user operation, obtaining the first correction degree parameter; or, The first correction degree parameter is obtained based on the deflection angle of the facial image in the first image.
5. The method according to claim 4, characterized in that The acquiring the first correction degree parameter in response to the received user operation includes: Display correction degree parameter adjustment options in the user interface; determining the first correction degree parameter in response to a received adjustment operation on a correction degree parameter adjustment option; displaying the first correction degree parameter in the user interface; The method further comprises: The first rectified image is displayed in the user interface.
6. The method according to claim 2, characterized in that The AI network is trained based on multiple training images containing facial images and corresponding shooting parameters, wherein the multiple training images containing facial images correspond to different shooting distances and / or different corresponding shooting parameters.
7. The method according to claim 2, characterized in that The correcting the distorted portion of the facial image in the first image based on the at least one first correction parameter to obtain a first corrected image includes: Based on the facial image, acquiring facial motion information; updating the correction parameters corresponding to the grid points in the first correction grid based on the facial motion information and the correction parameters corresponding to the grid points in the second correction grid, wherein the second correction grid is obtained based on a previous frame image adjacent to the first image and the shooting parameters corresponding to the previous frame image; Based on the updated first correction grid, the distorted portion of the facial image in the first image is corrected to obtain the first corrected image.
8. The method according to any one of claims 1 to 7, characterized in that Before correcting the distorted portion of the facial image in the first image based on the at least one first correction parameter to obtain the first corrected image, the method further includes: Obtaining a depth value of each pixel in the first image; Based on the depth value, determining whether the shooting scene corresponding to the facial image is a planar image; Based on the judgment result, it is determined whether to correct the distorted portion of the facial image in the first image based on the at least one first correction parameter.
9. The method according to any one of claims 1 to 8, characterized in that The shooting distance corresponding to the first image is a first shooting distance; and the method further includes: Obtaining at least one second correction parameter based on a second image and the first shooting parameters, wherein the second image is captured by the camera based on the first shooting parameters, the second image includes a facial image, the shooting distance corresponding to the second image is a second shooting distance, the second shooting distance is less than the first shooting distance, the at least one second correction parameter corresponds one-to-one to at least one pixel in the second image, the degree of distortion of the facial image in the second image is greater than the degree of distortion of the facial image in the first image, and the second correction parameter corresponding to the distorted portion of the facial image in the second image is greater than the first correction parameter corresponding to the distorted portion of the facial image in the first image; Based on the at least one second correction parameter, the distorted portion of the facial image in the second image is corrected to obtain a second corrected image.
10. The method according to any one of claims 1 to 9, characterized in that The first correction parameter is used to indicate a correction displacement of a corresponding pixel point in the first image.
11. The method according to any one of claims 1 to 10, characterized in that The first shooting parameter includes a field of view FOV parameter.
12. The method according to any one of claims 1 to 11, characterized in that The first image is a preview image captured by the camera in real time.
13. The method according to claim 12, characterized in that The method further comprises: Correcting the thumbnail corresponding to the preview image based on the at least one first correction parameter to obtain a corrected thumbnail; and / or, The image saved in the album corresponding to the preview image is corrected based on the at least one first correction parameter to obtain a corrected saved image.
14. The method according to claim 1, wherein The acquiring at least one first correction parameter based on the first image and the first shooting parameter includes: performing downsampling processing on the first image; The at least one first correction parameter is acquired based on the downsampled first image and the first shooting parameter.
15. An image processing device, characterized in that: include: a first acquisition module, configured to acquire at least one first correction parameter based on a first image and first shooting parameters, wherein the first image is captured by a camera based on the first shooting parameters, the first image includes a face image, the at least one first correction parameter corresponds one-to-one to at least one pixel in the first image, and the first correction parameter is associated with a shooting distance between the camera and the face; The correction module is used to correct the distorted part of the face image in the first image based on the at least one first correction parameter to obtain a first corrected image.
16. The device according to claim 15, characterized in that The first acquisition module is specifically configured to: The first image and the first shooting parameter are input into an AI network, and a first correction grid is obtained as output by the AI network; wherein at least one grid point in the first correction grid corresponds one-to-one to at least one pixel point in the first image, and each grid point in the first correction grid corresponds to one of the first correction parameters.
17. The device according to claim 16, characterized in that The device further comprises: A second acquisition module is used to obtain a first correction degree parameter; The correction module is specifically used to: Based on the first correction degree parameter, updating a first correction parameter corresponding to each grid point in the first correction grid; Based on the updated first correction grid, the distorted portion of the facial image in the first image is corrected to obtain the first corrected image.
18. The device according to claim 17, characterized in that The second acquisition module is specifically configured to: In response to the received user operation, obtaining the first correction degree parameter; or, The first correction degree parameter is obtained based on the deflection angle of the facial image in the first image.
19. The device according to claim 18, characterized in that The second acquisition module is specifically configured to: determining the first correction degree parameter in response to a received adjustment operation on a correction degree parameter adjustment option in a user interface; displaying the first correction degree parameter in the user interface; The device further comprises: A display module is configured to display the first corrected image in the user interface.
20. The device according to claim 16, wherein The AI network is trained based on multiple training images containing facial images and corresponding shooting parameters, wherein the multiple training images containing facial images correspond to different shooting distances and / or different corresponding shooting parameters.
21. The device according to claim 16, characterized in that The correction module is specifically used to: Based on the facial image, acquiring facial motion information; updating the first correction parameters corresponding to the grid points in the first correction grid based on the facial motion information and the correction parameters corresponding to the grid points in the second correction grid, wherein the second correction grid is obtained based on a previous frame image adjacent to the first image and the shooting parameters corresponding to the previous frame image; Based on the updated first correction grid, the distorted portion of the facial image in the first image is corrected to obtain the first corrected image.
22. The device according to any one of claims 15 to 21, characterized in that The correction module is further used to: Obtaining a depth value of each pixel in the first image; Based on the depth value, determining whether the shooting scene corresponding to the facial image is a planar image; Based on the judgment result, it is determined whether to correct the distorted portion of the facial image in the first image based on the at least one first correction parameter.
23. The device according to any one of claims 15 to 22, characterized in that The shooting distance corresponding to the first image is a first shooting distance; The first acquisition module is further configured to acquire at least one second correction parameter based on a second image and the first shooting parameters, wherein the second image is captured by the camera based on the first shooting parameters, the second image includes a facial image, the shooting distance corresponding to the second image is a second shooting distance, the second shooting distance is less than the first shooting distance, the at least one second correction parameter corresponds one-to-one to at least one pixel in the second image, the degree of distortion of the facial image in the second image is greater than the degree of distortion of the facial image in the first image, and the second correction parameter corresponding to the distorted portion of the facial image in the second image is greater than the first correction parameter corresponding to the distorted portion of the facial image in the first image; The correction module is further configured to correct the distorted portion of the facial image in the second image based on the at least one second correction parameter to obtain a second corrected image.
24. The device according to any one of claims 15 to 23, characterized in that The first correction parameter is used to indicate a correction displacement of a corresponding pixel point in the first image.
25. The device according to any one of claims 15 to 24, characterized in that The first shooting parameter includes a field of view FOV parameter.
26. The device according to any one of claims 15 to 25, characterized in that The first image is a preview image captured by the camera in real time.
27. The device according to claim 26, characterized in that The correction module is further used to: Correcting the thumbnail corresponding to the preview image based on the at least one first correction parameter to obtain a corrected thumbnail; and / or, The image saved in the album corresponding to the preview image is corrected based on the at least one first correction parameter to obtain a corrected saved image.
28. The device according to claim 15, characterized in that The first acquisition module is further configured to: performing downsampling processing on the first image; The at least one first correction parameter is acquired based on the downsampled first image and the first shooting parameter.
29. A graphical user interface on a computer device, characterized in that: The computer device has a display screen, a camera, a memory, and one or more processors for executing one or more instructions stored in the memory, wherein: In response to the received first operation, displaying an image preview interface on the display screen; A first corrected image is displayed in the image preview interface; wherein the first corrected image is obtained by correcting the distorted portion of the facial image in the first image based on at least one first correction parameter; the first image is captured by the camera based on a first shooting parameter, the first image includes a facial image, the at least one first correction parameter corresponds one-to-one to at least one pixel point in the first image, the first correction parameter is obtained based on the first image and the first shooting parameter, and the first correction parameter is associated with the shooting distance between the camera and the face.
30. The graphical user interface according to claim 29, wherein: The image preview interface also includes correction parameter adjustment options, including: In response to receiving a second operation on the correction degree parameter adjustment option, a first correction degree parameter is displayed in the image preview interface, where the first correction degree parameter is used to indicate the correction degree of the distorted part of the facial image by the at least one correction parameter.
31. An image processing device, characterized in that: include: one or more processors; Memory; and one or more computer programs, wherein the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the apparatus performs the method according to any one of claims 1 to 13.
32. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 14.
33. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 14.
34. A chip, characterized in that: The electronic device comprises one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of the electronic device and send the signal to the processor, wherein the signal includes a computer instruction stored in the memory; when the processor executes the computer instruction, the electronic device executes the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Method and apparatus for face image correction
CN105405104A
Image correction method and device and electronic equipment
CN110660034A
Image transformation method and device
CN113850726A
Image processing device, and control method of image processing device
JP2008109305A