Image processing method and related device

By dividing the face images captured by the front camera and targeted image enhancement processing, the problem of poor image quality is solved and a clearer and more delicate face images are achieved.

CN120070225AActive Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311578815.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-30
Estimated Expiration
2043-11-22

AI Technical Summary

Technical Problem

Face images taken by front cameras often have problems with poor image quality or lack of details, such as unclear hair or eyebrows, rough skin texture, which affects the user experience.

Method used

An image processing method is proposed, by dividing the face image area and processing images in different regions using different image enhancement methods. Specifically, the first area (such as eyebrows, hair, eyes) adopts edge and contour processing, and the second area (such as skin) adopts structural processing, thereby improving the clarity and detail fidelity of the image.

Benefits of technology

By performing targeted image enhancement processing on different areas, the clarity and detail fidelity of face images are significantly improved, making hair and eyebrows clear, and the skin delicate and soft, improving image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070225A_ABST
    Figure CN120070225A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and a related device, and is applied to the technical field of terminals. The method comprises the steps of obtaining a first image in response to an operation for triggering photographing; performing region division on the portrait in the first image to obtain a plurality of regions; processing the image of the first area in a first image enhancement mode, and processing the image of the second area in a second image enhancement mode to obtain a second image; wherein the first region and the second region are regions in the plurality of regions, and the definition of the image processed by adopting the first image enhancement mode is different from the definition of the image processed by adopting the second image enhancement mode. Therefore, the image quality of the face image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to an image processing method and related devices. Background Art

[0002] The front camera of an electronic device can capture a user's face image. Currently, problems such as poor image quality or missing image details may occur in the face image captured by the front camera. For example, the hair or eyebrows in the face image are not clear, and the skin texture is rough, which affects the user experience. Summary of the Invention

[0003] Embodiments of this application provide an image processing method and related devices, which are applied to the technical field of terminals. In the embodiments of this application, it is beneficial to improve the image quality of face images.

[0004] In a first aspect, an image processing method is proposed in the embodiments of this application, including: in response to an operation for triggering a photo taking, obtaining a first image; dividing the portrait in the first image into multiple regions; processing the image of the first region in a first image enhancement manner and processing the image of the second region in a second image enhancement manner to obtain a second image; where the first region and the second region are both regions among the multiple regions, and the image clarity after being processed in the first image enhancement manner is different from the image clarity after being processed in the second image enhancement manner.

[0005] The image processing method provided in the embodiments of this application can perform different image enhancements on different regions in the portrait, which is beneficial to perform different processing according to the clarity requirements of different regions and can improve the image quality.

[0006] In a possible implementation, the first image enhancement manner is used to process the edge and / or contour of the image of the first region, and the second image enhancement manner is used to process the structure of the image of the second region.

[0007] The image of the first region has higher requirements for the edge and / or contour, and the image of the second region has higher requirements for the structure. Different processing can be performed on the first region and the second region, which is beneficial to improve the image quality.

[0008] In a possible implementation, the image of the first region includes the image of one or more parts of the eyebrows, hair, or eyes in the portrait, and the image of the second region includes the image of the skin in the portrait. In this way, different processing on different regions can make the hair and eyebrows clear and the face skin delicate and soft, which is beneficial to improve the image quality.

[0009] In a possible implementation, processing the image of the first region in the above first image enhancement manner and processing the image of the second region in the second image enhancement manner to obtain a second image includes: inputting the first image and the first mask image into a first model to obtain the second image. The first mask image is an image obtained by partitioning the portrait in the first image. In the first model, the first mask image is used as a guidance map to extract the features of the image in the first region and the image features in the second region, so as to process the image in the first region in the first image enhancement manner and process the image in the second region in the second image enhancement manner.

[0010] The first mask image is an image obtained by partitioning the portrait in the first image. The first mask image includes multiple regions and can be used as a guidance map to extract the features of the image in the first region and the image features in the second region, so as to process the image in the first region in the first image enhancement manner and process the image in the second region in the second image enhancement manner. In this way, it is beneficial to perform different image enhancement processes on different regions.

[0011] In a possible implementation, the first model is used to extract the first features of the first image based on the information indicated in the first mask image, downsample the first features and then perform feature parallel connection with the features of the first mask image to obtain second features, and based on the second features, process the image in the first region in the first image enhancement manner and process the image in the second region in the second image enhancement manner. In this way, the first model incorporates the features of the first mask image and uses the first mask image as a guidance to perform different image enhancement processes on different regions.

[0012] In a possible implementation, the first model includes a shallow network, an intermediate network, and a deep network; the shallow network is used to extract the first features of the first image based on the information indicated in the first mask image; the intermediate network is used to downsample the first features and then perform feature parallel connection with the features of the first mask image to obtain second features; the deep network is used to downsample the second features; or, the shallow network is used to extract the first features of the first image based on the information indicated in the first mask image; the intermediate network is used to downsample the first features; the deep network is used to downsample the features obtained from the intermediate network and then perform feature parallel connection with the features of the first mask image.

[0013] When the first model includes a shallow network, an intermediate network, and a deep network, the features of the first masked image can be in the intermediate network or the deep network. The features of the first masked image can be feature-parallel combined with the features of the layer where they are located, which is beneficial for the network to better extract the structural and detailed features of the first image, or in other words, is beneficial for the network to better extract the low-frequency and high-frequency features of the input image, so as to facilitate different image enhancement processing for different regions.

[0014] In a possible implementation, the above method further includes: the first model is trained in the following manner: inputting the first image sample and the second masked image into the source model to obtain the output image of the source model. The second masked image is an image obtained by dividing the portrait in the second image sample into regions. The first image sample is obtained by processing the image in the first region of the second image sample in a first blurring manner and processing the image in the second region of the second image sample in a second blurring manner. The clarity of the image processed by the first blurring manner is different from the blurriness of the image processed by the second blurring manner; when the target loss function satisfies the preset model convergence condition, the first model is obtained, where the target loss function is related to the first loss function and the second loss function. The first loss function is the difference between the image in the first region of the output image of the source model and the image in the first region of the second image sample, and the second loss function is the difference between the image in the second region of the output image of the source model and the image in the second region of the second image sample. In this way, different regions correspond to different loss functions, which is beneficial for training different regions separately and beneficial for the trained first model to perform different image enhancement processing on different regions.

[0015] In a possible implementation, the difference between the number of pixel points included in the image of the first region in the first image sample and the number of pixel points included in the image of the second region in the first image sample is less than or equal to a preset value. The number of pixel points included in the image of the first region in the first image sample is the same as or close to the number of pixel points included in the image of the second region in the first image sample, which is beneficial for the model to be able to learn well both in the first region and the second region during the training process and beneficial for reducing the probability that the training effect of one region is good while that of the other region is poor.

[0016] In a possible implementation, the relationship between the target loss function and the first loss function and the second loss function satisfies the following formula:

[0017] L = W 1 *L 1 +W 2 *L 2

[0018] Among them, L is used to represent the target loss function, L i is used to represent the first loss function, L j is used to represent the second loss function, W 1 is used to represent the weight of the first loss function, W 2 is used to represent the weight of the second loss function, W 1 and W 2 are both constants.

[0019] W 1 and W 2 can be the same or different, and the embodiments of the present application do not make any limitations in this regard. If W 1 and W 2 are different, different regions can be trained to different degrees, which is beneficial to improving the training effect of the model.

[0020] In a possible implementation, before obtaining the first image in response to an operation for triggering photographing, the above method further includes: in response to an operation of the user turning on the front camera, displaying a first interface, where the first interface includes an image collected by the front camera and a first control; obtaining the first image in response to an operation for triggering photographing, including: obtaining the first image collected by the front camera in response to an operation of the user triggering the first control.

[0021] The probability of the front camera capturing a face image is relatively high. The embodiments of the present application can perform different image enhancement processes on different regions in the portrait, which is beneficial to obtaining an image with higher quality. In addition, the proportion of the portrait in the face image captured by the front camera in the image is relatively high, which can better divide the portrait into regions and perform different image enhancement processes on different regions, which is beneficial to better exerting the role of the method provided by the embodiments of the present application.

[0022] In a second aspect, embodiments of the present application provide an electronic device, which may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device may be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) electronic device, augmented reality (AR) electronic device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on.

[0023] In a third aspect, embodiments of the present application provide an electronic device, including a processor and a memory. The memory is used to store code instructions, and the processor is used to run the code instructions to execute the method described in any possible implementation manner of the first aspect.

[0024] In a fourth aspect, embodiments of the present application provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method described in any possible implementation manner of the first aspect is implemented.

[0025] In a fifth aspect, embodiments of the present application provide a computer program product including a computer program. When the computer program is run, the computer is caused to execute the method described in any possible implementation manner of the first aspect.

[0026] In a sixth aspect, embodiments of the present application provide a chip or a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instruction to execute the method described in any possible implementation manner of the first aspect. Among them, the communication interface in the chip may be an input / output interface, a pin, a circuit, etc.

[0027] In a possible implementation, the chip or chip system described above in the embodiments of the present application further includes at least one memory, and instructions are stored in the at least one memory. The memory may be a storage unit inside the chip, for example, a register, a cache, etc., or may be a storage unit of the chip (for example, a read-only memory, a random access memory, etc.).

[0028] It should be understood that the technical solutions of the second to sixth aspects of the embodiments of the present application correspond to those of the first aspect of the embodiments of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram of capturing a face image;

[0030] Figure 2 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application;

[0031] Figure 3 is a schematic software structure block diagram of an electronic device provided by an embodiment of the present application;

[0032] Figure 4 is a schematic diagram of an image processing method provided by an embodiment of the present application;

[0033] Figure 5 is a schematic diagram of an input image and a first mask image provided by an embodiment of the present application;

[0034] Figure 6 is a schematic diagram of another image processing method provided by an embodiment of the present application;

[0035] Figure 7 is a schematic diagram of yet another image processing method provided by an embodiment of the present application;

[0036] Figure 8 is a schematic diagram of a chip provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To facilitate a clear description of the technical solutions of the embodiments of the present application, the following explanations are made first:

[0038] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first image and the second image are only used to distinguish different images, and no limitation is imposed on their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily mean different.

[0039] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0040] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0041] It should be noted that "when... " in the embodiments of the present application can be at the instant when a certain situation occurs, or within a period of time before and after a certain situation occurs. The embodiments of the present application do not make specific limitations on this. In addition, the display interface provided in the embodiments of the present application is only an example, and the display interface can also include more or less content.

[0042] The front camera of the electronic device can capture a user's face image. Since the distance between the user and the front camera is relatively close, the proportion of the user's face in the captured image is relatively high, and high requirements are placed on the clarity and texture details of the face. Currently, problems such as unclear hair or eyebrows and rough skin texture may occur in the face image captured by the front camera, affecting the user experience.

[0043] Exemplarily, Figure 1 shows a schematic diagram of capturing a face image. As Figure 1 shown, in response to the user's operation of turning on the front camera, the electronic device can display Figure 1 interface a in Figure 1 . As shown in interface a in Figure 1 , the interface includes an image captured by the front camera, a photographing control 101, and a control 102 for viewing the captured image. In response to the user's operation of triggering the photographing control 101, the electronic device captures an image through the front camera and can display Figure 1As shown in the b interface in [Figure 0], the interface includes a thumbnail 103 of an image captured by the front camera. The user can trigger a control 102 for viewing the captured image to view the image captured by the front camera. In response to the user's operation of triggering the control 102 for viewing the captured image, the electronic device can display Figure 1 the c interface in [Figure 0]. As Figure 1 shown in the c interface in [Figure 0], the interface includes a face image captured by the front camera. In the face image captured by the front camera, the eyebrows are unclear and the skin is rough.

[0044] The reason for such a problem is that when processing the image captured by the front camera, the same denoising process is performed on the entire image, resulting in the same denoising intensity for the face, hair, and background, causing problems such as unclear hair or eyebrows and rough skin texture in the face image, reducing the image quality, and affecting the user experience. Herein, the background refers to the part of the image other than the portrait.

[0045] In view of this, the embodiments of the present application provide an image processing method and related device, which can perform different degrees of processing on the facial features, hair, and background of the face in the face image, making the hair and eyebrows clear and the facial skin delicate and smooth, which is beneficial to improving the image quality.

[0046] The image processing method provided by the embodiments of the present application can be applied to an electronic device. For example, devices such as mobile phones or tablets. For ease of understanding, the hardware structure of the electronic device in the embodiments of the present application will be introduced below.

[0047] Exemplarily, Figure 2 is a schematic diagram of the hardware structure of an electronic device provided by the embodiments of the present application. As Figure 2 shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, and a display screen 194, etc.

[0048] Optionally, the above-mentioned sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0049] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0050] The camera 193 may include a front camera and a rear camera, and both the front camera and the rear camera can capture facial images. The processor 110 responds to the user's triggering operation of capturing an image, obtains an image captured by the front camera or the rear camera, processes the image by the method provided in the embodiment of the present application, and stores the processed image in the internal memory 121.

[0051] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The layered architecture can adopt an Android system, an Apple (IOS) system, or other operating systems, which are not limited in the embodiments of the present application. The following takes the Android system of the layered architecture as an example to exemplify the software architecture of the electronic device provided in the embodiments of the present application.

[0052] Figure 3 A schematic diagram of the software architecture of an electronic device provided in an embodiment of the present application. The layered architecture divides the software system of the electronic device into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system can be divided into four layers, from top to bottom: application layer (applications), application framework layer (application framework), hardware abstraction layer (hardware abstraction layer, HAL), kernel layer (kernel) and hardware layer. The application layer can include a series of application packages, and the application layer runs applications by calling the application programming interface (application programming interface, API) provided by the application framework layer. Figure 3 As shown, the application package may include applications such as camera and gallery.

[0053] The application framework layer provides APIs and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, Figure 3 as shown, the application framework layer may include a camera access interface and a view system. Among them, the camera access interface can be used to provide an application programming interface and a programming framework for the camera application. The view system includes visual controls, for example, it may include controls for taking pictures (such as the picture-taking control 101 shown in interface a above Figure 1 and controls for viewing the captured images (such as the control 102 for viewing the captured images shown in interface a above Figure 1 ).

[0054] The purpose of the HAL layer is to abstract the hardware, which can provide a unified interface for querying hardware devices for the upper-layer applications, or can also provide data storage services for the upper-layer applications. As Figure 3 shown, the HAL layer may include a camera hardware abstraction layer and a camera algorithm library. Among them, the camera hardware abstraction layer can provide virtual hardware for the camera device. The camera algorithm library may include the running code and data for implementing the image processing method provided in the embodiments of the present application.

[0055] The kernel layer is the layer between the hardware and the software. As Figure 3 shown, the kernel layer may include one or more of the following: a camera device driver, a digital signal processor driver, and an image processor driver, etc. Among them, the camera device driver is used to drive the sensor of the camera to collect images and drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process the images. The image processor driver is used to drive the graphics processor to process the images.

[0056] The hardware layer may include hardware such as a camera and a display screen.

[0057] It should be understood that in some embodiments, the layers that implement the same function may be called other names, or the layer that can implement the functions of multiple layers may be regarded as one layer, or the layer that can implement the functions of multiple layers may be divided into multiple layers. The embodiments of the present application do not limit this.

[0058] Next, in combination with the software structure shown above Figure 3 the image processing method in the embodiments of the present application will be specifically described:[[]]END]]

[0059] In response to an operation by the user to open the camera application, such as an operation of clicking on the camera application icon, the camera application calls the camera access interface in the application framework layer to start the camera application, and then sends an instruction to start the camera by calling the camera device in the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the kernel layer. This camera device driver can start the corresponding camera to capture images. In response to an operation by the user to trigger a photo-taking, the camera device driver can control the camera to capture images and obtain the original images.

[0060] The camera can transmit the captured original images to the camera hardware abstraction layer through the camera device driver. The camera hardware abstraction layer can send the original images to the camera algorithm library.

[0061] The camera algorithm library stores program codes for implementing the image processing method provided in the embodiments of the present application. Based on a digital signal processor and an image processor, the camera algorithm library executes the above codes to implement the different degrees of processing of the facial features, hair, and background in the face image as introduced above, making the hair and eyebrows clear and the facial skin delicate and soft, and obtaining images of relatively high quality. The camera algorithm library can send the images of relatively high quality to the camera hardware abstraction layer. The camera hardware abstraction layer can transmit them to the display screen for display.

[0062] The above Figure 2 and Figure 3 introduced the electronic device in the embodiments of the present application. Next, the method provided in the embodiments of the present application will be introduced.

[0063] Figure 4 shows a schematic diagram of an image processing method provided in the embodiments of the present application. As Figure 4 shown, the method includes the following steps:

[0064] S401. The electronic device processes the input image to obtain a first mask image.

[0065] The input image can be an image obtained through a camera. The input image may or may not include a portrait, and the embodiments of the present application do not make any limitations.

[0066] There are various different implementation manners for the electronic device to obtain the input image.

[0067] In a possible implementation manner, the electronic device may include a front camera and a rear camera, and the electronic device can obtain the input image through the front camera or the rear camera. The electronic device can obtain one or more frames of input images captured by the front camera or the rear camera, and the embodiments of the present application do not make any limitations in this regard.

[0068] In some scenarios, an electronic device is installed with a camera application. The camera application includes a photographing control. In response to a user's operation of opening the camera application, the front camera or the rear camera is turned on. In response to a user's operation of triggering the photographing control, an input image can be obtained through the front camera or the rear camera.

[0069] In another possible implementation, the electronic device can receive an image obtained through a camera from another device. In this way, the electronic device may not include a camera and can have the function of processing images.

[0070] The first mask image can also be referred to as a mask map. The first mask image is obtained by dividing the input image into regions. The first mask image can include multiple regions. Different regions among these multiple regions do not overlap. Each region among these multiple regions can include one or more parts. The clarity requirements for the parts included in different regions can be different, and the clarity requirements for the parts included in the same region can be the same or similar. The number of regions and the parts included in the regions can be preset.

[0071] Exemplarily, if the clarity requirements for hair, eyebrows, and eyes are the same or similar, and the clarity requirements for hair, eyebrows, and eyes are different from those for the skin, then the first mask image can include a first region and a second region. The first region includes hair, eyebrows, and eyes, and the second region includes the skin.

[0072] In some implementations, the electronic device processes the input image to obtain the background and the portrait of the input image, segments the portrait to obtain each part of the portrait, and divides each part into regions to obtain the first mask image.

[0073] Exemplarily, the process by which the electronic device processes the input image to obtain the first mask image can include: The electronic device processes the input image based on a portrait segmentation model to obtain the region where the portrait is located and the region where the background is located, and processes the portrait in the input image based on a face segmentation model to obtain the regions where hair, eyebrows, eyes, ears, nose, mouth, and skin are located, and divides each part into regions to obtain the first mask image. The embodiments of the present application do not specifically limit the segmentation model used to obtain the first mask image. The number of segmentation models can be one or multiple, and the embodiments of the present application do not limit this.

[0074] In a possible implementation, the first mask image includes one or more of the following regions: the region where the face is located, the region where the background is located, the region where the hair is located, the region where the eyebrows are located, the region where the eyes are located, the region where the ears are located, the region where the nose is located, the region where the mouth is located, or the region where the skin is located. The skin may include the skin at one or more positions of the cheeks, forehead, or neck. In some implementations, the skin may also include the skin at the position of the nose, which is not limited in the embodiments of the present application.

[0075] Exemplarily, Figure 5 shows a schematic diagram of an input image and a first mask image. As Figure 5 shown, the electronic device processes the input image to obtain a first mask image divided into multiple regions. The first mask image includes the region where the face is located, the region where the background is located, the region where the hair is located, the region where the eyebrows are located, the region where the eyes are located, the region where the ears are located, the region where the nose is located, the region where the mouth is located, or the region where the skin is located.

[0076] In some examples, in response to the user's operation of triggering the photographing control, multiple frames of face images may be obtained through the front camera or the rear camera, and the terminal device may process these multiple frames of face images to obtain a frame of the first mask image.

[0077] S403. The electronic device inputs the input image and the first mask image into an image processing model to obtain a first output image, where the first output image is used to represent the image output by the image processing model.

[0078] The image processing model may also be referred to as an image generation model, which is not limited in the embodiments of the present application. The image processing model is used to perform different enhancement processes on different regions of the portrait in the input image. For example, the portrait in the input image may be divided into multiple regions, and there are no overlapping regions in each of these multiple regions. The image processing model may perform image enhancement processes with different degrees on different regions to obtain an image with higher quality. If one region among these multiple regions includes hair, eyebrows, and eyes, and another region includes skin, the image processing model may perform image enhancement with different degrees on these two regions, so that the hair and eyebrows in the first output image are clear, and the facial skin is delicate and soft.

[0079] In some implementations, as Figure 4As shown in the figure, the image processing model provided by the embodiment of the present application may include five layers of networks. The first layer of network and the fifth layer of network may be the shallow layers of the image processing model, the second layer of network and the fourth layer of network are the middle layers of the image processing model, and the third layer of network is the deep layer of the image processing model. The first layer of network and the fifth layer of network may be skip connections, and the second layer of network and the fourth layer of network may be skip connections. From the first layer of network to the third layer of network, it can be understood as an encoding process, and from the third layer of network to the fifth layer of network, it can be understood as a decoding process. From the first layer of network to the second layer of network, and from the second layer of network to the third layer of network, both can be understood as a downsampling process or a dimensionality reduction process. From the third layer of network to the fourth layer of network, and from the fourth layer of network to the fifth layer of network, both can be understood as an upsampling process or a dimensionality increase process.

[0080] In some examples, the resolution of the image obtained by the first layer of network may be 64*256*256, the resolution of the image obtained by the second layer may be 128*128*128, the resolution of the image obtained by the third layer may be 256*64*64, the resolution of the image obtained by the fourth layer may be 128*128*128, and the resolution of the image obtained by the fifth layer may be 64*256*256.

[0081] In the image generation model of the embodiment of the present application, the features of the first mask image are imported. The import method may be one or more of an attention mechanism, a spatial feature transformation, or a sparse fast Fourier transform (SFT). The embodiment of the present application does not limit the specific implementation method of the import. The features of the first mask image may be in one or more of the first layer of network to the fifth layer of network. In Figure 4 the example shown, the features of the first mask image are in the second layer of network and the fourth layer of network.

[0082] In Figure 4 the example shown, the first layer of network may be used to perform convolution on the first mask image and the input image to extract the first feature from the input image based on the information indicated by the first mask image. The second layer of network imports the features of the first mask image and may be used to obtain the first feature from the first layer of network, and after downsampling the first feature, perform feature concatenation (concat or confloat) with the features of the first mask image to obtain the second feature. The third layer of network may be used to obtain the second feature from the second layer of network and perform downsampling on the second feature. The fourth layer of network may be used to perform upsampling on the features obtained by the third layer of network. The fifth layer of network may be used to perform upsampling on the features obtained by the fourth layer of network.

[0083] The features of the first mask image can be in the middle layer or the deep layer. The features of the first mask image can be concatenated (concat or confloat) with the features of the layer where it is located, which is beneficial for the network to better extract the structural and detailed features of the input image, or rather, beneficial for the network to better extract the low-frequency and high-frequency features of the input image.

[0084] The image processing method provided by the embodiments of the present application processes the input image to obtain the first mask image. When the image processing model extracts the features in the input image, guided by the first mask image, it is beneficial for the image processing model to better extract the features of different regions in the portrait, and then different regions of the portrait, such as facial features, hair, and background, can be processed to different degrees, making the hair and eyebrows clear, and the facial skin delicate and soft, which is beneficial for obtaining an image with higher quality.

[0085] Optionally, before the electronic device executes S403, the input image and the first mask image can be respectively sliced, and the sliced input image and the sliced first mask image are input into the image processing model to obtain the sliced first output image, and the sliced first output image is fused to obtain the first output image. In this way, it is beneficial to improve the processing efficiency.

[0086] Exemplarily, the resolutions of the input image and the first mask image can both be 2000*3000. The electronic device can divide the input image into 50 blocks to obtain an input image including 50 blocks, and divide the first mask image into 50 blocks to obtain a first mask image including 50 blocks. The electronic device can input the input image including 50 blocks and the first mask image including 50 blocks into the first model, and the first model can output a first output image including 50 blocks. The electronic device can fuse the first output image including 50 blocks to obtain the first output image.

[0087] The above Figure 4 The image processing model involved is obtained based on training samples. The training process of the image processing model will be introduced in detail below.

[0088] Exemplarily, Figure 6 shows a schematic diagram of an image processing method provided by the embodiments of the present application. This method can be executed by an electronic device or by a server, and the embodiments of the present application do not make a limitation on this. As Figure 6 shown, this method includes the following steps:

[0089] S601. Process the face image with clear texture to obtain the second mask image.

[0090] A face image with clear texture is used to represent a face image with high image quality. The face image with clear texture can be captured by a high-definition camera. For example, the face image with clear texture can be captured by a single-lens reflex camera.

[0091] Processing the face image with clear texture to obtain a second mask image can be understood as dividing the face image with clear texture into regions to obtain the second mask image.

[0092] In some examples, the electronic device processes the face image with clear texture based on a portrait segmentation model to obtain the region where the portrait is located and the region where the background is located, and processes the portrait in the face image with clear texture based on a face segmentation model to obtain each part of the portrait, and divides each part into regions to obtain the second mask image. The segmentation model used in this step may be the same as or different from the segmentation model used in S402 above, and the embodiments of the present application do not limit this.

[0093] The second mask images corresponding to different face images with clear texture may include the same regions or different regions, and the embodiments of the present application do not limit this.

[0094] Exemplarily, if the second mask images corresponding to different face images with clear texture include the same regions, when the number of face images with clear texture is 10,000, the second mask images corresponding to these 10,000 images all include regions such as the face, background, hair, eyebrows, eyes, ears, nose, mouth, and skin.

[0095] If the second mask images corresponding to different face images with clear texture include different regions, when the number of face images with clear texture is 10,000, among these 10,000 images, there may be 1,000 images whose corresponding second mask images all include regions such as the face, background, hair, eyebrows, eyes, ears, nose, mouth, and skin, there may be 4,000 images whose corresponding second mask images all include the regions where the eyebrows and hair are located, there may be 2,000 images whose corresponding second mask images all include the region where the eyebrows are located, there may be 2,000 images whose corresponding second mask images all include the region where the hair is located, and there may be 1,000 images whose corresponding second mask images all include the regions where the eyebrows, hair, and skin are located.

[0096] S602: Based on the second mask image, different degrees of degradation can be performed on different regions in the face image with clear texture to obtain a blurred image.

[0097] A higher degree of degradation can be referred to as strong degradation, which can be understood as a higher degree of blurriness. A lower degree of degradation can be called weak degradation, which can be understood as a lower degree of blurriness. Strong degradation can be performed on parts that require high texture clarity, while weak degradation can be done on parts that require delicacy and softness or non-key areas of concern.

[0098] Exemplarily, parts that require high texture clarity can include eyebrows, hair, and eyes. For the areas of eyebrows, hair, and eyes in a face image with clear texture, strong degradation can be carried out to increase the degree of blur. Parts that require delicacy and softness can include the skin. For the area of the skin in a face image with clear texture, weak degradation can be performed to reduce the degree of blur. Non-key areas of concern can include the background. For the background area in a face image with clear texture, weak degradation can be done.

[0099] Performing different degrees of degradation on different regions in a face image with clear texture is beneficial for training the model to perform different degrees of image enhancement on different regions.

[0100] S603: Input the second mask image and the blurred image into the initial model to obtain a second output image.

[0101] It can be understood that the initial model is the model before the above image processing model is trained.

[0102] The first layer network in the initial model can perform convolution on the second mask image and the blurred image to obtain Feature 1. Feature 1 is used to represent the features extracted from the face image based on the indication of the second mask image, and Feature 1 can be transmitted to the second layer network of the initial model. The second layer network in the initial model imports the features of the second mask image. The second layer network in the initial model can perform downsampling on the extracted Feature 1 and then perform feature parallel connection with the features of the second mask image, and transmit it to the third layer network in the initial model. The third layer network in the initial model reduces the dimension of the obtained features and transmits them to subsequent layers for decoding to obtain the second output image.

[0103] In the initial model, importing the features of the second mask image and using the second mask image as a guidance map for model training is beneficial for the model to utilize the information in the second mask image to perform image enhancement on images in different regions of the blurred image.

[0104] In the blurred image, if the number of pixel points occupied by a region is small, the number of blurred images of that region is large. If the number of pixel points occupied by a region is large, the number of blurred images of that region is small. This can make the number of pixel points included in images of different regions the same or similar.

[0105] Exemplarily, the eyebrows occupy fewer pixel points in the human face, so the blurred images input to the initial model can include more blurred images including eyebrows; the skin occupies more pixel points in the human face, so the blurred images input to the initial model can include fewer blurred images including skin.

[0106] In this way, the pixel points included in the images of different regions are the same or similar, which is conducive to reducing the probability that the training effect of one region is good and the training effect of another region is poor.

[0107] It should be noted that the above S601 and S602 are optional. Those skilled in the art can obtain the second mask image and the blurred image through other devices, and input these data into the electronic device, and the electronic device can input these data into the initial model for training.

[0108] S604. Calculate the loss function.

[0109] The electronic device can separately compare each region included in the second output image with each region of the texture-clear face image, and calculate the loss function. Each region can correspond to a loss value. The loss function is related to the loss values of each region, and different loss values can correspond to different weights. Among them, the calculation methods of the loss values corresponding to different regions can be the same or different, and the embodiments of the present application do not limit this. If the calculation methods of the loss values corresponding to different regions can be different, it is beneficial to obtain images with higher quality.

[0110] Exemplarily, a region includes hair and eyebrows. The calculation method of the loss value corresponding to the region where the hair and eyebrows are located can be calculated using the high-frequency loss method. For example, the electronic device can calculate the loss value corresponding to the region where the hair and eyebrows are located through one or more of generative adversarial networks (GAN), visual geometry group (VGG), or gradient. Among them, the gradient calculation method refers to calculating the gradient of two images and using the gradient to calculate the loss function, which is a simplified calculation method of VGG. Another region includes skin. The calculation method of the loss value corresponding to the region where the skin is located can be calculated using the low-frequency loss method. For example, the electronic device can calculate the loss value corresponding to the region where the skin is located through the absolute error loss function (L1). In this way, the high-frequency loss is beneficial for the model to learn the image edge and image contour to make the hair and eyebrows clearer, and the low-frequency loss is beneficial for the model to learn the structure of the image to make the skin smoother and softer.

[0111] In some implementations, the weight of the area where the eyebrows and hair are located is W 1 , and the loss value of the area where the eyebrows and hair are located is L 1 , the weight of the area where the skin is located is W 2 , and the loss value of the area where the skin is located is L 2 , the loss function can be W 1 *L 1 +W 2 *L 2 . The high-frequency loss is more difficult to learn than the low-frequency loss. Therefore, the weight of the area where the eyebrows and hair are located can be greater than the weight of the area where the skin is located, which is beneficial to improving the accuracy of model training.

[0112] S605. If the loss function does not meet the training requirements, update the parameters in the initial model, and repeat the above S603 and S604 until the loss function meets the training requirements.

[0113] The image processing method provided by the embodiments of the present application processes the face image to obtain a second mask image. When the image processing model extracts features in the face image, guided by the second mask image, it is beneficial for the image processing model to extract different features from different regions, and then different degrees of image enhancement processing can be performed on different regions in the portrait, such as facial features, hair, and background, making the hair and eyebrows clear and the facial skin delicate and soft, which is beneficial to obtaining an image with higher quality.

[0114] The training process of the model provided by the embodiments of the present application is described above. Next, the image processing method of the embodiments of the present application will be described in detail in combination with the step flow.

[0115] Exemplarily, Figure 7 shows a schematic flowchart of an image processing method provided by the embodiments of the present application. As Figure 7 shown, the method includes:

[0116] S701. In response to an operation for triggering a photo taking, obtain a first image.

[0117] The electronic device may include a camera, and the first image may be obtained by the camera of the electronic device in response to an operation for triggering a photo taking. The first image may be obtained by the front camera or the rear camera. The embodiments of the present application do not make any limitation thereto.

[0118] The first image may, for example, refer to the input image in the example shown above Figure 4 .

[0119] S702. Divide the portrait in the first image into multiple regions.

[0120] There is no overlapping part between each of these multiple regions. Parts of the portrait with the same or similar clarity requirements can be in the same region.

[0121] In some implementations, these multiple regions may include two regions, one region including hair, eyebrows, and eyes, and the other region including skin.

[0122] In other implementations, these multiple regions may include three regions, the first region including the background, the second region including hair, eyebrows, and eyes, and the third region including skin.

[0123] In still other implementations, these multiple regions may include four regions, the first region including the background, the second region including hair, eyebrows, and eyes, the third region including skin, and the fourth region including the region other than the background, hair, eyebrows, eyes, and skin.

[0124] S703. Process the image of the first region in a first image enhancement manner and process the image of the second region in a second image enhancement manner to obtain a second image; wherein, both the first region and the second region are regions among the multiple regions, and the clarity of the image processed by the first image enhancement manner is different from the clarity of the image processed by the second image enhancement manner.

[0125] It can be understood that in the embodiments of the present application, the first region and the second region are taken as examples for illustration. The electronic device may also, while processing the images of the first region and the second region, process the image of the third region in a third image enhancement manner.

[0126] The image processing method provided by the embodiments of the present application can perform different image enhancements on different regions in the portrait, which is beneficial to performing different processes according to the clarity requirements of different regions and can improve the image quality.

[0127] Optionally, the first image enhancement manner is used to process the edge and / or the contour of the image of the first region, and the second image enhancement manner is used to process the structure of the image of the second region.

[0128] The first image enhancement manner processes the edge and / or the contour of the image of the first region, which can be understood as the first image enhancement manner processes the high-frequency details of the image of the first region. The second image enhancement manner processes the structure of the image of the second region, which can be understood as the first image enhancement manner processes the low-frequency details of the image of the first region. If the image of the third region is processed by a third image enhancement manner, the third image enhancement manner can process one or more of the texture, color, shape, or contrast of the image of the third region.

[0129] The images in the first region have a higher requirement for high-frequency details, and the images in the second region have a higher requirement for low-frequency details. Different processing can be performed on the first region and the second region, which is beneficial to improving the image quality.

[0130] Optionally, the images in the first region include images of one or more parts of the eyebrows, hair, or eyes in a portrait, and the images in the second region include images of the skin in a portrait.

[0131] One or more parts of the eyebrows, hair, or eyes have a higher requirement for clarity. Clear skin will result in rough texture, so the skin has a lower requirement for clarity. Performing different processing on different regions can make the hair and eyebrows clear and the facial skin delicate and soft, which is beneficial to improving the image quality.

[0132] Optionally, processing the images in the first region by the first image enhancement method and processing the images in the second region by the second image enhancement method to obtain a second image includes: inputting the first image and the first mask image into a first model to obtain the second image. The first mask image is an image obtained by dividing the regions of the portrait in the first image. In the first model, the first mask image is used as a guidance map to extract the features of the images in the first region and the features of the images in the second region, so as to process the images in the first region by the first image enhancement method and process the images in the second region by the second image enhancement method.

[0133] The first model can perform different image enhancement processing on different regions of the image. In the above Figure 4 shown example, the first model can be an image processing model, the first image can be an input image, and the first mask image can be obtained through step S402. The second image can be the first output image. In the above Figure 4 In, the structure of the first model is a Unet network structure, which is a possible implementation manner and is not limited here.

[0134] The first mask image is an image obtained by dividing the regions of the portrait in the first image. The first mask image includes multiple regions and can be used as a guidance map to extract the features of the images in the first region and the features of the images in the second region, so as to process the images in the first region by the first image enhancement method and process the images in the second region by the second image enhancement method. In this way, it is beneficial to perform different image enhancement processing on different regions.

[0135] Optionally, the first model is used to extract the first feature of the first image based on the information indicated in the first masked image, downsample the first feature, and then perform feature concatenation with the feature of the first masked image to obtain a second feature. Based on the second feature, image processing of the first region is performed in a first image enhancement manner, and image processing of the second region is performed in a second image enhancement manner.

[0136] The first masked image includes multiple regions. The first model can extract the features of the images of the respective regions in the first image based on the positions of the respective regions in the first masked image to obtain the first feature, and can downsample the first feature, and then perform feature concatenation on the downsampled first feature and the feature of the first masked image to obtain a second feature, so as to implement different image enhancement processing for different regions. Herein, downsampling may also be referred to as dimensionality reduction, and the embodiments of the present application do not make any limitation thereto.

[0137] In this way, the first model incorporates the features of the first masked image and uses the first masked image as a guide to perform different image enhancement processing on different regions.

[0138] Optionally, the first model includes a shallow network, an intermediate network, and a deep network; the shallow network is used to extract the first feature of the first image based on the information indicated in the first masked image; the intermediate network is used to downsample the first feature and then perform feature concatenation with the feature of the first masked image to obtain a second feature; the deep network is used to downsample the second feature; or, the shallow network is used to extract the first feature of the first image based on the information indicated in the first masked image; the intermediate network is used to downsample the first feature; the deep network is used to downsample the feature obtained from the intermediate network and then perform feature concatenation with the feature of the first masked image.

[0139] When the first model includes a shallow network, an intermediate network, and a deep network, the feature of the first masked image may be in the intermediate network or the deep network. The feature of the first masked image can be concatenated with the feature of the layer where it is located, which is beneficial for the network to better extract the structural and detailed features of the first image, or rather, beneficial for the network to better extract the low-frequency and high-frequency features of the input image, so as to implement different image enhancement processing for different regions.

[0140] Optionally, the above method further includes: The first model is trained as follows: The first image sample and the second mask image are input into the source model to obtain the output image of the source model. The second mask image is an image obtained by partitioning the portrait in the second image sample. The first image sample is obtained by processing the image of the first region in the second image sample in a first blurring manner and processing the image of the second region in the second image sample in a second blurring manner. The clarity of the image processed by the first blurring manner is different from the blurriness of the image processed by the second blurring manner. When the target loss function meets the preset model convergence condition, the first model is obtained. The target loss function is related to the first loss function and the second loss function. The first loss function is the difference between the image of the first region in the output image of the source model and the image of the first region in the second image sample. The second loss function is the difference between the image of the second region in the output image of the source model and the image of the second region in the second image sample.

[0141] In the above Figure 6 In the example shown, the second image sample can be a face image with clear texture, the first image sample can be a blurred image obtained through S602, the second mask image can be obtained through S601, the source model can be an initial model, and the output image of the source model is the second output image. The blurring manner can be understood as degradation. The first image sample is obtained by degrading the images of different regions in the second image sample to different degrees. The target loss function can be the loss function shown above Figure 6 . Each region in the images of different regions can correspond to a loss function, and the target loss function can be related to the loss functions corresponding to these regions. The preset model convergence condition can be referred to as the training requirement, which is not limited in the embodiments of the present application.

[0142] In this way, different regions correspond to different loss functions, which is beneficial to separately training for different regions and beneficial to enabling the first model obtained by training to perform different image enhancement processing on different regions.

[0143] Optionally, the difference between the number of pixel points included in the image of the first region in the first image sample and the number of pixel points included in the image of the second region in the first image sample is less than or equal to a preset value.

[0144] The preset value is used to represent a critical value. When the difference between the number of pixels contained in the image of the first region in the first image sample and the number of pixels included in the image of the second region in the first image sample is less than or equal to the preset value, it can be stated that the number of pixels contained in the image of the first region in the first image sample is the same as or close to the number of pixels included in the image of the second region in the first image sample. When the difference between the number of pixels contained in the image of the first region in the first image sample and the number of pixels included in the image of the second region in the first image sample is greater than the preset value, it can be stated that the number of pixels contained in the image of the first region in the first image sample is quite different from the number of pixels included in the image of the second region in the first image sample.

[0145] The number of pixels contained in the image of the first region in the first image sample being the same as or close to the number of pixels included in the image of the second region in the first image sample is conducive to enabling the model to learn well both in the first region and the second region during the training process, and is conducive to reducing the probability that the training effect of one region is good while that of the other region is poor.

[0146] Optionally, the relationship between the target loss function and the first loss function and the second loss function satisfies the following formula:

[0147] L = W 1 *L 1 +W 2 *L 2

[0148] Wherein, L is used to represent the target loss function, L i is used to represent the first loss function, L j is used to represent the second loss function, W 1 is used to represent the weight of the first loss function, W 2 is used to represent the weight of the second loss function, W 1 and W 2 are both constants.

[0149] W 1 and W 2 can be the same or different, and the embodiments of the present application do not make any limitations in this regard. If W 1 and W 2 are different, different regions can be trained to different degrees, which is conducive to improving the training effect of the model.

[0150] In some implementations, for the image of the first region, learn the edge and / or the contour of the image of the first region; for the image of the second region, learn the structure of the image of the second region. Then, for the weight W of the loss function corresponding to the first region 1It can be greater than the weight W of the loss function corresponding to the second region 2 Edges and contours are high-frequency features, while structures are low-frequency features. High-frequency features are more difficult to learn than low-frequency features, and a larger weight can be beneficial to improving the training effect of the model.

[0151] Optionally, before obtaining the first image in response to an operation for triggering a photo, the above method further includes: in response to an operation by the user to turn on the front camera, displaying a first interface, where the first interface includes an image captured by the front camera and a first control; obtaining the first image in response to an operation for triggering a photo, including: in response to an operation by the user to trigger the first control, obtaining the first image captured by the front camera.

[0152] The first interface can be as shown in the a interface above Figure 1 The first image captured by the front camera can be as shown in the a interface above Figure 1 The first control can be Figure 1 The photo-taking control 101 in the above

[0153] For the scenario of performing different image enhancement processes on different regions of an image, it can be a scenario of capturing an image through the front camera. The probability of the front camera capturing a face image is relatively high. In the embodiments of the present application, different image enhancement processes can be performed on different regions in a portrait, which is beneficial to obtaining an image with higher quality. Additionally, in the face image captured by the front camera, the proportion of the portrait in the image is relatively high, which can better divide the portrait into regions and perform different image enhancement processes on different regions, facilitating the greater exertion of the role of the method provided in the embodiments of the present application.

[0154] The method of the embodiments of the present application has been described above. Next, the device for executing the above method provided by the embodiments of the present application will be described. Those skilled in the art can understand that the method and the device can be combined and referenced with each other, and the related device provided by the embodiments of the present application can execute the steps in the above method.

[0155] To implement the above functions, the device for implementing the above image processing method includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the method steps of the examples described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0156] The embodiments of the present application can divide the device for implementing the image processing method into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0157] Figure 8 It is a schematic structural diagram of a chip provided by an embodiment of the present application. As Figure 8 shown, the chip 80 includes one or more (including two) processors 801, a communication line 802, a communication interface 803, and a memory 804.

[0158] In some embodiments, the memory 804 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.

[0159] The above-described image processing method described in the embodiments of the present application can be applied to the processor 801 or implemented by the processor 801. The processor 801 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above image processing method can be completed by the hardware integrated logic circuit or software-form instructions in the processor 801. The above processor 801 may be a general-purpose processor (for example, a microprocessor or a conventional processor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The processor 801 can implement or execute the various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0160] The steps of the image processing method disclosed in the embodiments of the present application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. Among them, the software module can be located in mature storage media in the art such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage media is located in the memory 804, and the processor 801 reads the information in the memory 804 and combines its hardware to complete the steps of the above method.

[0161] The processor 801, the memory 804, and the communication interface 803 can communicate with each other through the communication line 802.

[0162] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory in advance, or downloaded and installed in the memory in the form of software.

[0163] The embodiments of the present application also provide a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by a computer, or a data storage device such as a server or a data center including one or more available media integrated. For example, the available medium can include magnetic media (such as floppy disks, hard disks, or magnetic tapes), optical media (such as digital versatile discs (DVDs)), or semiconductor media (such as solid state disks (SSDs)).

[0164] The embodiments of the present application also provide a computer-readable storage medium. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. A computer-readable medium may include a computer storage medium and a communication medium, and may also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0165] As a possible design, the computer-readable medium may include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; the computer-readable medium may include a magnetic disk storage or other magnetic disk storage devices. Moreover, any connecting line can also be appropriately referred to as a computer-readable medium. For example, if software is transmitted using coaxial cables, fiber optic cables, twisted pairs, DSL, or wireless technologies (such as infrared, radio, and microwave) from a website, server, or other remote source, then the coaxial cables, fiber optic cables, twisted pairs, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, magnetic disks and optical discs include optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks usually reproduce data magnetically, while optical discs use lasers to optically reproduce data.

[0166] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

Claims

1. An image processing method, characterized in that, it includes: responding to an operation for triggering a photo taking, and acquiring a first image; dividing the region of the portrait in the first image to obtain multiple regions; processing the image of the first region in a first image enhancement manner and processing the image of the second region in a second image enhancement manner to obtain a second image; wherein, both the first region and the second region are regions among the multiple regions, and the clarity of the image processed by the first image enhancement manner is different from the clarity of the image processed by the second image enhancement manner.

2. The method according to claim 1, characterized in that, the first image enhancement manner is used to process the edge and / or the contour of the image of the first region, and the second image enhancement manner is used to process the structure of the image of the second region.

3. The method according to claim 1 or 2, characterized in that, the image of the first region includes the image of one or more parts of eyebrows, hair or eyes in the portrait, and the image of the second region includes the image of the skin in the portrait.

4. The method according to any one of claims 1 to 3, characterized in that, the processing the image of the first region in a first image enhancement manner and processing the image of the second region in a second image enhancement manner to obtain a second image includes: inputting the first image and a first mask image into a first model to obtain the second image, the first mask image is an image obtained by dividing the region of the portrait in the first image, and in the first model, using the first mask image as a guidance map to extract the features of the image of the first region and the image features of the second region, so as to realize processing the image of the first region in the first image enhancement manner and processing the image of the second region in the second image enhancement manner.

5. The method according to claim 4, characterized in that, the first model is used to extract a first feature of the first image based on the information indicated in the first mask image, downsample the first feature and then perform feature parallel connection with the features of the first mask image to obtain a second feature, and realize processing the image of the first region in the first image enhancement manner based on the second feature and realize processing the image of the second region in the second image enhancement manner based on the second feature.

6. The method according to claim 5, characterized in that, the first model includes a shallow network, an intermediate network and a deep network; the shallow network is used to extract a first feature of the first image based on the information indicated in the first mask image; the intermediate network is used to downsample the first feature and then perform feature parallel connection with the features of the first mask image to obtain a second feature; the deep network is used to downsample the second feature; or, the shallow network is used to extract a first feature of the first image based on the information indicated in the first mask image; The intermediate layer network is used to downsample the first feature; the deep layer network is used to downsample the feature obtained from the intermediate layer network and then perform feature parallel connection with the feature of the first mask image.

7. The method according to any one of claims 4 to 6, wherein, the method further includes: The first model is trained in the following manner: Input the first image sample and the second mask image into the source model to obtain the output image of the source model. The second mask image is an image obtained by partitioning the portrait in the second image sample. The first image sample is obtained by processing the image of the first region in the second image sample in a first blurring manner and processing the image of the second region in the second image sample in a second blurring manner. The clarity of the image processed by the first blurring manner is different from the blurriness of the image processed by the second blurring manner; When the target loss function satisfies the preset model convergence condition, the first model is obtained, wherein the target loss function is related to the first loss function and the second loss function. The first loss function is the difference between the image of the first region in the output image of the source model and the image of the first region in the second image sample. The second loss function is the difference between the image of the second region in the output image of the source model and the image of the second region in the second image sample.

8. The method according to claim 7, wherein, The difference between the number of pixel points included in the image of the first region in the first image sample and the number of pixel points included in the image of the second region in the first image sample is less than or equal to a preset value.

9. The method according to claim 7 or 8, wherein, The relationship between the target loss function and the first loss function and the second loss function satisfies the following formula: L = W 1 *L 1 +W 2 *L 2 Among them, L is used to represent the target loss function, L i is used to represent the first loss function, L j is used to represent the second loss function, W 1 is used to represent the weight of the first loss function, W 2 is used to represent the weight of the second loss function, the W 1 and the W 2 are both constants.

10. The method according to any one of claims 1 to 9, wherein, Before obtaining the first image in response to the operation for triggering photographing, the method further includes: In response to the operation of the user turning on the front camera, display a first interface, where the first interface includes the image collected by the front camera and a first control; The obtaining the first image in response to the operation for triggering photographing includes: In response to the operation of the user triggering the first control, obtain the first image collected by the front camera.

11. An electronic device, wherein, including: a processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1 - 10.

12. A chip system, wherein, including: a processor, configured to read the instructions stored in the memory. When the processor executes the instructions, the chip system implements the method according to any one of claims 1 - 10.

13. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1-10 is implemented.

14. A computer program product, characterized in that it includes a computer program, and when the computer program is run, it causes a computer to execute the method described in any one of claims 1-10.

Citation Information

Patent Citations

  • Image denoising method and device, storage medium and terminal equipment

    CN114596226A

  • Image enhancement method, training method of image enhancement model and related equipment

    CN114627034A

  • Image processing method and related device

    CN115170455A

  • Deep-learning-based automatic skin retouching

    US20190251674A1

  • Techniques for image attribute editing using neural networks

    US20230162330A1