Image beauty processing method, device, storage medium and electronic device
The three-dimensional grid features are extracted through deep neural networks and generated information matrix, which solves the problem of unsatisfactory beauty effects in the existing technology under complex lighting and various skin conditions, and achieves more flexible and efficient beauty treatment.
Patent Information
- Application Number
- CN202111176571.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-10-09
AI Technical Summary
The prior art is difficult to achieve ideal beauty effects when dealing with complex lighting conditions and a variety of skin conditions.
A pre-trained deep neural network is used to generate an information matrix by extracting features based on a three-dimensional grid, and the beauty face image is processed to generate a beauty face image.
Improves the flexibility of image beauty processing, is suitable for a variety of lighting conditions and skin conditions, improves beauty effects, and reduces calculation and implementation costs.
Smart Images

Figure CN113902611B_ABST
Abstract
Description
Background Art
[0002] Beauty enhancement refers to using image processing technology to beautify the portraits in images or videos to better meet the aesthetic needs of users.
[0003] In related technologies, image beauty enhancement processing usually includes multiple fixed algorithm processes, such as calculating based on artificially designed image features, spatial filtering processing, layer fusion, etc. However, in actual shooting scenarios, there may be complex and diverse lighting conditions, and the skin conditions of the shooting objects are diverse. Using the above methods cannot better handle different situations, resulting in an unsatisfactory beauty enhancement effect.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present disclosure provides an image beauty enhancement processing method, an image beauty enhancement processing device, a computer-readable storage medium, and an electronic device, thereby at least to some extent improving the image beauty enhancement effect.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be learned in part through the practice of the present disclosure.
[0007] According to a first aspect of the present disclosure, there is provided an image beauty enhancement processing method, including: obtaining a face image to be beautified; extracting three-dimensional grid-based features from the face image to be beautified through a pre-trained deep neural network, and generating an information matrix according to the extracted features, where the three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and the pixel value domain of the face image to be beautified; using the information matrix to process the face image to be beautified to obtain a beautified face image corresponding to the face image to be beautified.
[0008] According to a second aspect of the present disclosure, there is provided an image beauty enhancement processing device, including: an image acquisition module configured to obtain a face image to be beautified; an information matrix generation module configured to extract three-dimensional grid-based features from the face image to be beautified through a pre-trained deep neural network, and generate an information matrix according to the extracted features, where the three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and the pixel value domain of the face image to be beautified; a beauty enhancement processing module configured to use the information matrix to process the face image to be beautified to obtain a beautified face image corresponding to the face image to be beautified.
[0009] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the image beautification processing method of the first aspect and its possible implementation manners as described above.
[0010] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the image beautification processing method of the first aspect and its possible implementation manners by executing the executable instructions.
[0011] The technical solution of the present disclosure has the following beneficial effects:
[0012] Based on the image beautification processing method of the present disclosure, on the one hand, through the processing of a deep neural network, blemish removal or other beautification functions are achieved, replacing multiple fixed algorithmic processes in the related art, increasing the flexibility of image beautification processing, being applicable to diverse lighting conditions or skin conditions, improving the image beautification effect, and reducing the time consumption and memory occupancy. On the other hand, the deep neural network in this solution outputs an information matrix and does not directly output the beautified image, thereby reducing the computational amount of the deep neural network, facilitating the implementation of a lightweight network, and reducing the implementation cost of the solution.
[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0015] Figure 1 A schematic diagram showing a system architecture in this exemplary embodiment;
[0016] Figure 2 A schematic diagram showing the structure of an electronic device in this exemplary embodiment;
[0017] Figure 3 A flowchart showing an image beautification processing method in this exemplary embodiment;
[0018] Figure 4 A flowchart showing the generation of a face image to be beautified in this exemplary embodiment;
[0019] Figure 5Flowchart showing a method for obtaining a stable bounding box of a face to be determined in this exemplary embodiment;
[0020] Figure 6 Schematic diagram showing a method for combining original face sub-images in this exemplary embodiment;
[0021] Figure 7 Schematic diagram showing a deep neural network and image beauty processing in this exemplary embodiment;
[0022] Figure 8 Flowchart showing a method for obtaining an information matrix in this exemplary embodiment;
[0023] Figure 9 Flowchart showing a method for obtaining a beauty face image based on a beauty information matrix in this exemplary embodiment;
[0024] Figure 10 Flowchart showing a method for training a deep neural network in this exemplary embodiment;
[0025] Figure 11 Schematic diagram showing a gradient processing method for a boundary region in this exemplary embodiment;
[0026] Figure 12 Schematic flowchart showing a method for image beauty processing in this exemplary embodiment;
[0027] Figure 13 Schematic diagram showing the structure of an image beauty processing apparatus in this exemplary embodiment. Detailed implementation manners
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0029] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus the repeated description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0030] Portrait blemish removal is part of image beauty processing, usually the first-stage processing in image beauty processing. Portrait blemish removal includes, but is not limited to, freckle and acne removal, eye bag removal, dirty corner of the mouth treatment, light and shadow leveling, dry lip line treatment, etc. After portrait blemish removal, processing such as skin smoothing, skin tone adjustment, facial feature deformation, and brightness adjustment can be continued.
[0031] In the related art, the effect of portrait blemish removal depends on the calculation of image features designed manually. However, the calculation of image features designed manually is difficult to handle all situations in practical applications, and it is usually difficult to accurately and fully detect blemishes on the skin, resulting in incomplete removal of portrait blemishes. Moreover, there is also a problem that the skin looks unrealistic after portrait blemish removal in the related art. For example, after a mole on the face is removed, there is a contrast with the surrounding skin, making it look unnatural.
[0032] In view of the above one or more problems, an exemplary embodiment of the present disclosure provides an image beauty processing method. The following combines Figure 1 to exemplarily illustrate the system architecture and application scenario of the operating environment of this exemplary embodiment.
[0033] Figure 1 A schematic diagram of the system architecture is shown. The system architecture 100 may include a terminal 110 and a server 120. Among them, the terminal 110 may be a terminal device such as a smart phone, a tablet computer, a desktop computer, a laptop computer, etc., and the server 120 generally refers to a background system that provides image beauty-related services in this exemplary embodiment, which may be a single server or a cluster formed by multiple servers. A connection may be formed between the terminal 110 and the server 120 through a wired or wireless communication link for data interaction.
[0034] In one embodiment, the terminal 110 can capture or obtain an image or video to be beautified in other ways and upload it to the server 120. For example, the user opens a beauty App (Application) on the terminal 110, selects an image or video to be beautified from the album, and uploads it to the server 120 for beautification. Or the user opens the beauty function in the live broadcast App on the terminal 110 and uploads the real-time captured video to the server 120 for beautification. The server 120 executes the above image beautification processing method to obtain a beautified image or video and returns it to the terminal 110.
[0035] In one embodiment, the server 120 can perform the training of the deep neural network, send the trained deep neural network to the terminal 110 for deployment. For example, the relevant data of the deep neural network is packaged in the update package of the above beauty App or live broadcast App, so that the terminal 110 can obtain the deep neural network by updating the App and deploy it locally. Furthermore, after the terminal 110 captures or obtains an image or video to be beautified in other ways, it can call the deep neural network to implement the beautification processing of the image or video by executing the above image beautification processing method.
[0036] In one embodiment, the terminal 110 can perform the training of the deep neural network. For example, it obtains the basic architecture of the deep neural network from the server 120 and trains it through the local data set, or obtains the data set from the server 120 and trains the locally constructed deep neural network, or trains the deep neural network completely independent of the server 120. Furthermore, the terminal 110 can call the deep neural network to implement the beautification processing of the image or video by executing the above image beautification processing method.
[0037] As can be seen from the above, the execution subject of the image beautification processing method in this exemplary embodiment can be the above terminal 110 or server 120, and the present disclosure does not limit this.
[0038] The exemplary embodiment of the present disclosure also provides an electronic device for executing the above deep neural network training method or image beautification processing method. The electronic device can be the above terminal 110 or server 120. Below, taking Figure 2 the mobile terminal 200 in Figure 2 as an example, the structure of the above electronic device is described exemplarily. Those skilled in the art should understand that except for the components specially used for mobile purposes,
[0039] such as Figure 2As shown in the figure, the mobile terminal 200 may specifically include: a processor 201, a memory 202, a bus 203, a mobile communication module 204, an antenna 1, a wireless communication module 205, an antenna 2, a display screen 206, a camera module 207, an audio module 208, a power module 209, and a sensor module 210.
[0040] The processor 201 may include one or more processing units. For example, the processor 201 may include an AP (Application Processor), a modulation and demodulation processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc. The deep neural network in this exemplary embodiment may run on a GPU, a DSP, or an NPU. The DSP and the NPU generally run the deep neural network with int-type data (integer type), and the GPU generally runs the deep neural network with float-type data (floating point type). In comparison, the power consumption of running on the DSP and the NPU is lower, the response speed is faster, and the accuracy is lower. The power consumption of running on the GPU is higher, the response speed is slower, and the accuracy is higher. In practical applications, a suitable processing unit can be selected to run the deep neural network according to the device performance and actual requirements. For example, when performing real-time beauty on the images in a video, since high speed is required, the DSP or the NPU can be selected to run the deep neural network.
[0041] The encoder can encode (i.e., compress) image or video data to form corresponding bitstream data, so as to reduce the bandwidth occupied by data transmission; the decoder can decode (i.e., decompress) the bitstream data of the image or video to restore the image or video data. The mobile terminal 200 can process images or videos in multiple coding formats, such as: image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), etc., and video formats such as MPEG (Moving Picture Experts Group) 1, MPEG2, H.263, H.264, HEVC (High Efficiency Video Coding), etc.
[0042] The processor 201 can be connected to the memory 202 or other components via the bus 203.
[0043] The memory 202 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 201 executes various functional applications and data processing of the mobile terminal 200 by running the instructions stored in the memory 202. The memory 202 can also store application data, such as storing files like images, videos, etc.
[0044] The communication function of the mobile terminal 200 can be implemented by a mobile communication module 204, antenna 1, a wireless communication module 205, antenna 2, a modulation and demodulation processor, a baseband processor, etc. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 204 can provide 2G, 3G, 4G, 5G, etc. mobile communication solutions applied to the mobile terminal 200. The wireless communication module 205 can provide wireless communication solutions such as wireless local area network, Bluetooth, near field communication, etc. applied to the mobile terminal 200.
[0045] The display screen 206 is used to implement the display function, such as displaying user interfaces, images, videos, etc. The camera module 207 is used to implement the shooting function, such as shooting a face image to be beautified or an original image to be beautified, etc. The audio module 208 is used to implement the audio function, such as playing audio, collecting voices, etc. The power module 209 is used to implement the power management function, such as charging the battery, powering the device, monitoring the battery status, etc. The sensor module 210 may include a depth sensor 2101, a pressure sensor 2102, a gyroscope sensor 2103, a barometric pressure sensor 2104, etc. to implement corresponding sensing and detection functions.
[0046] The following combines Figure 3 to describe the image beautification processing method in this exemplary embodiment. Figure 3 An exemplary process of the image beautification processing method is shown, which may include:
[0047] Step S310, obtaining a face image to be beautified;
[0048] Step S320, extracting three-dimensional grid-based features from the face image to be beautified through a pre-trained deep neural network, and generating an information matrix according to the extracted features. The three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and the pixel value domain of the face image to be beautified;
[0049] Step S330, processing the face image to be beautified using the information matrix to obtain a beautified face image corresponding to the face image to be beautified.
[0050] A deep neural network (DNN) is used to output an information matrix, and the information matrix is used to implement a beauty function. That is to say, the deep neural network is used to indirectly implement the beauty function. In this exemplary embodiment, the deep neural network can be trained to implement any one beauty function or a combination of any multiple beauty functions. The beauty functions include, but are not limited to, blemish removal (such as freckle removal, acne removal, eye bag removal), deformation, skin tone adjustment, skin smoothing, light and shadow adjustment, dirty corner of the mouth treatment, lip treatment, etc. Thus, Figure 3 The image beauty processing method can be used as the beauty processing in one stage, and before or after Figure 3 The image beauty processing method, other stages of beauty processing can be added. For example, the deep neural network is used to perform blemish removal on the face image to be beautified. After obtaining the face image to be beautified, it is processed through Figure 3 The image beauty processing method, and the obtained beautified face image is a blemish-removed beauty image. Subsequently, the blemish-removed beauty image can be further processed for personalized beauty to obtain the final beauty image.
[0051] Generally, blemish removal is necessary for image beauty, and the user's demand for blemish removal is relatively fixed. The general blemish-removed beauty processing process can be implemented through Figure 3 The image beauty processing method. In contrast, beauty functions such as skin smoothing, deformation, three-dimensional effect, skin tone adjustment, and light and shadow adjustment are not necessary, and the specific demands of users for these beauty functions also show personalized characteristics. These beauty functions can be called personalized beauty processing, which usually requires the user to make specific settings before processing. For example, the user selects one or more of these beauty functions and sets parameters such as the skin smoothing degree and deformation degree, and then the terminal or server processes according to the user's settings.
[0052] It should be noted that this disclosure does not limit the order of Figure 3 The image beauty processing and other beauty processing. For example, the image can be first processed for personalized beauty to obtain an intermediate beauty image, and then the intermediate beauty image is used as the face image to be beautified, and Figure 3 The image beauty processing method is executed, and the obtained beautified face image is the final output beauty image.
[0053] Based on the above image beauty processing method, on the one hand, the processing of the deep neural network is used to achieve blemish removal or other beauty functions, replacing multiple fixed algorithm processes in related technologies, increasing the flexibility of image beauty processing, being applicable to diverse lighting conditions or skin conditions, improving the image beauty effect, and reducing the time consumption and memory occupation. On the other hand, the deep neural network in this solution outputs an information matrix and does not directly output the beauty-enhanced image, thereby reducing the computational amount of the deep neural network, facilitating the implementation of a lightweight network, and reducing the implementation cost of the solution.
[0054] The following will specifically describe Figure 3 each step in.
[0055] Referring to Figure 3 , in step S310, a face image to be beautified is obtained.
[0056] Among them, the face image to be beautified can be the original image to be beautified captured, or the image obtained after a certain processing of the original image to be beautified. For example, if the original image to be beautified includes more image content other than the face, and these image contents do not need to be beautified, the face part can be intercepted from the original image to be beautified as the face image to be beautified.
[0057] In one implementation manner, the face image to be beautified can be one or more frames in a series of consecutive frames. Among them, the series of consecutive frames can be a video or a series of captured images, etc. The series of consecutive frames is the object to be beautified. Taking a video as an example, it can be a currently real-time captured or received video stream, or a completed captured or received complete video, such as a video stored locally. The present disclosure does not limit parameters such as the frame rate and image resolution of the video. For example, the video frame rate can be 30fps (frames per second), 60fps, 120fps, etc., and the image resolution can be 720P, 1080P, 4K, etc. and corresponding different aspect ratios. Each frame in the video can be beautified, or a part of the images can be selected from the video for beautification, and the images to be beautified are used as the above-mentioned original image to be beautified or face image to be beautified. For example, when receiving a video stream in real time, each received frame can be used as the face image to be beautified.
[0058] In one implementation manner, referring to Figure 4 shown, the above-mentioned obtaining of the face image to be beautified may include the following steps S410 and S420:
[0059] Step S410, extracting one or more original face sub-images from the original image to be beautified.
[0060] The original face sub-image is a sub-image obtained by cropping the face part in the original image to be beautified. The number of faces in the original image to be beautified in this exemplary embodiment is not limited. For example, when there are multiple faces in the original image to be beautified, multiple original face sub-images can be extracted, and through the processing of subsequent steps, multiple faces can be beautified simultaneously.
[0061] In one embodiment, the extraction of one or more original face sub-images from the original image to be beautified may include the following steps:
[0062] According to the face key points recognized in the original image to be beautified, generate one or more bounding boxes of faces in the original image to be beautified;
[0063] Retain the bounding boxes with an area greater than or equal to the face area threshold, and crop the images within the bounding boxes to obtain one or more original face sub-images.
[0064] Among them, the face key points may include the key parts of the face and the points on the face edge. The bounding box is a region in the image that encloses the face and has a certain geometric shape. The shape of the bounding box in this disclosure is not limited, such as it can be any shape like a rectangle, trapezoid, etc. The face key points of each face are within the bounding box. In one embodiment, the bounding box can be the smallest rectangle including the face key points.
[0065] Generally, all faces can be detected in the original image to be beautified through a face detection algorithm, which may include faces that do not need to be beautified (such as the faces of passers-by in the distance). Considering that in the scenario of image beautification, usually larger faces need to be beautified (the beautification effect of smaller faces is not obvious, so they usually do not need to be beautified), therefore, the bounding boxes can be filtered by the face area threshold. The face area threshold can be set according to experience or the size of the original image to be beautified. Exemplarily, the face area threshold can be the size of the original image to be beautified * 0.05; if the area of the bounding box is greater than or equal to the face area threshold, it is a face that needs to be beautified, and the bounding box is retained; if the area of the bounding box is less than the face area threshold, it is a face that does not need to be beautified, and the bounding box is deleted.
[0066] After the filtering of the bounding boxes is completed, the retained bounding boxes are the bounding boxes of valid faces. Crop the images within each bounding box to obtain the same number of original face sub-images as the number of bounding boxes.
[0067] In one embodiment, to facilitate subsequent combination of the original face sub-images, an upper limit on the number of original face sub-images can be set, that is, an upper limit on the number of bounding boxes is set. For example, it can be set to 4. If the number of remaining bounding boxes is greater than 4 after filtering by the above-mentioned face area threshold, 4 bounding boxes can be selected therefrom. For example, it can be the 4 bounding boxes with the largest areas, or the 4 bounding boxes closest to the center of the original image to be beautified. Correspondingly, 4 original face sub-images are intercepted, and the faces in other bounding boxes are not beautified; or multiple beautification processes can be performed. In this process, 4 bounding boxes are selected and the original face sub-images are intercepted for beautification, and in the next process, other bounding boxes are selected and the original face sub-images are intercepted for beautification, so as to complete the beautification of the faces in all bounding boxes in the original image to be beautified whose areas are greater than the face area threshold.
[0068] In one embodiment, before intercepting the image in the bounding box, the bounding box can also be enlarged so that the bounding box includes a small amount of area outside the face, which is convenient for subsequent gradient processing during image fusion. When performing the enlargement process, the bounding box can be enlarged in one or more directions according to a preset ratio. For example, the preset ratio is 1.1, and the bounding box is enlarged evenly in all directions so that the size of the enlarged bounding box is 1.1 times the original size. It should be noted that when enlarging the bounding box, if one or more boundaries of the bounding box reach the boundary of the original image to be beautified, the boundary of the bounding box stays at the boundary of the original image to be beautified.
[0069] For the case where the original image to be beautified is one or more frames in a series of consecutive frames of images, the extraction of the original face sub-images can be based on the information of other frames in the series of consecutive frames of images. In one embodiment, the face in the original image to be beautified is matched with the face in the reference frame image of the original image to be beautified, and a stable bounding box of the face in the original image to be beautified is determined according to the matching result; based on the stable bounding box of the face in the original image to be beautified, the original face sub-image is extracted from the original image to be beautified.
[0070] Among them, the bounding box of the initially detected face is called the basic bounding box. For example, it can be the smallest rectangle including the face key points, or the face frame obtained through relevant algorithms. The basic bounding box is optimized, such as expanded, position corrected, etc., and the optimized bounding box is called the stable bounding box.
[0071] In this exemplary embodiment, the original image to be beautified can be subjected to face detection to obtain relevant information about the face. The present disclosure does not limit the face detection algorithm. For example, the face key points can be detected through a specific neural network, including the key points of the face boundary, and the basic bounding box of the face is generated according to the key points of the face boundary, and the stable bounding box is obtained through optimization.
[0072] The reference frame image can be any one of the above-mentioned consecutive multiple frame images for which a stable face bounding box has been determined or beauty processing has been completed. For example, when performing frame-by-frame beauty processing on a video, the previous frame image of the original image to be beautified can be used as the reference frame image. By matching the face in the original image to be beautified with the face in the reference frame image, the stable face bounding box in the original image to be beautified can be determined based on the stable face bounding box in the reference frame image.
[0073] In one implementation, referring to Figure 5 As shown, the above-mentioned matching of the face in the original image to be beautified with the face in the reference frame image of the original image to be beautified and determining the stable face bounding box in the original image to be beautified according to the matching result may include the following steps S510 to S530:
[0074] Step S510: Detect the face in the original image to be beautified, denoted as the face to be determined, and match the face to be determined with the determined face in the reference frame image of the original image to be beautified.
[0075] Among them, the face to be determined refers to a face that needs to be beautified but for which a stable bounding box has not been determined, and can be regarded as a face with unknown identity. The determined face refers to a face for which a stable bounding box has been determined and can be regarded as a face with known identity. All the faces in the reference frame image for which a stable bounding box has been determined are determined faces. Correspondingly, the faces detected in the original image to be beautified are faces for which a stable bounding box has not been determined, that is, faces to be determined. Matching the face to be determined in the original image to be beautified with the determined face in the reference frame image can infer that there is a correlation between the stable bounding box of the face to be determined and the stable bounding box of the determined face that matches the face to be determined, and the stable bounding box of the face to be determined can be determined therefrom.
[0076] In one implementation, the face area threshold can be set according to experience or the size of the original image to be beautified; if the area of the basic bounding box of the face is greater than or equal to the face area threshold, it is a face that needs to be beautified, and the information such as the basic bounding box of the face can be retained, or the face can be denoted as the face to be determined; if the area of the basic bounding box of the face is less than the face area threshold, it is a face that does not need to be beautified, and the relevant information such as the basic bounding box of the face can be deleted and no subsequent processing is performed on it.
[0077] In one embodiment, to facilitate subsequent processing of the original face sub-images, such as combining the original face sub-images, or considering the limitations of device performance, an upper limit on the number of original face sub-images can be set, that is, an upper limit on the number of faces to be determined is set. For example, it can be set to 4. If the number of faces retained after filtering by the above face area threshold is greater than 4, then 4 faces to be determined can be further selected from them. For example, it can be the 4 faces with the largest areas, or the 4 faces closest to the center of the original image to be beautified. Then, 4 original face sub-images are intercepted correspondingly in the subsequent process, and no subsequent processing is performed on other faces. Alternatively, multiple beautification processes can be performed. In this processing, 4 faces are selected as the faces to be determined, and their corresponding original face sub-images are intercepted and beautified. In the next processing, other faces are selected as the faces to be determined, and their corresponding original face sub-images are intercepted and beautified, so as to complete the beautification of all faces within the basic bounding boxes with areas greater than the face area threshold in the image to be processed.
[0078] In one embodiment, to facilitate tracking and identifying faces in consecutive frames of images, an ID (Identity Document) can be assigned to each face. For example, starting from the first frame, an ID is assigned to each face. Subsequently, after detecting a face in each frame, each face is matched with the faces in the previous frame. If the match is successful, the face ID and other relevant information in the previous frame are inherited. If the match is unsuccessful, it is regarded as a new face and a new ID is assigned.
[0079] The present disclosure does not limit the method of matching the faces to be determined with the determined faces. For example, a face recognition algorithm can be used to perform recognition and comparison between each face to be determined and each determined face. If the similarity is higher than a preset similarity threshold, it is determined that the face to be determined and the determined face match successfully.
[0080] In one embodiment, it can be determined whether the face to be determined and the determined face match successfully according to the overlap degree (Intersection Over Union, IOU, also known as the intersection-to-union ratio) between the basic bounding box of the face to be determined and the basic bounding box of the determined face. The following provides an exemplary method for calculating the overlap degree:
[0081] Obtain the position of the basic bounding box of the face to be determined in the original image to be beautified, and the position of the basic bounding box of the determined face in the reference frame image. Count the number of pixel points with overlapping positions in the two basic bounding boxes, denoted as k1, and the number of pixel points with non-overlapping positions, denoted as k2 (representing the number of pixel points in the basic bounding box of the face to be determined that do not overlap with the basic bounding box of the determined face) and k3 (representing the number of pixel points in the basic bounding box of the determined face that do not overlap with the basic bounding box of the face not yet determined). Then the overlap degree of the two basic bounding boxes is:
[0082]
[0083] After determining the overlap degree, if the overlap degree reaches the preset overlap degree threshold, it is determined that the face to be determined matches the determined face successfully. The overlap degree threshold can be set according to experience and actual needs. For example, it can be set to 0.75.
[0084] In addition, the basic bounding box of the face to be determined or the basic bounding box of the determined face can be iteratively transformed by algorithms such as the ICP (Iterative Closest Point) algorithm, and the overlap degree of the two basic bounding boxes is calculated based on the number of pixel points with the same pixel values and the number of pixel points with different pixel values in the finally transformed basic bounding box of the face to be determined and the basic bounding box of the determined face, so as to determine whether the match is successful.
[0085] It should be noted that since there may be multiple faces to be determined in the original image to be beautified and multiple determined faces in the reference frame image, the matching calculation can be performed separately for each face to be determined and each determined face to obtain a similarity matrix or an overlap degree matrix. Then, algorithms such as the Hungarian algorithm can be used to achieve the global maximum matching, and it is determined whether they match successfully according to the similarity or overlap degree of each pair of faces to be determined and determined faces.
[0086] Step S520, if the face to be determined does not match the determined face successfully, expand the basic bounding box of the face to be determined according to the first preset parameter to obtain the stable bounding box of the face to be determined;
[0087] The fact that the face to be determined does not match the determined face successfully indicates that the face to be determined is a newly emerged face in consecutive multiple frames of images and no reference information can be obtained from the reference frame image. Therefore, on the basis of the basic bounding box of the face to be determined, appropriate expansion can be performed to obtain the stable bounding box. The first preset parameter is the expansion parameter for the basic bounding box of the newly emerged face, which can be determined according to experience or actual needs. For example, it can be to expand both the width and height of the basic bounding box by 1 / 4.
[0088] Suppose the basic bounding box of the face to be determined is represented as [bb0, bb1, bb2, bb3], where bb0 is the abscissa of the upper left point of the basic bounding box, bb1 is the ordinate of the upper left point of the basic bounding box, bb2 is the abscissa of the lower right point of the basic bounding box, bb3 is the ordinate of the lower right point of the basic bounding box, the width of the basic bounding box is w, and the height is h. Note that the pixel coordinates in the image usually have the upper left point of the image as (0, 0) and the lower right point as (Wmax, Hmax), where Wmax and Hmax represent the width and height of the image. Therefore, bb0 < bb2 and bb1 < bb3. Let E1 represent the first preset parameter. When the basic bounding box is expanded centered according to the first preset parameter (i.e., uniformly expanded up, down, left, and right), the size of the stable bounding box can be obtained as follows:
[0089]
[0090] where expand_w and expand_h are the width and height of the stable bounding box of the face to be determined respectively. It should be noted that if the expanded width expand_w exceeds the width Wmax of the original image to be beautified, then expand_w = Wmax; if the expanded height expand_h exceeds the height Hmax of the original image to be beautified, then expand_h = Hmax.
[0091] The center point coordinates of the stable bounding box are equal to the center point coordinates of the basic bounding box, that is:
[0092]
[0093] where center_x represents the x coordinate of the center point of the stable bounding box of the face to be determined, and center_y represents the y coordinate of the center point of the stable bounding box of the face to be determined.
[0094] Then the coordinates of the upper left point and the lower right point of the stable bounding box can be calculated as follows:
[0095]
[0096] where expand_bb0 is the abscissa of the upper left point of the stable bounding box, expand_bb1 is the ordinate of the upper left point of the stable bounding box, expand_bb2 is the abscissa of the lower right point of the stable bounding box, and expand_bb3 is the ordinate of the lower right point of the stable bounding box. Thus, the stable bounding box of the face to be determined is obtained. If the calculated coordinates exceed the boundaries of the original image to be beautified, the boundary coordinates of the original image to be beautified are used to replace the coordinates that exceed the boundaries. Finally, the expanded bounding box can be represented in the form of [expand_bb0, expand_bb1, expand_bb2, expand_bb3].
[0097] It should be added that the above coordinates usually adopt pixel coordinates in the image and are integers. Therefore, when calculating, float-type data can be used for calculation, and then rounding is performed, and the result is saved as int-type data. Exemplarily, when it comes to division operations, float-type data is used for calculation and the intermediate results are cached, and rounding is performed when calculating the final results (including the above expand_w, expand_h, center_x, center_y, expand_bb0, expand_bb1, expand_bb2, expand_bb3), and they are saved as int-type data.
[0098] For the center point coordinates, since saving int-type data will affect the accuracy of subsequent processing of other frames, int-type and float-type data can be saved. For example, the result calculated in formula (3) is saved as float-type data, as shown below:
[0099]
[0100] Among them, center_x_float and center_y_float represent the center point coordinates saved as float-type data, center_x and center_y represent the center point coordinates saved as int-type data, and int() represents the rounding operation.
[0101] Furthermore, to ensure the accuracy of the results, formula (4) can be changed to the following calculation method:
[0102]
[0103] Step S530, if the to-be-determined face matches the determined face successfully, then determine the stable bounding box of the to-be-determined face according to the stable bounding box of the determined face.
[0104] Generally, the to-be-determined face in the original image to be beautified does not change much compared with the determined face in the reference frame image it matches, which is reflected in that neither the position change nor the size change is large. Therefore, based on the stable bounding box of the determined face, appropriate position changes and size changes can be made to obtain the stable bounding box of the to-be-determined face.
[0105] In one implementation manner, the position change and size change of the stable bounding box of the determined face can be performed according to the position change parameter and size change parameter of the basic bounding box of the to-be-determined face relative to the basic bounding box of the determined face to obtain the stable bounding box of the to-be-determined face.
[0106] In one implementation, determining the stable bounding box of the face to be determined based on the stable bounding box of the determined face may include the following steps:
[0107] Based on a preset stability coefficient, the center point coordinates of the stable bounding box of the determined face and the center point coordinates of the basic bounding box of the face to be determined are weighted to obtain the center point coordinates of the stable bounding box of the face to be determined.
[0108] The above steps represent fusing the position of the stable bounding box of the determined face with the position of the basic bounding box of the face to be determined as the position of the stable bounding box of the face to be determined. When fusing, the center point coordinates of the two are weighted using a preset stability coefficient, and the preset stability coefficient can be the weight of the stable bounding box of the determined face, which can be determined according to experience or the actual scenario. Generally, in a scenario where the face moves faster, the preset stability coefficient is smaller. Exemplarily, in a live broadcast scenario, the face usually moves within a certain range with a small amplitude, and the preset stability coefficient can be set to 0.9. Then, the calculation of the center point coordinates of the stable bounding box of the face to be determined is as follows:
[0109]
[0110] Among them, pre_center_x represents the x coordinate of the center point of the stable bounding box of the determined face, and pre_center_y represents the y coordinate of the center point of the stable bounding box of the determined face. It can be seen that formula (7) represents weighting the two center point coordinates with the weight of the center point coordinates of the stable bounding box of the determined face being 0.9 and the weight of the center point coordinates of the basic bounding box of the face to be determined being 0.1 to obtain the center point coordinates of the stable bounding box of the face to be determined.
[0111] Similar to the above formula (5), the center point coordinates of int type and float type data can be saved, and then there is:
[0112]
[0113] Among them, pre_center_x_float is the float type data of the saved pre_center_x, and pre_center_y_float is the float type data of the saved pre_center_y.
[0114] By the above method of calculating the center point coordinates through weighting, in essence, a mechanism of momentum update for the center point coordinates is adopted, which can avoid excessive movement of the center point coordinates of the stable bounding box of the same face from the reference frame image to the original image to be beautified, resulting in jitter of the subsequently intercepted original face sub-image and affecting the beautification effect.
[0115] In one implementation, determining the stable bounding box of the face to be determined based on the stable bounding box of the determined face may include the following steps:
[0116] If the size of the basic bounding box of the face to be determined is greater than the product of the size of the stable bounding box of the determined face and the first magnification factor, expand the size of the stable bounding box of the determined face according to the second preset parameter to obtain the size of the stable bounding box of the face to be determined;
[0117] If the size of the basic bounding box of the face to be determined is less than the product of the size of the stable bounding box of the determined face and the second magnification factor, reduce the size of the stable bounding box of the determined face according to the third preset parameter to obtain the size of the stable bounding box of the face to be determined; the first magnification factor is greater than the second magnification factor;
[0118] If the size of the basic bounding box of the face to be determined is less than the product of the size of the stable bounding box of the determined face and the first magnification factor and greater than the product of the size of the stable bounding box of the determined face and the second magnification factor, take the size of the stable bounding box of the determined face as the size of the stable bounding box of the face to be determined.
[0119] The above steps indicate that according to the comparison result between the size of the basic bounding box of the face to be determined and the size of the stable bounding box of the determined face, calculations are performed in three cases respectively. The first magnification factor and the second magnification factor can be integer magnification factors or non-integer magnification factors. In one implementation, the first magnification factor is greater than or equal to 1, and the second magnification factor is less than 1. Exemplarily, the first magnification factor can be 1, and the second magnification factor can be 0.64.
[0120] When performing calculations, comparisons and calculations can be performed separately for the width and height. For example, if the comparison result of the width belongs to the first case above and the comparison result of the height belongs to the second case, calculate the width and height of the stable bounding box of the face to be determined in the two cases respectively.
[0121] Assume that the first magnification factor is t1 and the second magnification factor is t2. The calculation of the width is described as follows:
[0122] First case: If w > pre_expand_w·t1, and E2 represents the second preset parameter, then:
[0123] expand_w = pre_expand_w + pre_expand_w·E2 (9)
[0124] Second case: If w < pre_expand_w·t2, and E3 represents the third preset parameter, then:
[0125] expand_w = pre_expand_w - pre_expand_w·E3 (10)
[0126] In the third case, if pre_expand_w·t2 < w < pre_expand_w·t1, then:
[0127] expand_w = pre_expand_w (11)
[0128] For the height, it can also be calculated respectively according to the above three cases to obtain expand_h.
[0129] Generally, in a sequence of consecutive video frames, as long as the face does not approach the camera rapidly, move away from the camera rapidly, or move out of the frame, the size of the face will not change drastically, and thus the third case above is satisfied. At this time, the size of the stable bounding box of the face to be determined is made equal to the size of the stable bounding box of the face that has been determined, that is, the size of the stable bounding box is kept unchanged. The first and second cases above are both situations where the size of the face changes drastically. In the first case, the face expands drastically. At this time, the size of the stable bounding box of the face that has been determined is appropriately expanded according to the second preset parameter to obtain the size of the stable bounding box of the face to be determined. The second preset parameter can be determined according to experience and the actual scenario. In the second case, the face shrinks drastically. At this time, the size of the stable bounding box of the face that has been determined is appropriately reduced according to the third preset parameter to obtain the size of the stable bounding box of the face to be determined. The third preset parameter can be determined according to experience and the actual scenario.
[0130] If the expanded width expand_w exceeds the width Wmax of the original image to be beautified, then expand_w = Wmax; if the expanded height expand_h exceeds the height Hmax of the original image to be beautified, then expand_h = Hmax.
[0131] Through the calculations of the above three cases, it is possible to avoid excessive changes in the size of the stable bounding box of the same face from the reference frame image to the original image to be beautified, so as to prevent the original face sub-image intercepted subsequently from jittering and affecting the beautification effect.
[0132] After obtaining the center point coordinates and size of the stable bounding box of the face to be determined respectively, the coordinates of the upper left point and the lower right point of the stable bounding box can be calculated. If the calculated coordinates exceed the boundary of the original image to be beautified, the boundary coordinates of the original image to be beautified are used to replace the coordinates that exceed the boundary. Finally, the stable bounding box can be expressed in the form of [expand_bb0, expand_bb1, expand_bb2, expand_bb3].
[0133] As can be seen from the above, when the face to be determined matches the determined face successfully, the stable bounding box of the face to be determined is determined according to the stable bounding box of the determined face, so that the face to be determined inherits the information of the stable bounding box of the determined face to a certain extent, thereby ensuring a certain continuity and stability of the stable bounding boxes of faces between different frame images, preventing drastic position or size changes, and further ensuring the consistency of the face beautification effect during subsequent beautification processing, and preventing the beautified face from flashing due to drastic changes in the face.
[0134] In one implementation, after obtaining the stable bounding box of the face to be determined, relevant parameters of its stable bounding box can be saved, and the face to be determined can be marked as a determined face for use in matching the face to be determined and determining the stable bounding box in subsequent frames.
[0135] After obtaining the stable bounding box of the face in the original image to be beautified, the image within the stable bounding box can be intercepted to obtain the original face sub-image. When there are stable bounding boxes of multiple faces in the original image to be beautified, the original face sub-image corresponding to each face can be intercepted.
[0136] Step S420: Combine the original face sub-images based on the input image size of the deep neural network to generate the face image to be beautified.
[0137] The input image size is the image size that matches the input layer of the deep neural network. In this exemplary implementation, the original face sub-images are combined into an original face combined image, and the size of the original face combined image is the input image size. This exemplary implementation does not limit the size and aspect ratio of the input image size. Exemplarily, the ratio of the long side to the short side of the input image size can be set to be close to
[0138] In one implementation, the deep neural network can be a fully convolutional network, and the fully convolutional network can process images of different sizes. In this case, the deep neural network has no requirements for the input image size, and the size affects the amount of calculation, memory occupancy, and beautification fineness. The input image size can be determined according to the beautification fineness set by the user or the performance of the terminal device. Thus, this deep neural network can be deployed on devices with different performances such as high, medium, and low, with a wide range of applications, without the need to deploy different deep neural networks for different devices, reducing the network training cost. Exemplarily, considering that lightweight calculation is suitable on mobile terminals, the input image size can be determined as a relatively small value, such as 640 in width * 448 in height.
[0139] After obtaining the size of the input image, it is necessary to combine the original face sub-images into a face image to be beautified with this size. The specific combination method is related to the number of original face sub-images. In one implementation manner, the above-mentioned input image size based on the deep neural network combines the original face sub-images to generate a face image to be beautified, which may further include the following steps:
[0140] According to the number of original face sub-images, divide the input image size into sub-image sizes corresponding one by one to the original face sub-images;
[0141] Respectively transform the corresponding original face sub-images based on each sub-image size;
[0142] Combine the transformed original face sub-images to generate a face image to be beautified.
[0143] The following combines Figure 6 for example. Figure 6 In which Q represents the number of original face sub-images, Figure 6 Exemplary ways of input image size division and image combination when Q is 1 to 4 are respectively shown. Assume that the input image size is 640 in width * 448 in height. When Q is 1, the sub-image size is also 640 in width * 448 in height; when Q is 2, the sub-image size is half of the input image size, that is, 320 in width * 448 in height; when Q is 3, the sub-image sizes are 0.5, 0.25, 0.25 of the input image size respectively, that is, 320 in width * 448 in height, 320 in width * 224 in height, 320 in width * 224 in height; when Q is 4, the sub-image sizes are all 0.25 of the input image size, that is, 320 in width * 224 in height. Respectively transform each original face sub-image to be consistent with the sub-image size. It should be particularly noted that when the sub-image sizes are inconsistent, such as the case of Q being 3, the original face sub-images and the sub-image sizes can be corresponded one by one according to the size order of the original face sub-images and the size order of the sub-image sizes, that is, the largest original face sub-image corresponds to the largest sub-image size, and the smallest original face sub-image corresponds to the smallest sub-image size. After transforming the original face sub-images, then combine the transformed original face sub-images according to the Figure 6 shown method to generate a face image to be beautified.
[0144] In one implementation manner, when Q is an even number, the input image size can be divided into Q equal parts to obtain Q identical sub-image sizes. Specifically, Q can be decomposed into the product of two factors, that is, Q = q 1 *q 2 , so that q 1 / q 2 's ratio is the same as the aspect ratio of the input image size (such as )As close as possible, divide the width of the input image size into q 1 equal parts, and divide the height into q 2 equal parts. When Q is odd, divide the input image size into Q + 1 equal parts to obtain Q + 1 identical sub-image sizes. Combine two of the sub-image sizes into one sub-image size, and keep the remaining Q - 1 sub-image sizes unchanged, thus obtaining Q sub-image sizes.
[0145] In one embodiment, the size ratio (or area ratio) of the original face sub-image can be calculated first, such as it can be S 1 : S 2 : S 3 : …: S Q , and then divide the input image size into Q sub-image sizes according to this ratio.
[0146] After determining the sub-image size corresponding to each original face sub-image, the original face sub-image can be transformed based on the sub-image size. In one embodiment, the above-mentioned transformation of the corresponding original face sub-image based on each sub-image size respectively may include any one or more of the following:
[0147] ① When the size relationship between the width and height of the original face sub-image is different from the size relationship between the width and height of the sub-image size, rotate the original face sub-image by 90 degrees. That is to say, in the original face sub-image and the sub-image size, if both the width is greater than the height or both the width is less than the height, then the size relationship between the width and height of the original face sub-image and the sub-image size is the same, and there is no need to rotate the original face sub-image; otherwise, the size relationship between the width and height of the original face sub-image and the sub-image size is different, and the original face sub-image needs to be rotated by 90 degrees (either clockwise or counterclockwise). For example, when the sub-image size is 320 in width * 448 in height, that is, the width is less than the height, if the original face sub-image is in the case where the width is greater than the height, then rotate the original face sub-image by 90 degrees.
[0148] In one embodiment, in order to maintain the angle of the face in the original face sub-image, the original face sub-image may not be rotated either.
[0149] ② When the size of the original face sub-image is larger than the sub-image size, downsample the original face sub-image according to the sub-image size. Among them, the size of the original face sub-image being larger than the sub-image size means that the width of the original face sub-image is greater than the width of the sub-image size, or the height of the original face sub-image is greater than the height of the sub-image size. In the image beauty scene, the original image to be beautified is generally a clear image captured by a terminal device, and its size is relatively large. Therefore, it is a relatively common situation that the size of the original face sub-image is larger than the sub-image size, that is, usually the original face sub-image needs to be downsampled.
[0150] Downsampling can be implemented by methods such as bilinear interpolation and nearest neighbor interpolation, and the present disclosure does not limit this.
[0151] After downsampling, at least one of the width and height of the original face sub-image is aligned with the sub-image size, which specifically includes the following situations:
[0152] Both the width and height of the original face sub-image are the same as the sub-image size;
[0153] The width of the original face sub-image is the same as the width of the sub-image size, and the height is less than the height of the sub-image size;
[0154] The height of the original face sub-image is the same as the height of the sub-image size, and the width is less than the width of the sub-image size.
[0155] It should be noted that if the above rotation has been performed on the original face sub-image to obtain a rotated original face sub-image, then when the size of the original face sub-image is larger than the sub-image size, it is downsampled according to the sub-image size, and the specific implementation method is the same as the downsampling method of the above original face sub-image, so it will not be elaborated here.
[0156] On the contrary, when the size of the original face sub-image (or the rotated original face sub-image) is less than or equal to the sub-image size, the processing step of downsampling can be omitted.
[0157] ③ When the size of the original face sub-image is less than the sub-image size, the original face sub-image is filled according to the difference between the original face sub-image and the sub-image size, so that the size of the filled original face sub-image is equal to the sub-image size. Among them, the size of the original face sub-image being less than the sub-image size means that at least one of the width and height of the original face sub-image is less than the sub-image size, and the other is not greater than the sub-image size, which specifically includes the following situations:
[0158] The width of the original face sub-image is less than the width of the sub-image size, and the height is also less than the height of the sub-image size;
[0159] The width of the original face sub-image is less than the width of the sub-image size, and the height is equal to the height of the sub-image size;
[0160] The height of the original face sub-image is less than the height of the sub-image size, and the width is equal to the width of the sub-image size.
[0161] When filling, a preset pixel value can be used, usually a pixel value with a large color difference from the human face, such as (R0, G0, B0), (R255, G255, B255), etc.
[0162] Generally, it can be filled around the original face sub-image. For example, the center of the original face sub-image is made to coincide with the center of the sub-image size, and the difference part around the original face sub-image is filled so that the size of the original face sub-image after filling is the same as the sub-image size. Of course, it is also possible to align the original face sub-image with one side edge of the sub-image size and fill the other side. The present disclosure does not limit this.
[0163] It should be noted that if at least one of the above-mentioned rotation and downsampling processes has been performed on the original face sub-image to obtain the original face sub-image that has undergone at least one of the rotation and downsampling processes, then when the size of this original face sub-image is smaller than the sub-image size, it is filled according to the difference between it and the sub-image size. The specific implementation method is the same as the filling method of the above-mentioned original face sub-image, so it will not be elaborated here.
[0164] The above ① - ③ are three commonly used transformation methods, and any one or more of them can be used according to actual needs. For example, each original face sub-image is processed successively by ①, ②, and ③, and the processed original face sub-images are combined into the face image to be beautified.
[0165] In the above transformation, the direction, size, etc. of the original face sub-image are changed, which is for the convenience of unified processing by the deep neural network. Subsequently, an inverse transformation needs to be performed on the beautified face image to make it restored to be consistent with the direction, size, etc. of the original face sub-image to adapt to the size of the image to be processed. Therefore, the corresponding transformation information can be saved, including but not limited to: the direction and angle of rotation of each original face sub-image, the downsampling ratio, and the coordinates of the filled pixels. This facilitates subsequent inverse transformation according to this transformation information.
[0166] After combining the transformed original face sub-images, the combination information can be saved, including but not limited to the size of each original face sub-image (i.e., the corresponding sub-image size) and its position in the face image to be beautified, the arrangement method and order of each original face sub-image. Subsequently, the beautified face image can be split according to this combination information to obtain each individual beautified face sub-image.
[0167] The above has described how to obtain the face image to be beautified. Continuing to refer to Figure 3 , in step S320, based on the pre-trained deep neural network, features based on a three-dimensional grid are extracted from the face image to be beautified, and an information matrix is generated according to the extracted features. The three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and pixel value domain of the face image to be beautified.
[0168] Among them, the spatial domain of the face image to be beautified is the two-dimensional space where the image plane of the face image to be beautified is located, having two dimensions. The first dimension can be, for example, the width direction of the image, and the second dimension can be, for example, the height direction of the image. The pixel value range refers to the numerical range of the pixel values of the face image to be beautified. For example, it can be [0, 255], or if the pixel values are normalized, the pixel value range is [0, 1]. Taking the pixel value range as the third dimension, together with the above-mentioned first dimension and second dimension, a three-dimensional space is formed. In this exemplary embodiment, the three-dimensional space can be pre-divided, including dividing the spatial domain and dividing the pixel value range, to obtain a three-dimensional grid. The two-dimensional projection of the three-dimensional grid on the spatial domain is called the spatial domain grid; the one-dimensional projection of the three-dimensional grid on the pixel value range is called the value range partition. Exemplarily, a region of 16 pixels * 16 pixels can be used as the spatial domain grid, and [0, 1 / 8), [1 / 8, 1 / 4), [1 / 4, 3 / 8), etc. (dividing [0, 1] into 8 partitions evenly) can be used as the value range partition, so as to obtain a three-dimensional grid. Thus, features based on the three-dimensional grid can be extracted from the face image to be beautified, and an information matrix can be generated according to the extracted features. The information matrix is a parameter matrix for beautifying the face image to be beautified.
[0169] In one embodiment, the structure of the deep neural network can refer to Figure 7 as shown, including four main parts: a basic convolutional layer, a grid feature convolutional layer, a local feature convolutional layer, and an output layer. Each part can further include multiple intermediate layers. The grid feature convolutional layer and the local feature convolutional layer are two parallel parts between the basic convolutional layer and the output layer.
[0170] Referring to Figure 8 as shown, the above-mentioned process of extracting features based on the three-dimensional grid from the face image to be beautified by the pre-trained deep neural network and generating an information matrix according to the extracted features can include the following steps S810 to S840:
[0171] Step S810, perform downsampling convolution processing on the face image to be beautified according to the size of the spatial domain grid through the basic convolutional layer to obtain a basic feature image.
[0172] Downsampling convolution processing refers to reducing the image size through convolution to achieve the downsampling effect. A convolutional layer with a stride greater than 1 can be used to implement downsampling convolution processing. Combining Figure 7For example, the dimension of the face image to be processed is (B, W, H, C). B represents the number of images, and one or more face images to be beautified can be input into the deep neural network for processing. Therefore, B can be any positive integer. W represents the image width, H represents the image height, and C represents the number of image channels. When the face image to be beautified is an RGB image, C is 3. The size of the spatial grid is 16 pixels * 16 pixels. The basic convolutional layer may include four 3*3 convolutional layers with a stride of 2 (3*3 represents the size of the convolutional kernel, which is only exemplary and can be replaced with other sizes). After being processed by it, the height and width of the face image to be beautified are both reduced to 1 / 16. Of course, in the present disclosure, convolutional layers with other numbers and strides can also be set to achieve the same downsampling effect. For example, the above four 3*3 convolutional layers with a stride of 2 can be replaced with two 5*5 convolutional layers with a stride of 4. In addition, the basic convolutional layer may further include one or more 3*3 convolutional layers with a stride of 1 (3*3 represents the size of the convolutional kernel, which is only exemplary and can be replaced with other sizes) for further extracting features from the image after downsampling convolution without changing the image size, resulting in a basic feature image. Of course, setting a convolutional layer with a stride of 1 is not necessary. The dimension of the basic feature image is (B, W / 16, H / 16, k1), where k1 is the number of channels of the basic feature image, representing the dimension of the features, which is related to the number of convolutional kernels of the last convolutional layer in the basic convolutional layer and is not limited in the present disclosure. A pixel point in the basic feature image corresponds to 16 pixels * 16 pixels in the face image to be beautified.
[0173] As can be seen from the above, the processing process of the basic convolutional layer is to gradually extract features within the range of each spatial grid in the face image to be beautified, represent features of different dimensions in different channels through the setting of the convolutional kernel, and finally obtain a basic feature image, which is the feature image of the face image to be beautified at the scale of the spatial grid.
[0174] Step S820: Extract the features within the spatial grid from the basic feature image through the grid feature convolutional layer to obtain a grid feature image.
[0175] The basic feature image extracted by the basic convolutional layer reflects the basic features in the face image to be beautified. The grid feature convolutional layer can further extract deeper features within the range of the spatial grid to obtain a grid feature image. Combining Figure 7For example, the grid feature convolution layer may include one or more 3*3 convolution layers with a stride of 1 (3*3 represents the convolution kernel size, which is only exemplary and can also be replaced with other sizes) for further extracting grid features from the basic convolution image without changing the image size, thereby obtaining a grid feature image. The dimension of the grid feature image is (B, W / 16, H / 16, k2), where k2 is the number of channels of the grid feature image, representing the dimension of the feature, and is related to the number of convolution kernels of the last convolution layer in the grid feature convolution layer. This disclosure does not make any limitations in this regard.
[0176] Step S830: Extract features between spatial domain grids from the basic feature image through the local feature convolution layer to obtain a local feature image.
[0177] Based on the basic feature image extracted by the basic convolution layer, the local feature convolution layer can further extract deeper features between spatial domain grids to obtain a local feature image. Compared with the above-mentioned grid feature image, the local feature image is the feature within a local range of multiple spatial domain grids, and its scale is relatively larger. Figure 7 For example, the local feature convolution layer may include a downsampling layer, one or more 3*3 convolution layers with a stride of 1 (3*3 represents the convolution kernel size, which is only exemplary and can also be replaced with other sizes), and an upsampling layer. Among them, the downsampling layer may be a 2*2 (or other size) pooling layer, and maximum pooling, average pooling, etc. can be used. The features within a local range of 2*2 spatial domain grids are fused through pooling; of course, downsampling can also be implemented in other ways (such as convolution with a stride greater than 1). Furthermore, features are extracted from the downsampled feature image through the convolution layer, which are the features between spatial domain grids. The upsampling layer may be a 2*2 (or other size) transposed convolution layer, which realizes upsampling by performing transposed convolution on the feature image output by the convolution layer, thereby restoring the image size before downsampling (i.e., W / 16*H / 16); of course, upsampling can also be implemented in other ways (such as interpolation). After upsampling, a local feature image is obtained, and its dimension is (B, W / 16, H / 16, k3), where k3 is the number of channels of the local feature image, representing the dimension of the feature, and is related to the number of convolution kernels of the last convolution layer in the local feature convolution layer. This disclosure does not make any limitations in this regard.
[0178] Step S840: Perform dimensionality conversion on the grid feature image and the local feature image according to the number of value range partitions through the output layer to obtain an information matrix.
[0179] The grid feature image and the local feature image reflect the features of the face image to be beautified at different scales. The output layer can merge the grid feature image and the local feature image and then perform dimensional transformation. The merging methods include, but are not limited to, addition, concatenation (concat), etc. As described above, in the dimensions of the grid feature image and the local feature image, the image sizes W / 16 and H / 16 correspond to the size of the spatial grid, while the number of channels k2 and k3 is related to the number of convolutional kernels. Through dimensional transformation, the number of channels is matched with the number of value range partitions, so that the output information matrix corresponds to the three-dimensional grid. Combining Figure 7 For example, the output layer can include a concatenation layer and one or more 1*1 convolutional layers with a stride of 1 (1*1 represents the size of the convolutional kernel, which is only exemplary and can also be replaced with other sizes). The concatenation layer is used to concatenate the grid feature image and the local feature image to obtain a concatenated feature image with a dimension of (B, W / 16, H / 16, k2 + k3). The convolutional layer is used to perform dimensional transformation on the concatenated feature image to obtain an information matrix G with a dimension of (B, W / 16, H / 16, G_z * G_n). G_z is the number of value range partitions. For example, when dividing the pixel value range into 8 equal parts during the division of the three-dimensional grid, G_z is 8; G_n is the dimension of the sub-information matrix gi corresponding to each three-dimensional grid (i.e., the number of elements in the sub-information matrix gi), and i represents the ordinal number of the three-dimensional grid.
[0180] In one implementation, the information matrix G obtained in step S840 can be regarded as a set of sub-information matrices gi. For each face image to be beautified, the deep neural network can output its corresponding information matrix G, including W / 16 * H / 16 * G_z sub-information matrices gi, and W / 16 * H / 16 * G_z is exactly the number of three-dimensional grids, that is, the information matrix G includes the sub-information matrix gi corresponding to each three-dimensional grid.
[0181] The above illustrates how to obtain the information matrix. Continuing to refer to Figure 3 , in step S330, the face image to be beautified is processed using the information matrix to obtain the beautified face image corresponding to the face image to be beautified.
[0182] Generally, the pixel values of the face image to be beautified can be multiplied by the information matrix to achieve numerical transformation of the pixel values and obtain the beautified face image.
[0183] In one implementation, the information matrix can include a reference information matrix corresponding to each three-dimensional grid, and this reference information matrix is equivalent to the above-mentioned sub-information matrix gi. Referring to Figure 9 As shown, the above-mentioned process of processing the face image to be beautified using the information matrix to obtain the beautified face image corresponding to the face image to be beautified can include the following steps S910 and S920:
[0184] Step S910: Interpolate the reference information matrix based on the face image to be beautified to obtain a beautification information matrix corresponding to each pixel point of the face image to be beautified.
[0185] The reference information matrix can be the reference information for beautifying all pixel points within a three-dimensional grid, which can be regarded as a summary of the information required for beautifying all pixel points within the three-dimensional grid. The beautification information matrix is the specific information for beautifying each pixel point. The reference information matrix can further correspond to the reference point of the three-dimensional grid. For example, the reference point can be the center point of the three-dimensional grid. Since each pixel point of the face image to be beautified is distributed at different positions within its respective three-dimensional grid and has an offset relative to the reference point within the three-dimensional grid, the reference information matrix can be interpolated to obtain a beautification information matrix corresponding to each pixel point of the face image to be beautified.
[0186] In one implementation, interpolation can be performed on one or more reference information matrices according to the offset of each pixel point of the face image to be beautified relative to the center point of one or more three-dimensional grids to obtain a beautification information matrix corresponding to each pixel point of the face image to be beautified. Exemplarily, assume that the width of the face image to be beautified is 128 and the height is also 128, and the size of the spatial domain grid is 16 pixels * 16 pixels. Then both the first dimension and the second dimension of the three-dimensional space are equally divided into 8 parts; the pixel value range [0, 1] is also equally divided into 8 value range partitions, so the three-dimensional space is divided into 8 * 8 * 8 three-dimensional grids. Represent the three-dimensional grid located at the upper left corner of the face image to be beautified with {0, 0, 0} and having a pixel value in the range [0, 1 / 8). The center point coordinates of this three-dimensional grid are (8, 8, 1 / 16); obtain the pixel points within this three-dimensional grid in the face image to be beautified, calculate the offset of each pixel point from the center point, including the offset amounts in the first dimension, the second dimension, and the third dimension, and perform trilinear interpolation based on the reference information matrix of the {0, 0, 0} three-dimensional grid and the reference information matrices of its adjacent three-dimensional grids {1, 0, 0}, {0, 1, 0}, {0, 0, 1} according to the offset amounts to obtain a beautification information matrix corresponding to each pixel point in the {0, 0, 0} three-dimensional grid. It should be noted that if the three-dimensional grid is not on the boundary, trilinear interpolation can be performed based on the reference information matrix of this three-dimensional grid and the reference information matrices of its adjacent 6 three-dimensional grids to obtain a beautification information matrix corresponding to each pixel point in this three-dimensional grid.
[0187] It should be understood that the present disclosure does not limit the specific interpolation algorithm. For example, a non-linear interpolation algorithm can also be used.
[0188] As can be seen from the above, when performing interpolation, it is necessary to calculate the offset between the pixel value of the pixel point and the pixel value of the reference point, that is, the offset between the pixel point and the reference point in the third dimension. When the face image to be beautified is a single-channel image, the pixel value of the face image to be beautified can be directly used for calculation. When the face image to be beautified is a multi-channel image, it is difficult to calculate based on the pixel values of the multi-channel and the pixel values of the reference point. Based on this, in one implementation, the above-mentioned interpolation of the reference information matrix based on the face image to be beautified to obtain the beautification information matrix corresponding to each pixel point of the face image to be beautified may include the following steps:
[0189] When the face image to be beautified is a multi-channel image, convert the face image to be beautified into a single-channel reference value image;
[0190] Interpolate the reference information matrix based on the reference value image to obtain the beautification information matrix corresponding to each pixel point of the face image to be beautified.
[0191] Among them, the reference value image is an image that represents the multi-channel of the face image to be beautified through a single channel. For example, when the face image to be beautified is an RGB image, the reference value image can be its corresponding grayscale image, and the grayscale can adopt a normalized value, and the value range is [0, 1]. Note that the reference value image and the above-mentioned reference frame image are different concepts.
[0192] In one implementation, the following formula can be used to convert the face image to be beautified into a single-channel reference value image:
[0193]
[0194] Among them, R, G, and B are the normalized pixel values of each pixel point in the face image to be beautified; n represents dividing the value ranges of R, G, and B into n partitions, and j represents the ordinal number of the partition; a rj 、a gj 、a bj are the conversion coefficients of each partition of R, G, and B respectively, which can be determined according to experience or actual needs; shift rj 、shift gj 、shift bj are the conversion thresholds set in each partition of R, G, and B respectively, indicating that only the pixel values greater than the conversion threshold are converted, and the conversion threshold can be set according to experience or actual needs; guidemap r 、guidemap g 、guidemap b are the single-channel images of R, G, and B after partition conversion respectively; g r 、g g 、g bThey are the fusion coefficients of R, G, and B respectively, which can be empirical coefficients; guidemap bias is the offset added after fusion, which can also be determined according to experience; guidemap z is the reference value image, and its value range is [0, 1].
[0195] In one implementation, the above-mentioned a can also be obtained through pre-set model training rj 、a gj 、a bj 、shift rj 、shift gj 、shift bj 、g r 、g g 、g b 、guidemap bias and other parameters. By setting the initial values of the model, the value range of the finally obtained reference value image satisfies [0, 1].
[0196] Reference Figure 7 As shown, interpolation can be performed on the benchmark information matrix in the information matrix G based on the reference value image to obtain the beauty information matrix corresponding to each pixel point of the face image to be beautified. These beauty information matrices can be used as a set, and its dimension is (B, W, H, G_n).
[0197] Step S920: According to the beauty information matrix corresponding to each pixel point of the face image to be beautified, process each pixel point of the face image to be beautified respectively to obtain the beautified face image.
[0198] The pixel value of each pixel point can be multiplied by the corresponding beauty information matrix to obtain the processed pixel value, thereby forming the beautified face image. Exemplarily, the pixel value of pixel point i is represented as the pixel value vector [r, g, b], and its corresponding beauty information matrix is:[[]]
[0199]
[0200] Then there is the following relationship:[[]]
[0201]
[0202] Among them, [r′ g′ b′] represents the beautified pixel value.
[0203] In one implementation, the above-mentioned process of processing each pixel point of the face image to be beautified respectively according to the beauty information matrix corresponding to each pixel point of the face image to be beautified to obtain the beautified face image may include:[[]]
[0204] Add a new channel to the face image to be beautified according to the dimension of the beauty information matrix, and set the new channel to a preset value;
[0205] Multiply the pixel value vector of each pixel of the face image to be beautified by the beauty information matrix corresponding to each pixel to obtain the beautified face image; the pixel value vector of each pixel is a vector formed by the values of each channel of each pixel.
[0206] Among them, the dimension of the beauty information matrix represents the number of rows and columns of the beauty information matrix. It can be seen from formula (13) that it is necessary to perform a cross product operation on the pixel value vector of each pixel and the beauty information matrix, indicating that the dimension of the pixel value vector needs to be the same as the number of rows of the beauty information matrix. And the dimension of the pixel value vector is equal to the number of channels of the face image to be beautified. Therefore, if the number of channels of the face image to be beautified is not equal to (generally less than) the number of rows of the beauty information matrix, a new channel can be added to the face image to be beautified. For the added new channel, a preset value can be filled, such as 1. Thus, it is equivalent to converting the pixel value vector of each pixel in the face image to be beautified into a homogeneous vector.
[0207] Exemplarily, assume that the beauty information matrix corresponding to pixel point i is:
[0208]
[0209] That is, the number of rows of this beauty information matrix is 4, and the face image to be beautified is an RGB image with 3 channels. Therefore, a new channel needs to be added, and the new channel is uniformly filled with the value 1. Then the pixel value vector of pixel point i is [r, g, b, 1], so as to satisfy the following relationship:
[0210]
[0211] Thus, through the processing of the information matrix, a beautified face image is obtained, and its dimension is (B, W, H, C). In formulas (13) and (14), C = 3. The dimension of the beautified face image is the same as that of the face image to be beautified, indicating that the beautification processing process of this exemplary embodiment does not change the image dimension.
[0212] If the pixel values are normalized before the face image to be beautified is input into the deep neural network, after obtaining the beautified face image, the pixel values can be denormalized. For example, the pixel values in the range of [0, 1] can be uniformly multiplied by 255 to obtain pixel values in the range of [0, 255].
[0213] In one embodiment, the image beautification processing method may further include the training process of the deep neural network. Refer to Figure 10 As shown, it may include the following steps S1010 to S1030:
[0214] Step S1010: Input the sample image to be beautified into the deep neural network to be trained, so as to output a sample information matrix.
[0215] Step S1020: Process the sample image to be beautified by using the sample information matrix to obtain a beautified sample image corresponding to the sample image to be beautified.
[0216] Step S1030: Update the parameters of the deep neural network based on the difference between the labeled image corresponding to the sample image to be beautified and the beautified sample image.
[0217] The deep neural network can indirectly implement the combination of different beautification functions. In this exemplary embodiment, a beautified image dataset corresponding to different beautification functions can be obtained according to actual needs to train the required deep neural network. For example, if it is necessary to train a deep neural network for blemish removal, a sample image to be beautified with blemishes is obtained, and through manual blemish removal processing, the corresponding labeled image (Ground truth) is obtained, thereby constructing a beautified image dataset for blemish removal; if it is necessary to train a deep neural network for blemish removal + deformation, a sample image to be beautified with blemishes is obtained, and through manual blemish removal and deformation processing, the corresponding labeled image is obtained, thereby constructing a beautified image dataset for blemish removal + deformation. Of course, it is also possible to first obtain the labeled image and, through reverse processing, obtain the sample image to be beautified. For example, an image of a flawless face is obtained, and it is processed by adding blemishes, reverse deformation (which refers to the processing opposite to the deformation in beautification. For example, in beautification, "face slimming" deformation is often performed, and reverse deformation can be to widen the face), etc., to obtain the sample image to be beautified, and the image of the flawless face is used as its corresponding labeled image to construct a beautified image dataset for blemish removal + deformation. It can be seen that in this exemplary embodiment, different beautified image datasets can be constructed to train a deep neural network for any one or more combinations of beautification functions.
[0218] In one embodiment, multiple face images can be combined to obtain a sample image to be beautified, and the artificially beautified images corresponding to the multiple face images can be combined to obtain a labeled image corresponding to the sample image to be beautified, and then the sample image to be beautified and the labeled image are added to the beautified image dataset. In other words, the beautified image dataset can include different types such as single-face images, multi-face images, and combined-face images.
[0219] The structure of the deep neural network can refer to the above Figure 7The content of this part will not be elaborated further. The sample image to be beautified is input into the deep neural network, and the corresponding sample information matrix is output. Then, the sample information matrix is used to process the sample image to be beautified, and the beautified sample image corresponding to the sample image to be beautified is obtained. The processing process can refer to the content of steps S320 and S330. Since the deep neural network is not trained or not fully trained at this time, the sample information matrix cannot process the sample image to be beautified with high quality, resulting in a difference between the obtained beautified sample image and the labeled image. Based on this difference, a loss function can be constructed, and then the parameters of the deep neural network can be updated by backpropagation according to the loss function value to realize the training of the deep neural network. Generally, when the accuracy rate or other metrics of the deep neural network on the test set in the beautified image dataset reach the preset standard, it is determined that the training is completed.
[0220] Based on Figure 10 As can be seen from the method steps shown above, the dataset used to train the deep neural network in this exemplary embodiment can be an ordinary beautified image dataset. The sample images to be beautified and their corresponding labeled images in it are relatively easy to obtain, and there is no need to specifically set labels for the information matrix, making the solution highly practical.
[0221] In one embodiment, if the face image to be beautified is composed of the original face sub-images in the original image to be beautified, after obtaining the beautified face image, the beautified face sub-images corresponding to the original face sub-images can be split from the beautified face image. The beautified face sub-image is the beautified image corresponding to a single face. Among them, when splitting the combined beautified face image, the above-saved combined information can be used to split sub-images with specific positions and specific sizes from the combined beautified face image, that is, the beautified face sub-images, and the beautified face sub-images correspond one-to-one with the original face sub-images.
[0222] In one embodiment, the original face sub-images in the original image to be beautified can be replaced with the corresponding beautified face sub-images to obtain the target beautified image corresponding to the original image to be beautified, thereby realizing the beautification processing of the face in the original image to be beautified.
[0223] In one embodiment, if the original face sub-images are transformed when combining them into the face image to be beautified, the inverse transformation can be correspondingly performed on the obtained beautified face sub-images, including removing the filled pixels, upsampling, rotating 90 degrees backward, etc., so that the direction, size, etc. of the inverse-transformed beautified face sub-images are consistent with those of the original face sub-images, so that a 1:1 replacement can be performed in the original image to be beautified to obtain the target beautified image.
[0224] The beautified face sub-image is the face sub-image after beautification processing, usually the face sub-image with a relatively high degree of beautification. In one implementation, in order to increase the realism of the beautified face sub-image, before replacing the original face sub-image in the original image to be beautified with the corresponding beautified face sub-image, the beautified face sub-image can be subjected to beautification weakening processing using the original face sub-image. Beautification weakening processing refers to reducing the degree of beautification of the beautified face sub-image to increase... Two exemplary methods of beautification weakening processing are provided below:
[0225] Method 1: According to the set beautification degree parameter, the original face sub-image is fused into the beautified face sub-image. Among them, the beautification degree parameter can be the beautification intensity parameter under a specific beautification function, such as the degree of blemish removal. In this exemplary implementation, the beautification degree parameter can be the parameter set for the current time, the system default parameter, or the parameter used in the previous beautification, etc. After determining the beautification degree parameter, the original face sub-image and the beautified face sub-image can be fused with the beautification degree parameter as the ratio. For example, assuming that the range of the degree of blemish removal is 0-100 and the currently set value is a, refer to the following formula:
[0226]
[0227] where image_blend represents the fused image, image_ori represents the original face sub-image, and image_deblemish represents the beautified face sub-image. When a is 0, it means no blemish removal is performed, and the original face sub-image is completely used; when a is 100, it means complete blemish removal, and the beautified face sub-image is completely used. Therefore, formula (15) represents that through fusion, an image between the original face sub-image and the beautified face sub-image is obtained. The larger a is, the closer the obtained image is to the beautified face sub-image, that is, the higher the degree of beautification and the more obvious the beautification effect.
[0228] It should be noted that if the original face sub-image is transformed when combining the original face sub-images into the face image to be beautified, the inverse transformation can be performed on the split beautified face sub-image. The original face sub-image and the beautified face sub-image have the following relationship: the original face sub-image before transformation is consistent with the beautified face sub-image after inverse transformation in terms of direction, size, etc.; the original face sub-image after transformation is consistent with the beautified face sub-image before inverse transformation in terms of direction, size, etc. Therefore, when fusing the original face sub-image and the beautified face sub-image using the above formula (15), the original face sub-image before transformation and the beautified face sub-image after inverse transformation can be fused, or the original face sub-image after transformation and the beautified face sub-image before inverse transformation can be fused.
[0229] Method 2: Fuse the high-frequency image of the original face sub-image into the beautified face sub-image. The high-frequency image refers to an image containing high-frequency information such as detailed textures in the original face sub-image.
[0230] In one implementation, the high-frequency image can be obtained in the following way:
[0231] When combining one or more of the above original face sub-images based on the input image size of the deep neural network, if the original face sub-image is downsampled, then the downsampled face sub-image obtained after downsampling is upsampled to obtain an upsampled face sub-image;
[0232] According to the difference between the original face sub-image and the upsampled face sub-image, obtain the high-frequency image of the original face sub-image.
[0233] Among them, the resolution of the downsampled face sub-image is lower than that of the original face sub-image. Generally, in the process of downsampling, it is inevitable to lose the high-frequency information of the image. Upsample the downsampled face sub-image so that the obtained upsampled face sub-image has the same resolution as the original face sub-image. It should be noted that if the original face sub-image is rotated before downsampling, after upsampling the downsampled face sub-image, it can also be rotated in the reverse direction so that the obtained upsampled face sub-image has the same direction as the original face sub-image.
[0234] Upsampling can use methods such as bilinear interpolation and nearest neighbor interpolation. Although upsampling can restore the resolution, it is difficult to completely restore the lost high-frequency information. That is, the upsampled face sub-image can be regarded as the low-frequency image of the original face sub-image. Therefore, determine the difference between the original face sub-image and the upsampled face sub-image. For example, the original face sub-image can be subtracted from the upsampled face sub-image, and the result is the high-frequency information of the original face sub-image. Form an image with the subtracted value, that is, the high-frequency image of the original face sub-image.
[0235] In another implementation, the high-frequency information can also be extracted by filtering the original face sub-image to obtain a high-frequency image.
[0236] When fusing the above high-frequency image into the beautified face sub-image, the direct addition method can be used to superimpose the high-frequency image onto the beautified face sub-image, so that the beautified face sub-image increases high-frequency information such as detailed textures and is more realistic.
[0237] Since the original face sub-image and the upsampled face sub-image are usually very similar, in the high-frequency image obtained based on their difference, the pixel values are generally small. For example, the values of each RGB channel do not exceed 4. However, for the mutation positions in the original face sub-image, such as small moles on the face, they have strong high-frequency information. Therefore, the pixel values at the corresponding positions in the high-frequency image may be relatively large. When fusing the high-frequency image to the original face sub-image, the pixel values at these positions may have an adverse effect, such as generating sharp edges like "mole marks", resulting in an unnatural visual experience.
[0238] In view of the above problems, in one implementation, the image beauty processing method may further include the following steps:
[0239] Determine defect points in the high-frequency image;
[0240] Adjust the pixel values within a preset area around the above defect points in the high-frequency image to a preset value range.
[0241] Among them, the defect points are pixel points with strong high-frequency information, and the points with relatively large pixel values in the high-frequency image can be determined as defect points. Or, in one implementation, the defect points can be determined in the following way:
[0242] Subtract the beautified face sub-image from the corresponding original face sub-image to obtain the difference of each pixel point;
[0243] When it is determined that the difference of a certain pixel point meets the preset defect condition, the pixel point corresponding to this pixel point in the high-frequency image is determined as a defect point.
[0244] Among them, the preset defect condition is used to measure the difference between the beautified face sub-image and the original face sub-image to determine whether each pixel point is a defect point to be removed. In the defect removal process, small moles, pimples, etc. on the face are usually removed and the skin color of the face is filled. At this position, the difference between the beautified face sub-image and the original face sub-image is very large. Therefore, the preset defect condition can be set to identify defect points.
[0245] Exemplarily, the preset defect condition may include: the differences of all channels are greater than the first color difference threshold, and at least one of the differences of all channels is greater than the second color difference threshold. The first color difference threshold and the second color difference threshold may be empirical thresholds. For example, when the above channels include RGB, the first color difference threshold may be 20, and the second color difference threshold may be 40. Thus, after obtaining the difference between each pixel point in the beautified face sub-image and in the original face sub-image, the specific differences of the three RGB channels in the difference are judged to determine whether the differences of each channel are all greater than 20, and whether at least one of the channel differences is greater than 40. When these two conditions are met, it indicates that the preset defect condition is satisfied, and the pixel points at the corresponding positions in the high-frequency image are determined as defect points.
[0246] After determining the defect points, a preset area around the defect points can be further determined in the high-frequency image. For example, it can be a 5*5 pixel area centered on the defect point, and the specific size can be determined according to the size of the high-frequency image, which is not limited in this disclosure. The pixel values within the preset area are adjusted to a preset numerical range. The preset numerical range is generally a relatively small numerical range, which can be determined according to experience and actual needs. When adjusting, the pixel values usually need to be decreased. Exemplarily, the preset numerical range can be -2 to 2, and the pixel values around the defect points may exceed -5 to 5. Adjusting them to -2 to 2 actually performs a limiting process. Thereby, sharp edges such as "mole marks" can be weakened, and a more natural visual feeling can be increased.
[0247] The above describes two beautification weakening processing methods. This exemplary embodiment can adopt these two beautification weakening processing methods simultaneously. For example, first, the original face sub-image and the beautified face sub-image are fused through Method 1. On this basis, then the high-frequency image is superimposed through Method 2 to obtain a beautified face sub-image after beautification weakening processing. This beautified face sub-image has both a good beautification effect and a sense of reality.
[0248] In one embodiment, when replacing the original face sub-image in the image to be processed with the corresponding beautified face sub-image, the following steps may also be performed:
[0249] Perform a fade processing on the boundary area between the un-replaced area in the original image to be beautified and the beautified face sub-image, so that the boundary area forms a smooth transition.
[0250] Among them, the un-replaced area in the original image to be beautified is the area in the original image to be beautified except for the original face sub-image. The boundary area between the above un-replaced area and the beautified face sub-image actually includes two parts: the boundary area in the un-replaced area adjacent to the beautified face sub-image, and the boundary area in the beautified face sub-image adjacent to the un-replaced area. In this exemplary embodiment, gradient processing can be performed on any one of the two parts, or gradient processing can be performed on both parts simultaneously.
[0251] Reference Figure 11 As shown, a certain proportion (such as 10%) of the boundary area can be determined in the beautified face sub-image, which extends inward from the edge of the beautified face sub-image. It should be noted that the boundary area usually needs to avoid the face part to prevent changing the color of the face part during the gradient processing. For example, by intercepting the original face sub-image through the above stable bounding box, a certain distance is provided between the face in the original face sub-image and the boundary, and thus a certain distance is also provided between the face in the beautified face sub-image and the boundary. In this way, during the gradient processing, the face part can be better avoided. After determining the boundary area, obtain the inner edge color of the boundary area, denoted as the first color; obtain the inner edge color of the un-replaced area, denoted as the second color; then perform gradient processing on the boundary area with the first color and the second color. Thus, the boundary between the un-replaced area and the beautified face sub-image is a gradient color area ( Figure 11 the diagonal area in), so as to form a smooth transition and prevent color mutation, resulting in an unharmonious visual experience.
[0252] It should be noted that when there are multiple beautified face sub-images, each beautified face sub-image can be used to replace the corresponding original face sub-image in the image to be processed, and gradient processing of the boundary area is performed to obtain a target beautified image, making it have a natural and harmonious visual experience.
[0253] In one embodiment, after obtaining the beautified face image, the beautified face image can be subjected to beautification weakening processing using the face image to be beautified. The beautification weakening processing can specifically refer to the above two methods, and thus will not be elaborated here. When the face image to be beautified includes only one face or the face image to be beautified is equivalent to the original image to be beautified, the beautification weakening processing can be directly performed without splitting the beautified face image. When the face image to be beautified is composed of multiple original face sub-images, the beautified face image can also be first subjected to beautification weakening processing as a whole, and then split into beautified face sub-images, so as to avoid performing beautification weakening processing on each beautified face sub-image separately and improve efficiency.
[0254] Figure 12 shows a schematic flow of the image beautification processing method, including:
[0255] Step S1201: Obtain the original image to be beautified. For example, the current frame image in the video can be used as the original image to be beautified.
[0256] Step S1202: Extract the original face sub-image from the original image to be beautified.
[0257] Step S1203: Divide the input image size of the deep neural network into multiple sub-image sizes according to the number of original face sub-images. Downsample the original face sub-images according to the sub-image sizes, and rotation, padding, etc. can also be performed to obtain the downsampled face sub-images corresponding to each original face sub-image.
[0258] Step S1204: Upsample the downsampled face sub-images. If rotation, padding, etc. were also performed when obtaining the downsampled face sub-images, reverse rotation, removing padding, etc. can also be performed to obtain the upsampled face sub-images, whose resolution is the same as that of the corresponding original face sub-images.
[0259] Step S1205: Subtract the original face sub-image from the corresponding upsampled face sub-image to obtain the high-frequency image of the original face sub-image.
[0260] Step S1206: Combine the downsampled face sub-images into an image to be beautified.
[0261] Step S1207: Input the image to be beautified into the deep neural network and output the information matrix.
[0262] Step S1208: Process the image to be beautified using the information matrix to obtain the beautified face image.
[0263] Step S1209: Split the beautified face image into beautified face sub-images corresponding one by one to the original face sub-images.
[0264] Step S1210: Fuse the beautified face sub-images and the corresponding original face sub-images according to the beautification degree parameter, and then add the high-frequency image of the original face sub-image to obtain the face sub-image to be replaced.
[0265] Step S1211: Fuse the face sub-image to be replaced into the original image to be beautified. Specifically, the part of the original face sub-image in the original image to be beautified can be replaced by the face sub-image to be replaced, and color gradient processing at the edge can be performed so that the face in the original image to be beautified is replaced by the beautified face, and finally the target beautified image is obtained. Subsequent personalized beautification processing can also be performed.
[0266] An exemplary embodiment of the present disclosure also provides an image beautification processing device. Referring to Figure 13 As shown, the image beautification processing device 1300 may include:
[0267] An image acquisition module 1310, configured to acquire a face image to be beautified;
[0268] An information matrix generation module 1320, configured to extract 3D grid-based features from the face image to be beautified through a pre-trained deep neural network, and generate an information matrix according to the extracted features, where the 3D grid is obtained by dividing the 3D space formed by the spatial domain and the pixel value domain of the face image to be beautified;
[0269] A beauty processing module 1330, configured to process the face image to be beautified by using the information matrix to obtain a beautified face image corresponding to the face image to be beautified.
[0270] In one implementation, the deep neural network includes a basic convolutional layer, a grid feature convolutional layer, a local feature convolutional layer, and an output layer. The above-mentioned extraction of 3D grid-based features from the face image to be beautified through a pre-trained deep neural network and generation of an information matrix according to the extracted features include:
[0271] Performing downsampling convolution processing on the face image to be beautified by the basic convolutional layer according to the size of the spatial grid, to obtain a basic feature image, where the spatial grid is the 2D projection of the 3D grid in the spatial domain;
[0272] Extracting features within the spatial grid from the basic feature image by the grid feature convolutional layer to obtain a grid feature image;
[0273] Extracting features between the spatial grids from the basic feature image by the local feature convolutional layer to obtain a local feature image;
[0274] Performing dimensionality conversion on the grid feature image and the local feature image by the output layer according to the number of value domain partitions to obtain an information matrix, where the value domain partition is the 1D projection of the 3D grid in the pixel value domain.
[0275] In one implementation, the information matrix includes a reference information matrix corresponding to each 3D grid. The above-mentioned processing of the face image to be beautified by using the information matrix to obtain a beautified face image corresponding to the face image to be beautified includes:
[0276] Interpolating the reference information matrix based on the face image to be beautified to obtain a beautified information matrix corresponding to each pixel point of the face image to be beautified;
[0277] Processing each pixel point of the face image to be beautified respectively according to the beautified information matrix corresponding to each pixel point of the face image to be beautified to obtain a beautified face image.
[0278] In one implementation, interpolating the reference information matrix based on the face image to be beautified to obtain a beautification information matrix corresponding to each pixel of the face image to be beautified includes:
[0279] When the face image to be beautified is a multi-channel image, converting the face image to be beautified into a single-channel reference value image;
[0280] Interpolating the reference information matrix based on the reference value image to obtain a beautification information matrix corresponding to each pixel of the face image to be beautified.
[0281] The reference information matrix corresponds to the center point of the three-dimensional grid; interpolating the reference information matrix based on the face image to be beautified to obtain a beautification information matrix corresponding to each pixel of the face image to be beautified includes:
[0282] Interpolating one or more reference information matrices according to the offset of each pixel of the face image to be beautified relative to the center point of one or more three-dimensional grids to obtain a beautification information matrix corresponding to each pixel of the face image to be beautified.
[0283] In one implementation, processing each pixel of the face image to be beautified respectively according to the beautification information matrix corresponding to each pixel of the face image to be beautified to obtain a beautified face image includes:
[0284] Adding a new channel to the face image to be beautified according to the dimension of the beautification information matrix and setting the new channel to a preset value;
[0285] Multiplying the pixel value vector of each pixel of the face image to be beautified by the beautification information matrix corresponding to each pixel respectively to obtain a beautified face image; the pixel value vector of each pixel is a vector formed by the values of each channel of each pixel.
[0286] In one implementation, the image beautification processing device 1300 may further include a network training module, configured to:
[0287] Inputting the face sample image to be beautified into a deep neural network to be trained to output a sample information matrix;
[0288] Processing the face sample image to be beautified by using the sample information matrix to obtain a beautified sample image corresponding to the face sample image to be beautified;
[0289] Updating the parameters of the deep neural network based on the difference between the labeled image corresponding to the face sample image to be beautified and the beautified sample image.
[0290] In one implementation, obtaining the face image to be beautified includes:
[0291] Extract one or more original face sub-images from the original image to be beautified;
[0292] Combine the original face sub-images based on the input image size of the deep neural network to generate a face image to be beautified;
[0293] After obtaining the beautified face image, the method further includes:
[0294] Split the beautified face sub-images corresponding to the original face sub-images from the beautified face image.
[0295] In one implementation, combining the original face sub-images based on the input image size of the deep neural network to generate a face image to be beautified includes:
[0296] According to the number of original face sub-images, divide the input image size into sub-image sizes corresponding one by one to the original face sub-images;
[0297] Transform the corresponding original face sub-images respectively based on each sub-image size;
[0298] Combine the transformed original face sub-images to generate a face image to be beautified.
[0299] In one implementation, respectively transforming the corresponding original face sub-images based on each sub-image size includes any one or more of the following:
[0300] When the size relationship between the width and height of the original face sub-image is different from the size relationship between the width and height of the sub-image size, rotate the original face sub-image by 90 degrees;
[0301] When the size of the original face sub-image or the rotated original face sub-image is larger than the sub-image size, downsample the original face sub-image or the rotated original face sub-image according to the sub-image size;
[0302] When the size of the original face sub-image or the original face sub-image processed by at least one of rotation and downsampling is smaller than the sub-image size, pad the original face sub-image according to the difference between the size of the original face sub-image and the sub-image size, or pad the original face sub-image processed by at least one of rotation and downsampling according to the difference between the size of the original face sub-image processed by at least one of rotation and downsampling and the sub-image size.
[0303] In one implementation, the beauty processing module 1330 is further configured to:
[0304] Replace the original face sub-images in the original image to be beautified with the corresponding beautified face sub-images to obtain the target beautified image corresponding to the original image to be beautified.
[0305] In one implementation, the beauty processing module 1330 is further configured to:
[0306] Before replacing the original face sub-image in the original image to be beautified with the corresponding beautified face sub-image, perform beauty weakening processing on the beautified face sub-image using the original face sub-image.
[0307] In one implementation, the above-mentioned beauty weakening processing of the beautified face sub-image using the original face sub-image includes:
[0308] Fuse the original face sub-image into the beautified face sub-image according to the set beauty degree parameter.
[0309] In one implementation, the above-mentioned beauty weakening processing of the beautified face sub-image using the original face sub-image includes:
[0310] Fuse the high-frequency image of the original face sub-image into the beautified face sub-image.
[0311] In one implementation, the beauty processing module 1330 is further configured to:
[0312] Before fusing the high-frequency image of the original face sub-image into the beautified face sub-image, determine the defect points in the high-frequency image, and adjust the pixel values within a preset area around the defect points in the high-frequency image to within a preset numerical range.
[0313] In one implementation, the beauty processing module 1330 is further configured to:
[0314] When replacing the original face sub-image in the original image to be beautified with the corresponding beautified face sub-image, perform a fade processing on the boundary area between the un-replaced area in the original image to be beautified and the beautified face sub-image, so that the boundary area forms a smooth transition.
[0315] In one implementation, the beautified face image includes a blemish-removed beautified image. The beauty processing module 1330 is further configured to:
[0316] After obtaining the blemish-removed beautified image, perform personalized beauty processing on the blemish-removed beautified image to obtain the final beautified image.
[0317] The specific details of each part in the above device have been described in detail in the implementation of the method part. For the details not disclosed, please refer to the implementation content of the method part, and thus will not be elaborated here.
[0318] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium, which can be implemented in the form of a program product. The program product includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. In an alternative embodiment, the program product can be implemented as a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0319] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0320] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0321] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0322] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partly on the user's device, execute as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or execute entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0323] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0324] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here. After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0325] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only defined by the appended claims.
Claims
1. An image beauty processing method, characterized in that, it includes: Obtain a face image to be beautified; Extract three-dimensional grid-based features from the face image to be beautified through a pre-trained deep neural network and output an information matrix, where the three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and pixel value domain of the face image to be beautified; Process the face image to be beautified using the information matrix to obtain a beautified face image corresponding to the face image to be beautified; Wherein, the information matrix includes a reference information matrix corresponding to each three-dimensional grid, and the reference information matrix includes reference information for beauty processing of all pixel points within the three-dimensional grid; the process of using the information matrix to process the face image to be beautified to obtain a beautified face image corresponding to the face image to be beautified includes: Interpolate the reference information matrix based on the face image to be beautified to obtain a beauty information matrix corresponding to each pixel point of the face image to be beautified; Process each pixel point of the face image to be beautified according to the beauty information matrix corresponding to each pixel point of the face image to be beautified to obtain the beautified face image.
2. The method according to claim 1, characterized in that, The deep neural network includes a basic convolutional layer, a grid feature convolutional layer, a local feature convolutional layer, and an output layer; the process of extracting three-dimensional grid-based features from the face image to be beautified through a pre-trained deep neural network and outputting an information matrix includes: Perform downsampling convolution processing on the face image to be beautified by the basic convolutional layer according to the size of the spatial grid, where the spatial grid is the two-dimensional projection of the three-dimensional grid on the spatial domain, to obtain a basic feature image; Extract features within the spatial grid from the basic feature image through the grid feature convolutional layer to obtain a grid feature image; Extract features between the spatial grids from the basic feature image through the local feature convolutional layer to obtain a local feature image; Perform dimensional conversion on the grid feature image and the local feature image by the output layer according to the number of value domain partitions to obtain the information matrix, where the value domain partition is the one-dimensional projection of the three-dimensional grid on the pixel value domain.
3. The method according to claim 1, characterized in that, The process of interpolating the reference information matrix based on the face image to be beautified to obtain a beauty information matrix corresponding to each pixel point of the face image to be beautified includes: When the face image to be beautified is a multi-channel image, convert the face image to be beautified into a single-channel reference value image; Interpolate the reference information matrix based on the reference value image to obtain a beauty information matrix corresponding to each pixel point of the face image to be beautified.
4. The method according to claim 3, characterized in that, The following formula is used to convert the face image to be beautified into a single-channel reference value image: (1); Wherein, R, G, and B are the normalized pixel values of each pixel point in the face image to be beautified; n represents dividing the value ranges of R, G, and B into n partitions, and j represents the ordinal number of the partition; a rj , a gj , a bj are the conversion coefficients of each partition of R, G, and B respectively, which can be determined according to experience or actual requirements; shift rj , shift gj , shift bj are the conversion thresholds set in each partition of R, G, and B respectively; guidemap r , guidemap g , guidemap b are the single-channel images of R, G, and B after partition conversion respectively; g r , g g , g b are the fusion coefficients of R, G, and B respectively; guidemap bias is the offset added after fusion; guidemap z is the reference value image, and its value range is [0, 1].
5. The method according to claim 1, characterized in that, The reference information matrix corresponds to the center point of the three-dimensional grid; interpolating the reference information matrix based on the to-be-beautified face image to obtain a beautification information matrix corresponding to each pixel point of the to-be-beautified face image, including: Interpolating one or more of the reference information matrices according to the offset of each pixel point of the to-be-beautified face image relative to the center point of one or more of the three-dimensional grids, to obtain a beautification information matrix corresponding to each pixel point of the to-be-beautified face image.
6. The method according to claim 1, wherein, processing each pixel point of the to-be-beautified face image respectively according to the beautification information matrix corresponding to each pixel point of the to-be-beautified face image to obtain the beautified face image, including: Adding a new channel to the to-be-beautified face image according to the dimension of the beautification information matrix, and setting the new channel to a preset value; Multiplying the pixel value vector of each pixel point of the to-be-beautified face image by the beautification information matrix corresponding to each pixel point respectively to obtain the beautified face image; the pixel value vector of each pixel point is a vector formed by the values of each channel of each pixel point.
7. The method according to claim 1, wherein, The method further includes: Inputting a to-be-beautified sample image into the to-be-trained deep neural network to output a sample information matrix; Processing the to-be-beautified sample image by using the sample information matrix to obtain a beautified sample image corresponding to the to-be-beautified sample image; Updating the parameters of the deep neural network based on the difference between the labeled image corresponding to the to-be-beautified sample image and the beautified sample image.
8. The method according to claim 1, wherein, Obtaining the to-be-beautified face image, including: Extracting one or more original face sub-images from the to-be-beautified original image; Combining the original face sub-images based on the input image size of the deep neural network to generate the to-be-beautified face image; After obtaining the beautified face image, the method further includes: Splitting the beautified face sub-images corresponding to the original face sub-images from the beautified face image.
9. The method according to claim 8, wherein, Combining the original face sub-images based on the input image size of the deep neural network to generate the to-be-beautified face image, including: Dividing the input image size into sub-image sizes corresponding one by one to the original face sub-images according to the number of the original face sub-images; Transforming the corresponding original face sub-images respectively based on each sub-image size; Combining the transformed original face sub-images to generate the to-be-beautified face image.
10. The method according to claim 9, wherein, Respectively transforming the corresponding original face sub-images based on each sub-image size includes any one or more of the following: When the size relationship between the width and height of the original face sub-image is different from the size relationship between the width and height of the sub-image size, rotating the original face sub-image by 90 degrees; When the size of the original face sub-image or the rotated original face sub-image is greater than the size of the sub-image, downsample the original face sub-image or the rotated original face sub-image according to the size of the sub-image; When the size of the original face sub-image or the original face sub-image that has undergone at least one of rotation and downsampling is less than the size of the sub-image, pad the original face sub-image according to the difference between the size of the original face sub-image and the size of the sub-image, or pad the original face sub-image that has undergone at least one of rotation and downsampling according to the difference between the size of the original face sub-image that has undergone at least one of rotation and downsampling and the size of the sub-image.
11. The method according to claim 8, wherein, the method further includes: Replacing the original face sub-image in the original image to be beautified with the corresponding beautified face sub-image to obtain the target beautified image corresponding to the original image to be beautified.
12. The method according to claim 11, wherein, Before replacing the original face sub-image in the original image to be beautified with the corresponding beautified face sub-image, the method further includes: Performing beautification weakening processing on the beautified face sub-image using the original face sub-image.
13. The method according to claim 12, wherein, The performing beautification weakening processing on the beautified face sub-image using the original face sub-image includes: Fusing the original face sub-image into the beautified face sub-image according to the set beautification degree parameter.
14. The method according to claim 12, wherein, The performing beautification weakening processing on the beautified face sub-image using the original face sub-image includes: Fusing the high-frequency image of the original face sub-image into the beautified face sub-image.
15. The method according to claim 14, wherein, the method further includes: When combining the original face sub-images extracted from the original image to be beautified according to the input image size of the deep neural network, if the original face sub-images are downsampled, upsample the downsampled face sub-images obtained after downsampling to obtain upsampled face sub-images, and the upsampled face sub-images have the same resolution as the original face sub-images; Obtain the high-frequency image of the original face sub-image according to the difference between the original face sub-image and the upsampled face sub-image.
16. The method according to claim 15, wherein, Before fusing the high-frequency image of the original face sub-image into the beautified face sub-image, the method further includes: Determine the defect points in the high-frequency image; Adjust the pixel values within a preset area around the defect points in the high-frequency image to a preset value range.
17. The method according to claim 11, wherein, When replacing the original face sub-image in the original image to be beautified with the corresponding beautified face sub-image, the method further includes: Perform a gradient processing on the boundary region between the un-replaced region in the original image to be beautified and the beautified face sub-image, so that the boundary region forms a smooth transition.
18. The method according to claim 1, wherein, the beautified face image includes a blemish-removed beautified image, and after obtaining the blemish-removed beautified image, the method further includes: Performing personalized beautification processing on the blemish-removed beautified image to obtain a final beautified image.
19. An image beautification processing apparatus, wherein, comprising: An image acquisition module configured to acquire a face image to be beautified; An information matrix generation module configured to extract three-dimensional grid-based features from the face image to be beautified through a pre-trained deep neural network and output an information matrix, where the three-dimensional grid is obtained by dividing the three-dimensional space formed by the spatial domain and the pixel value domain of the face image to be beautified; A beautification processing module configured to process the face image to be beautified by using the information matrix to obtain a beautified face image corresponding to the face image to be beautified; wherein, the information matrix includes a reference information matrix corresponding to each of the three-dimensional grids, and the reference information matrix includes reference information for beautification processing of all pixel points within the three-dimensional grid; the processing of the face image to be beautified by using the information matrix to obtain a beautified face image corresponding to the face image to be beautified includes: Interpolating the reference information matrix based on the face image to be beautified to obtain a beautification information matrix corresponding to each pixel point of the face image to be beautified; Processing each pixel point of the face image to be beautified respectively according to the beautification information matrix corresponding to each pixel point of the face image to be beautified to obtain the beautified face image.
20. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the method according to any one of claims 1 to 18.
21. An electronic device, wherein, comprising: A processor; and A memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method according to any one of claims 1 to 18 by executing the executable instructions.
Citation Information
Patent Citations
Image enhancement method, model training method and equipment
CN113066017A
Image beautifying processing method and device, storage medium and electronic equipment
CN113077397A