Image processing method, apparatus, device, medium, and program product
By training a contour smoothing network using a generative adversarial network, the target smoothing attribute is adaptively determined, and high-dimensional and low-dimensional features of facial images are corrected. This solves the problem of unstable smoothing effect in facial contour beautification and achieves natural and economical facial contour processing.
Patent Information
- Application Number
- CN202211339290.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Existing facial contouring technologies suffer from inaccurate facial key point detection, resulting in unstable and unnatural smoothing effects that fail to meet users' requirements for smooth facial contours.
Generative adversarial networks are used for training to generate a contour smoothing network. By adaptively determining the target smoothing attribute, the high-dimensional and low-dimensional features of the original facial image are corrected to achieve smoothing of the facial contour without changing the features of other regions in the image.
It improves image processing, making facial contours more natural, reducing the user's creation cost, and achieving smooth one-click changes to facial contours.
Smart Images

Figure CN115641276B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to an image processing method and device, equipment, medium and program product. BACKGROUND
[0002] Image processing technology is widely used in the scene of portrait or pet image beautification, which usually performs face beautification based on key points, such as face contour beautification. However, the existing face contour beautification may be unstable and unnatural due to inaccurate face key point detection, and cannot meet the user's requirements for face contour smoothing in the image. SUMMARY
[0003] The present disclosure provides an image processing method and device, equipment, storage medium and program product to solve the technical problem of poor face contour smoothing effect in a face image to some extent.
[0004] In a first aspect, the present disclosure provides an image processing method, comprising:
[0005] obtaining an original face image to be processed;
[0006] processing the original face image based on a first network to obtain a high-dimensional feature, a low-dimensional feature and a target smoothing attribute of the original face image;
[0007] The first network performs contour smoothing processing on the high-dimensional feature and the low-dimensional feature based on the target smoothing attribute to obtain a high-dimensional corrected feature and a low-dimensional corrected feature;
[0008] generating a target face image based on the high-dimensional corrected feature and the low-dimensional corrected feature.
[0009] In a second aspect, the present disclosure provides an image processing device, comprising:
[0010] an acquisition module configured to obtain an original face image to be processed;
[0011] a first network configured to process the original face image to obtain a high-dimensional feature, a low-dimensional feature and a target smoothing attribute of the original face image; the first network performs contour smoothing processing on the high-dimensional feature and the low-dimensional feature based on the target smoothing attribute to obtain a high-dimensional corrected feature and a low-dimensional corrected feature; and generates a target face image based on the high-dimensional corrected feature and the low-dimensional corrected feature.
[0012] A third aspect of this disclosure provides an electronic device, characterized in that it includes one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the programs including instructions for performing the method according to the first or second aspect.
[0013] A fourth aspect of this disclosure provides a non-volatile computer-readable storage medium comprising a computer program that, when executed by one or more processors, causes the processors to perform the method described in the first or second aspect.
[0014] A fifth aspect of this disclosure provides a computer program product including computer program instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.
[0015] As can be seen from the above, the image processing method, apparatus, device, medium and program product provided by this disclosure, based on the target smoothing attribute determined by the first network adaptively on the original facial image, corrects the high-dimensional and low-dimensional features of the original facial image to achieve smoothing of the facial contour in the image without changing the features of other areas in the image, making the processed image more natural, improving the image processing effect, and reducing the user's creation cost while beautifying the image. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the image processing architecture according to an embodiment of the present disclosure.
[0018] Figure 2 This is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.
[0019] Figure 3 This is a schematic diagram illustrating the principle of an image processing method according to an embodiment of the present disclosure.
[0020] Figure 4 This is a schematic diagram illustrating the principle of an image processing method according to an embodiment of the present disclosure.
[0021] Figure 5 This is a schematic flowchart of an image processing method according to an embodiment of the present disclosure.
[0022] Figure 6 This is a schematic diagram of an image processing apparatus according to an embodiment of the present disclosure. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0024] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0027] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0029] Figure 1A schematic diagram of an image processing architecture according to an embodiment of the present disclosure is shown. (Reference) Figure 1 The image processing architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 can be connected via a wired or wireless network 130. The server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, security services, and CDN.
[0030] Terminal 120 can be implemented in hardware or software. For example, when terminal 120 is implemented in hardware, it can be various electronic devices with a display screen and support page display, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal 120 is implemented in software, it can be installed in the electronic devices listed above; it can be implemented as multiple software programs or software modules (e.g., software programs or software modules used to provide distributed services) or as a single software program or software module, without specific limitations.
[0031] It should be noted that the image processing method provided in this application embodiment can be executed by the terminal 120 or by the server 110. It should be understood that... Figure 1 The number of terminals, networks, and servers shown is for illustrative purposes only and is not intended to be a limitation. Any number of terminals, networks, and servers can be used depending on implementation needs.
[0032] Figure 2 A schematic diagram of the hardware structure of an exemplary electronic device 200 provided in an embodiment of this disclosure is shown. For example... Figure 2 As shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. The processor 202, memory 204, network module 206, and peripheral interface 208 are interconnected within the electronic device 200 via the bus 210.
[0033] Processor 202 may be a central processing unit (CPU), image processor, neural network processor (NPU), microcontroller (MCU), programmable logic device, digital signal processor (DSP), application-specific integrated circuit (ASIC), or one or more integrated circuits. Processor 202 can be used to perform functions related to the techniques described in this disclosure. In some embodiments, processor 202 may also include multiple processors integrated as a single logic component. For example, such as... Figure 2 As shown, processor 202 may include multiple processors 202a, 202b and 202c.
[0034] Memory 204 can be configured to store data (e.g., instructions, computer code, etc.). Figure 2 As shown, the data stored in memory 204 may include program instructions (e.g., program instructions for implementing the image processing method of embodiments of this disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). Processor 202 may also access the program instructions and data stored in memory 204 and execute the program instructions to operate on the data to be processed. Memory 204 may include volatile or non-volatile storage devices. In some embodiments, memory 204 may include random access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid-state drive (SSD), flash memory, memory stick, etc.
[0035] Network module 206 can be configured to provide communication with other external devices to electronic device 200 via a network. This network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It is understood that the type of network is not limited to the specific examples described above. In some embodiments, network module 306 may include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.
[0036] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to enable information input and output. For example, peripheral devices may include input devices such as keyboards, mice, touchpads, touch screens, microphones, and various sensors, as well as output devices such as displays, speakers, vibrators, and indicator lights.
[0037] Bus 210 can be configured to transfer information between various components of electronic device 200 (e.g., processor 202, memory 204, network module 206, and peripheral interface 208), such as internal buses (e.g., processor-memory bus), external buses (USB port, PCI-E bus), etc.
[0038] It should be noted that although the architecture of the above-described electronic device 200 only shows the processor 202, memory 204, network module 206, peripheral interface 208, and bus 210, in specific implementations, the architecture of the electronic device 200 may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the architecture of the above-described electronic device 200 may only include the components necessary for implementing the embodiments of this disclosure, and does not necessarily include all the components shown in the figures.
[0039] To achieve better image quality, people often use beauty-enhancing apps to process images. These apps typically enhance images based on facial key points, such as detecting and adjusting these points to smooth facial contours. However, this method can result in poor and inconsistent smoothing effects due to inaccurate facial key point detection. Furthermore, smoothing facial contours may alter facial features, skin tone, and other characteristics, leading to excessive changes in the overall smoothed image that deviate significantly from the real face and fails to meet users' requirements for smooth facial contours. Therefore, improving the smoothing effect of facial contours in facial images has become a pressing technical problem.
[0040] In view of this, embodiments of this disclosure provide an image processing method, apparatus, device, storage medium, and program product. Based on the target smoothing attribute determined adaptively by a first network on the original facial image, the high-dimensional and low-dimensional features of the original facial image are corrected to achieve smoothing of the facial contours in the image without altering the features of other areas in the image. This makes the processed image more natural, improves the image processing effect, and reduces the user's creative cost while beautifying the image. Specifically, in image processing applications, based on the target smoothing attribute adaptively determined by the first network, the smoothness of the facial contours can be changed with a single click, making uneven contours in the image smooth and fluid.
[0041] See Figure 3 , Figure 3 A schematic diagram illustrating an image processing method according to an embodiment of the present disclosure is shown. Figure 3In this model, the image processing model 300 can employ a Generative Adversarial Network (GAN) architecture, comprising a generator network 310 and a discriminator network 320. The initial GAN can be trained using first training samples and a pre-defined supervision strategy to obtain the image processing model 300. The generator network 310 in the image processing model 300 can be used as a contour smoothing network 330 in practical applications to smooth facial contours in the original image to be processed, thus obtaining the smoothed target image.
[0042] In some embodiments, the first training sample may include at least one sample pair, each sample pair including an original training image A and a facial correction image B corresponding to the original training image A. In some embodiments, the original training image A may be feature-corrected based on a preset contour smoothing standard to obtain the corresponding facial correction image B. Specifically, a first number (e.g., 3000) of original training images A may be obtained first. The original training images A include facial images that meet image quality requirements, such as a resolution of not less than 1024*1024 pixels. The original training image A in the first training sample may be a face image, and it is possible to cover as many types of facial images as possible, such as male and female faces, faces of various age groups (e.g., 20-80 years old), and faces from various angles, thereby ensuring the richness of the training data and improving the accuracy of model training. Then, according to a preset facial contour effect standard, the facial image of the original training image A is manually feature-corrected to make the facial contour of the original training image A smooth, thereby obtaining the corresponding facial correction image B. It should be understood that when manually correcting facial contours, other facial features such as skin color, facial features, and skin texture can be left unprocessed. This ensures that the trained image processing model only improves the smoothness of the facial contours without excessively altering the face, affecting the realism of the facial image, and resulting in unnatural image processing effects. Furthermore, the original training image A in the first training sample can also be an animal facial image, such as a cat or dog's facial image.
[0043] In some embodiments, the generative adversarial network can be trained based on a first training sample and a preset supervision strategy to obtain a contour smoothing network. To ensure that the generative network 310 in the image processing model 300 can adaptively learn the network parameters for facial smoothing during training, a preset supervision strategy can be set, and the image processing model 300 can be trained in a supervised manner using the first training sample, thereby obtaining a generative network 310 that can adaptively match appropriate target smoothing attributes to the input image as a contour smoothing network 330.
[0044] In some embodiments, the preset supervision strategy may include setting illumination condition parameters for the input image of the generative adversarial network to simulate the illumination conditions of the input image. Further, these illumination condition parameters may be randomly set. During the training phase, the input image to the output adversarial network may be an image from the first training sample, such as the original training image A and the corresponding facial correction image B. Simulating the illumination conditions of the input image in the image processing model 300 based on the illumination condition parameters increases the richness of the input data, allowing the image processing model 300 to process more diverse input images during training, thereby improving the accuracy of the image processing model 300. Specifically, for example, a gamma correction algorithm can be used to simulate illumination; that is, the illumination condition parameters can be set by setting the relevant parameters of the gamma correction algorithm.
[0045] In some embodiments, the preset supervision strategy may include:
[0046] The generative network generates a corresponding first image based on the original training image in the first training sample.
[0047] The cross-entropy loss functions of the generator network and the discriminator network are calculated based on the first image and the original training image, respectively.
[0048] The first discriminant network parameters of the discriminant network are adjusted based on the cross-entropy loss function to maximize the cross-entropy loss function, and the first generator network parameters of the generator network are adjusted based on the cross-entropy loss function to minimize the cross-entropy loss function.
[0049] Specifically, the cross-entropy loss function V(D, G) of the generator network 310 and the discriminator network 320 can include the sum of a first expectation function with respect to a first logarithmic function and a second expectation function with respect to a second logarithmic function. The first logarithmic function includes a logarithmic function of the first discrimination result for the original training image x, and the second logarithmic function includes a logarithmic function of the difference between a first preset value (e.g., 1) and a second discrimination result generated for the original training image x. The training process of the image processing model 300 can be performed by training the generator network 310 and the discriminator network 320 separately and alternately. For example, the generator network 310 can be fixed first, and the discriminator network 320 can be trained to update the first discriminator network parameters of the discriminator network 320. At this time, the first discriminator network parameters of the discriminator network 320 can be adjusted so that the discriminator network 320 outputs 1 (i.e., D(B) = 1) when its input is a facial correction image B, and outputs 0 (i.e., D(A') = 0) when its input is a first image A'. Therefore, the training objective of image processing model 300 is to maximize the cross-entropy function V(D, G).
[0050] Then, the discriminator network 320 is fixed, and the generator network 310 is trained to update the first generator network parameters of the generator network 310. At this point, the first generator network parameters of the generator network 310 can be adjusted so that when the first image A' output by the generator network 310 is used as the input to the discriminator network 320, the discriminator network 320 outputs 1 (i.e., D(A') = 1). Therefore, the training objective of the image processing model 300 is to minimize the cross-entropy function V(D, G). Since the first network parameters of the discriminator network 320 remain unchanged, then E... x [logD(x)] remains unchanged, minimizing the cross-entropy function V(D, G).
[0051] This process of training the discriminator network 320 while keeping the generator network 310 fixed, and training the generator network 310 while keeping the discriminator network 320 fixed, is repeated until a Nash equilibrium is reached. This makes the results generated by the generator network 310 more realistic.
[0052] In some embodiments, the image processing model further includes a smoothing attribute discriminator associated with the generative network, and the preset supervision strategy may include:
[0053] The generative network generates a corresponding first image based on the original training image in the first training sample.
[0054] Based on the original training image, the first smoothing attribute of the original training image, the face correction image, the second smoothing attribute of the face correction image, and the first image, calculate the smoothing attribute loss function of the generator network and the smoothing attribute discriminator;
[0055] The parameters of the second generator network of the generator network are adjusted based on the smoothing attribute loss function so that the first image matches the second smoothing attribute.
[0056] Specifically, such as Figure 4 As shown, Figure 4 A schematic diagram of an image processing model according to an embodiment of the present disclosure is shown. Figure 4In the image processing model 300, a smoothing attribute discriminator 340 associated with the generator network 310 can also be set in the image processing model 300 to calculate the smoothing attribute loss function of the generator network 310 and the smoothing attribute discriminator 340. The original training image A and its first smoothing attribute R1, the face correction image B and its second smoothing attribute R2, and the first image A' from the first training samples can be used as paired input data to the smoothing attribute discriminator 340. This input data can be represented as an object-attribute pair (Image, attr), where Image represents the input image and attr represents the smoothing attribute. The smoothing attribute discriminator 340 performs a matching judgment on the input data (Image, attr). If Image and attr match, the judgment result is True; otherwise, it is False. For example, if imgae is the original training image A and attr is the value R1 corresponding to the original training image A, the output is True; if imgae is the original training image A and attr is the value R2 corresponding to the facial correction image B, the output is False. During training, the generator network 310 and the smooth attribute discriminator 340 are trained and updated separately and alternately.
[0057] When updating the smooth attribute discriminator 340, the generator network 310 is fixed. The parameters of the second discriminator network of the smooth attribute discriminator 340 are adjusted so that when the input of the smooth attribute discriminator 340 is (A, R1) or (B, R2), the corresponding output is Dattr(A, R1) = Dattr(B, R2) = 1 (i.e., True); when the input is (B, R1), (A, R2) or (A', R2), the corresponding output is Dattr(B, R1) = Dattr(A, R2) = Dattr(A', R2) = 0 (i.e., False). The cross-entropy function of the generator network 310 and the smooth attribute discriminator 340 can be used as the smooth attribute loss function V(DattrG), which includes the sum of a third expectation function with respect to a third logarithmic function and a fourth expectation function with respect to a fourth logarithmic function. The third logarithmic function includes the logarithmic function of the discrimination result of the original training image x and its attribute sttr, and the fourth logarithmic function includes the logarithmic function of the difference between a first preset value (e.g., 1) and the generated result of the original training image x and its attribute sttr. Therefore, the objective of updating the smooth attribute discriminator 340 is to maximize the smooth attribute loss function V(Dattr, G).
[0058] When updating the generator network 310, the smoothing attribute discriminator 340 is fixed. The second generator network parameters of the generator network 310 are adjusted so that when the input of the smoothing attribute discriminator 340 is (A', R2), the corresponding output is Dattr(A', R1) = Dattr(G(A), R1) = 1 (i.e., True). This ensures that the first image A' matches the second smoothing attribute of the facial correction image B, ultimately making the first image A' generated by the generator network 310 conform to the smoothing attribute of the facial correction image B. This gives the generator network 310 the characteristic of adaptively matching the smoothing attribute of the input image. In this way, the generator network 310 can adaptively match the smoothing attribute suitable for its input image to generate a contour smoothing image suitable for the input image, thereby improving the contour smoothing effect when the generator network 310 is used as a contour smoothing network.
[0059] In some embodiments, the preset supervision strategy may include:
[0060] The generative network generates a corresponding first image based on the original training image in the first training sample.
[0061] Calculate the feature correction loss function based on the first image and the facial correction image in the first training sample;
[0062] Based on the feature correction loss function, the parameters of the third generator network of the generator network and the parameters of the third discriminator network of the discriminator network are adjusted to minimize the feature correction loss function.
[0063] Furthermore, in some embodiments, calculating a feature correction loss function based on the first image and the facial correction image in the first training sample further includes:
[0064] Feature extraction is performed on the first image to obtain a first high-dimensional semantic feature and a first low-dimensional texture feature; and feature extraction is performed on the facial correction image in the first training sample to obtain a second high-dimensional semantic feature and a second low-dimensional texture feature.
[0065] A first high-dimensional feature loss function is calculated based on the first high-dimensional semantic features and the second high-dimensional semantic features, and a first low-dimensional feature loss function is calculated based on the first low-dimensional texture features and the second low-dimensional texture features;
[0066] The feature correction loss function is obtained by summing the first high-dimensional feature loss function and the second low-dimensional feature loss function.
[0067] High-dimensional semantic features refer to features obtained from deep networks in image processing models. These features are located near the output layer and are characterized by low resolution, small feature map size, high abstraction, and the inclusion of more global information. Low-dimensional texture features refer to features obtained from shallow networks in image processing models. These features are located near the input layer and are characterized by high resolution, large feature map size, the inclusion of more detailed information, and ease of alignment with the original training images. By correcting these two types of features, their advantages can be combined, thereby improving the training effect of the image processing model and the effect of facial contour smoothing.
[0068] Specifically, the generative network 310 generates a corresponding first image A' based on the original training image A. The visual processor in the image processing model 300 can extract features from the first image A' to obtain the first high-dimensional semantic feature F1_A' and the first low-dimensional texture feature F2_A'; and extract features from the facial correction image in the first training sample to obtain the second high-dimensional semantic feature F3 and the second low-dimensional texture feature F4. Then, the feature correction loss function L_F can be calculated as: first high-dimensional feature loss function l1(F1_A', F3) + first low-dimensional feature loss function l1(F2_A', F4), where l1 is the mean absolute error function.
[0069] It should be understood that the first generator network parameters, the second generator network parameters, and the third generator network parameters in this disclosure can all represent the model parameters of the generator network, and they can be the same or different; the first discriminant network parameters, the second discriminant network parameters, and the third discriminant network parameters can all represent the model parameters of the discriminant network, and they can be the same or different.
[0070] In some embodiments, the preset supervision strategy may include:
[0071] The loss weights of the original training image are determined based on the first smoothing property of the original training image and the second smoothing property of the face correction image.
[0072] Further, the loss weights of the original training image are determined based on a first smoothing attribute of the original training image and a second smoothing attribute of the face correction image, including:
[0073] The first smoothing attribute of the original training image and the second smoothing attribute of the face correction image are calculated based on the smoothing attribute algorithm.
[0074] The degree of smoothing change of the original training image is calculated based on the first smoothing attribute and the second smoothing attribute;
[0075] The loss weight of the original training image is determined based on the degree of smoothing change, wherein the loss weight of the original training image is proportional to the degree of smoothing change.
[0076] In some embodiments, calculating the degree of smoothing change of the original training image based on the first smoothing attribute and the second smoothing attribute may include:
[0077] Calculate the attribute difference between the first smoothing attribute and the second smoothing attribute;
[0078] The degree of smooth change is obtained based on the absolute value function of the attribute difference.
[0079] Specifically, the first smoothing attribute S of the original training image A can be calculated separately. A The second smoothing property S of the facial correction image B B Then the degree of smooth change of the original training image A = abs(S B -S A ), where abs is the absolute value function. Since the degree of smoothness change reflects the extent to which facial contour changes are needed for each training sample, the loss function weight should be larger for samples with severely uneven facial contours. By allocating loss weights, the model can increase its attention to samples with greater smoothness changes, while reducing overcorrection of samples with smaller smoothness changes, thus ensuring the smoothing effect of facial contours. Therefore, it can be determined that the loss weight of the original training image A is proportional to the degree of smoothness change.
[0080] After employing one or more of the aforementioned pre-defined supervision strategies, the generative adversarial network (GAN) is trained to obtain a trained image processing model. This trained model adaptively learns the network parameters for facial smoothing from training samples during training, enabling it to adaptively match smoothing properties suitable for the input image. The generative network is then used as a contour smoothing network for facial contour processing in practical applications, processing the facial contours of the user-input original image to make them smooth and fluid.
[0081] In some embodiments, the generative adversarial network is trained based on a first training sample and a preset supervision policy to obtain the contour smoothing network, and may further include:
[0082] A preliminary image processing model is obtained by training the generative adversarial network based on the first training sample and the preset supervision strategy.
[0083] The contour smoothing network is obtained by performing secondary training on the preliminary image processing model based on the second training samples and the preset supervision strategy.
[0084] Furthermore, in some embodiments, the second training samples can be obtained based on the preliminary image processing model, specifically including:
[0085] A training dataset comprising multiple facial images is obtained, and the facial images are input into the generative network in the generative adversarial network to obtain a second image;
[0086] The second image is input into the preliminary image processing model to obtain a third image corresponding to the second image after preliminary smoothing;
[0087] The second training sample is obtained based on the second image and the corresponding third image.
[0088] Since the amount of data in the first training sample is relatively small, the preliminary image processing model (including the preliminary generator network and the preliminary discriminator network) trained based on the first training sample is not very stable. In order to increase the stability of the generator image processing model, it is necessary to use the preliminary image processing model to process a large amount of data to obtain a large amount of secondary training dataset as the second training sample. Then, the preliminary image processing model is trained again using the second training sample and a preset supervision strategy to obtain a more stable image processing model. The generator network in this more stable image processing model is used as the contour smoothing network, thereby improving the stability of facial contour processing in practical applications.
[0089] Specifically, a large number of facial images (e.g., an open-source dataset of facial images) can be acquired and input into the initial generative network of an initial generative adversarial network to generate a second image D, resulting in a large dataset set_D. This large dataset set_D is then input into a preliminary image processing model trained based on the first training samples to obtain the output third image E. The second image D and the corresponding third image E form a new training data pair, serving as the second training samples. This pair is then used to perform a second training on the preliminary image processing model using a pre-defined supervision strategy, resulting in a new image processing model. The generative network in this new image processing model can be used as a contour smoothing network for practical applications (e.g.,...). Figure 4 Contour smoothing network 330 in the middle.
[0090] See Figure 5 , Figure 5 A schematic flowchart of an image processing method according to an embodiment of the present disclosure is shown. Figure 5 In the image processing method 500, the following steps may be included.
[0091] Step S510: Obtain the original facial image to be processed;
[0092] Step S520: Process the original facial image based on the first network to obtain the high-dimensional features, low-dimensional features and target smoothing properties of the original facial image;
[0093] Step S530: The first network performs contour smoothing processing on the high-dimensional features and the low-dimensional features based on the target smoothing attribute to obtain high-dimensional corrected features and low-dimensional corrected features;
[0094] Step S540: Generate a target facial image based on the high-dimensional correction features and the low-dimensional correction features.
[0095] Specifically, for the original image imageA to be processed, the user wants to smooth the facial contours in the original image imageA. Facial keypoint detection can be performed on the original image imageA to obtain facial keypoints P. Then, based on the facial keypoints P, facial cropping is performed on the original image imageA to obtain the original face image image_Face. The original face image image_Face is then input into a pre-trained first network (e.g., ...). Figure 3 In the contour smoothing network 330, the first network extracts features from the original facial image image_Face to obtain high-dimensional semantic features F1 and low-dimensional texture features F2. Since the trained first network can adaptively determine the target smoothing attribute matching the input data based on the input data, it can smooth the high-dimensional semantic features F1 and low-dimensional texture features F2 respectively to obtain high-dimensional semantic correction features F1' and low-dimensional texture correction features F2'. The contour smoothing network then generates the smoothed target facial image image_Face' from the high-dimensional semantic correction features F1' and low-dimensional texture correction features F2'. Furthermore, the target facial image image_Face' can be fused into the original image imageA to obtain the target image imageA' with smoothed facial contours. According to the image processing method of this disclosure, the high-dimensional semantic features and low-dimensional texture features of an image are corrected based on the adaptive smoothing properties of a contour smoothing network, so as to achieve smoothing of facial contours in the image without changing the features of other areas in the image. This makes the processed image more natural, improves the image processing effect, and reduces the user's creative cost while beautifying the image. Specifically, in image processing applications, the smoothness of facial contours can be changed with one click based on the adaptive smoothing properties of the contour smoothing network, making uneven contours in the image smooth and fluid.
[0096] In practical applications, users may need to beautify not only human images but also animal images (such as pets like cats and dogs). The method according to embodiments of this disclosure can smooth facial contours not only in human images but also in animal images, resulting in smoother and more natural facial contours in both human and animal images.
[0097] In some embodiments, the original facial image is processed based on a first network to obtain high-dimensional and low-dimensional features of the original facial image, including:
[0098] The first network extracts features from the original facial image to obtain the high-dimensional features (e.g., the high-dimensional semantic features F1 of the original facial image image_Face) and the low-dimensional features (e.g., the low-dimensional texture features F2 of the original facial image image_Face); wherein the high-dimensional features are semantic features and the low-dimensional features are texture features.
[0099] Specifically, high-dimensional features can refer to high-dimensional semantic features, and low-dimensional features can refer to low-dimensional texture features. The first network can extract features from the original face image image_Face to obtain the high-dimensional semantic features F1 and the low-dimensional texture features F2 of the original face image image_Face.
[0100] In some embodiments, the original facial image is processed based on a first network to obtain a target smoothing property of the original facial image, including:
[0101] The first network performs facial contour detection on the original facial image to obtain facial contour feature points of the original facial image, and determines the target smoothing attribute based on the facial contour feature points.
[0102] In some embodiments, determining the target smoothing property based on the facial contour feature points includes:
[0103] Target contour feature points are determined based on the facial contour feature points, and the target smoothing attribute is determined based on the target contour feature points; or,
[0104] The original smoothing attribute of the original facial image is determined based on the facial contour feature points, and the target smoothing attribute is determined based on the original smoothing attribute.
[0105] Specifically, the trained first network can adaptively determine the target smoothing attribute of the input data based on the input data. For example, it can first determine the target contour feature points of the input data and then obtain the corresponding target smoothing attribute based on the target contour feature points. Alternatively, it can first determine the original smoothing attribute from the original contour feature points of the input data and then determine the target smoothing attribute based on the original smoothing attribute.
[0106] In some embodiments, based on the target smoothing attribute, contour smoothing processing is performed on the high-dimensional features and the low-dimensional features to obtain high-dimensional corrected features and low-dimensional corrected features, including:
[0107] Based on the target smoothness attribute, the loss function weights for the high-dimensional features and the low-dimensional features are set to obtain the high-dimensional corrected features and the low-dimensional corrected features. In some embodiments, the generative adversarial network is trained based on a first training sample and a preset supervision policy to obtain the first network (e.g., Figure 3 Contour smoothing network 330);
[0108] The first training sample includes at least one sample pair, wherein the sample pair includes the original training image (e.g. Figure 3 The original training image A) and the corresponding facial correction image (e.g. Figure 3 The facial correction image (B) is obtained by smoothing the facial contours based on the original training image.
[0109] In some embodiments, the generative adversarial network includes a generative network (e.g., Figure 3 The generator network 310) and the discriminator network associated with the generator network (e.g., the discriminator network 310) Figure 3 The discriminant network 320 in the middle), the preset supervision strategy includes:
[0110] The generative network generates a corresponding first image (e.g., first image A') based on the original training images in the first training samples;
[0111] The cross-entropy loss function (e.g., V(D, G)) of the generator network and the discriminator network is calculated based on the first image and the original training image, respectively.
[0112] Based on the cross-entropy loss function, the parameters of the first discriminant network of the discriminant network are adjusted to maximize the cross-entropy loss function (e.g., ...). ), and adjust the first generator network parameters of the generator network based on the cross-entropy loss function to minimize the cross-entropy loss function (e.g. ).
[0113] In some embodiments, the generative adversarial network includes a generative network and a smoothing attribute discriminator associated with the generative network (e.g., Figure 4 The smooth attribute discriminator 340 in the middle), the preset supervision strategy includes:
[0114] The generative network generates a corresponding first image based on the original training image in the first training sample.
[0115] Based on the original training image, the first smoothing attribute of the original training image (e.g., the first smoothing attribute R1), the face correction image, the second smoothing attribute of the face correction image (e.g., the second smoothing attribute R2), and the first image, calculate the smoothing attribute loss function (e.g., V(Dattr, G)) of the generator network and the smoothing attribute discriminator;
[0116] The parameters of the second generator network are adjusted based on the smoothing attribute loss function so that the first image matches the second smoothing attribute (e.g., Datr(A', R1) = Datr(G(A), R1) = 1).
[0117] In some embodiments, the generative adversarial network includes a generative network and a smoothing attribute discriminator associated with the generative network, and the preset supervision strategy includes:
[0118] The generative network generates a corresponding first image based on the original training image in the first training sample.
[0119] Calculate a feature correction loss function (e.g., feature correction loss function L_F) based on the first image and the facial correction image in the first training sample;
[0120] The third generator network parameters of the generator network and the third discriminator network parameters of the discriminator network are adjusted based on the feature correction loss function to minimize the feature correction loss function.
[0121] In some embodiments, calculating a feature correction loss function based on the first image and the facial correction image in the first training sample further includes:
[0122] Feature extraction is performed on the first image to obtain a first high-dimensional semantic feature and a first low-dimensional texture feature; and feature extraction is performed on the facial correction image in the first training sample to obtain a second high-dimensional semantic feature and a second low-dimensional texture feature.
[0123] A first high-dimensional feature loss function is calculated based on the first high-dimensional semantic features and the second high-dimensional semantic features, and a first low-dimensional feature loss function is calculated based on the first low-dimensional texture features and the second low-dimensional texture features;
[0124] The feature correction loss function is obtained by summing the first high-dimensional feature loss function and the second low-dimensional feature loss function.
[0125] In some embodiments, the preset supervision strategy includes:
[0126] The loss weights of the original training image are determined based on the first smoothing property of the original training image and the second smoothing property of the face correction image.
[0127] In some embodiments, determining the loss weights of the original training image based on a first smoothing attribute of the original training image and a second smoothing attribute of the face correction image includes:
[0128] The first smoothing attribute of the original training image and the second smoothing attribute of the face correction image are calculated based on the smoothing attribute algorithm.
[0129] Based on the first smoothing property (e.g., the first smoothing property S) A ) and the second smoothing property (e.g., the second smoothing property S) B ) Calculate the degree of smoothness change of the original training image (e.g., smoothness change degree = abs(S) B -S A ));
[0130] The loss weight of the original training image is determined based on the degree of smoothing change, wherein the loss weight of the original training image is proportional to the degree of smoothing change.
[0131] In some embodiments, calculating the degree of smoothing change of the original training image based on the first smoothing attribute and the second smoothing attribute includes:
[0132] Calculate the attribute difference between the first smoothing attribute and the second smoothing attribute;
[0133] The degree of smooth change is obtained based on the absolute value function of the attribute difference.
[0134] In some embodiments, the generative adversarial network is trained based on a first training sample and a preset supervision policy to obtain the contour smoothing network, and further includes:
[0135] A preliminary image processing model is obtained by training the generative adversarial network based on the first training sample and the preset supervision strategy.
[0136] The contour smoothing network is obtained by performing secondary training on the preliminary image processing model based on the second training samples and the preset supervision strategy (e.g., ...). Figures 3-4 Contour smoothing network 330);
[0137] The second training sample is obtained based on the preliminary image processing model, and specifically includes:
[0138] A training dataset comprising multiple facial images is obtained, and the facial images are input into the generative adversarial network to obtain a second image (e.g., second image D);
[0139] The second image is input into the preliminary image processing model to obtain a third image (e.g., the third image E) that has undergone preliminary smoothing processing and corresponds to the second image;
[0140] The second training sample is obtained based on the second image and the corresponding third image.
[0141] In some embodiments, the preset supervision strategy includes: setting illumination condition parameters for the input image of the generative adversarial network to simulate the illumination conditions of the input image.
[0142] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0143] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0144] Based on the same technical concept, corresponding to any of the above embodiments, this disclosure also provides an image processing apparatus, see [link to relevant documentation]. Figure 6 The image processing apparatus includes:
[0145] The acquisition module is used to acquire the raw facial image to be processed;
[0146] A first network is used to process the original facial image to obtain high-dimensional features, low-dimensional features, and target smoothing attributes of the original facial image; the first network performs contour smoothing processing on the high-dimensional features and the low-dimensional features based on the target smoothing attributes to obtain high-dimensional correction features and low-dimensional correction features; and generates a target facial image based on the high-dimensional correction features and the low-dimensional correction features.
[0147] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0148] The apparatus of the above embodiments is used to implement the corresponding image processing method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0149] Based on the same technical concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the image processing method as described in any of the above embodiments.
[0150] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0151] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the image processing method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0152] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0153] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0154] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0155] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An image processing method, comprising: Obtain the raw facial image to be processed; The original facial image is processed based on a first network to obtain high-dimensional features, low-dimensional features, and target smoothing attributes of the original facial image; the first network includes a generator network and a smoothing attribute discriminator associated with the generator network; Based on the target smoothing attribute, the first network performs contour smoothing processing on the high-dimensional features and the low-dimensional features to obtain high-dimensional corrected features and low-dimensional corrected features; A target facial image is generated based on the high-dimensional correction features and the low-dimensional correction features; The generating network generates a corresponding first image based on the original training images in the first training samples; Based on the original training image, the first smoothing attribute of the original training image, the facial correction image, the second smoothing attribute of the facial correction image, and the first image, calculate the smoothing attribute loss function of the generator network and the smoothing attribute discriminator; Based on the smoothing attribute loss function, the first generator network parameters of the generator network and the first discriminator network parameters of the smoothing attribute discriminator are adjusted so that the first image matches the second smoothing attribute. The first training sample includes at least one sample pair, which includes an original training image and a corresponding facial correction image. The facial correction image is obtained by performing facial contour smoothing processing on the original training image.
2. The method according to claim 1, wherein, The original facial image is processed based on the first network to obtain high-dimensional and low-dimensional features of the original facial image, including: The first network extracts features from the original facial image to obtain the high-dimensional features and the low-dimensional features of the original facial image; wherein, the high-dimensional features are semantic features and the low-dimensional features are texture features.
3. The method according to claim 1, wherein, The original facial image is processed based on the first network to obtain the target smoothing properties of the original facial image, including: The first network performs facial contour detection on the original facial image to obtain facial contour feature points of the original facial image, and determines the target smoothing attribute based on the facial contour feature points.
4. The method according to claim 3, wherein, Determining the target smoothing attribute based on the facial contour feature points includes: Target contour feature points are determined based on the facial contour feature points, and the target smoothing attribute is determined based on the target contour feature points; or, The original smoothing attribute of the original facial image is determined based on the facial contour feature points, and the target smoothing attribute is determined based on the original smoothing attribute.
5. The method according to claim 1, wherein, Based on the target smoothing attribute, contour smoothing processing is performed on the high-dimensional features and the low-dimensional features to obtain high-dimensional corrected features and low-dimensional corrected features, including: Based on the target smoothing attribute, the loss function weights of the high-dimensional feature and the low-dimensional feature are set to obtain the high-dimensional corrected feature and the low-dimensional corrected feature.
6. The method according to claim 1, wherein, The first network also includes a discriminative network associated with the generating network; The generative network generates a corresponding first image based on the original training image in the first training sample. Feature extraction is performed on the first image to obtain a first high-dimensional semantic feature and a first low-dimensional texture feature; And perform feature extraction on the facial correction image in the first training sample to obtain the second high-dimensional semantic features and the second low-dimensional texture features; A first high-dimensional feature loss function is calculated based on the first high-dimensional semantic features and the second high-dimensional semantic features, and a first low-dimensional feature loss function is calculated based on the first low-dimensional texture features and the second low-dimensional texture features; The feature correction loss function is obtained by summing the first high-dimensional feature loss function and the first low-dimensional feature loss function; The second generator network parameters of the generator network and the second discriminator network parameters of the discriminator network are adjusted based on the feature correction loss function to minimize the feature correction loss function.
7. An image processing apparatus, comprising: The acquisition module is used to acquire the raw facial image to be processed; A first network is used to process the original facial image to obtain high-dimensional features, low-dimensional features, and target smoothing attributes of the original facial image; the first network includes a generator network and a smoothing attribute discriminator associated with the generator network; The first network performs contour smoothing processing on the high-dimensional features and the low-dimensional features based on the target smoothing attribute to obtain high-dimensional corrected features and low-dimensional corrected features; and generates a target facial image based on the high-dimensional corrected features and low-dimensional corrected features; The generating network generates a corresponding first image based on the original training images in the first training samples; Based on the original training image, the first smoothing attribute of the original training image, the facial correction image, the second smoothing attribute of the facial correction image, and the first image, calculate the smoothing attribute loss function of the generator network and the smoothing attribute discriminator; Based on the smoothing attribute loss function, the first generator network parameters of the generator network and the first discriminator network parameters of the smoothing attribute discriminator are adjusted so that the first image matches the second smoothing attribute. The first training sample includes at least one sample pair, which includes an original training image and a corresponding facial correction image. The facial correction image is obtained by performing facial contour smoothing processing on the original training image.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1 to 6.
10. A computer program product comprising computer program instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Texture enhancement method and device based on texture image, equipment and storage medium
CN111445410A