Image processing method, device, apparatus and storage medium

CN116703700BActive Publication Date: 2026-09-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210173342.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2026-09-29
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

现有方案,当原图和效果图之间的形变差异太大时,包括脸部边缘轮廓和五官大小、位置,最终结果图的脸部边缘和五官位置就会出现虚影,这是因为效果的形变程度太大,传统网络学习不到这么大的形变

Benefits of technology

[0067]本公开实施例公开了一种图像处理方法、装置、设备及存储介质。将原始图像输入生成对抗网络的生成器中,获得中间图像和第一像素变换信息;根据第一像素变换信息对中间图像进行像素变换,获得目标图像。本公开实施例提供的图像处理方法,利用生成对抗网络输出的第一像素变换信息对中间图像进行像素变换,获得目标图像,实现了对图像的大幅度形变处理,可以克服形变带来的虚影问题,从而提高图像形变的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116703700B_ABST
    Figure CN116703700B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an image processing method, device and equipment and a storage medium. An original image is input into a generator of a generative adversarial network to obtain an intermediate image and first pixel transformation information; and the intermediate image is subjected to pixel transformation according to the first pixel transformation information to obtain a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Technology

[0002] Current image applications (APPs) offer many special effects based on image algorithms. Some of these effects alter facial features and shape, such as slimming the face, making the user appear younger, or increasing weight. However, with existing solutions, when the distortion difference between the original image and the resulting image is too large—including the facial contours and the size and position of facial features—the final image will exhibit ghosting around the edges and features. This is because the distortion is too great for traditional networks to learn. Summary of the Invention

[0003] This disclosure provides an image processing method, apparatus, device, and storage medium to achieve large-scale deformation processing of facial images, which can overcome the ghosting problem caused by large deformation and thus improve the effect of facial image deformation.

[0004] In a first aspect, embodiments of this disclosure provide an image processing method, including:

[0005] The original image is input into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information;

[0006] The intermediate image is pixel-transformed based on the first pixel transformation information to obtain the target image.

[0007] Furthermore, the first pixel transformation information includes optical flow transformation information, affine transformation information, and / or perspective transformation information.

[0008] Furthermore, the optical flow transformation information is represented by an optical flow transformation matrix, where each element of the optical flow transformation matrix represents the positional offset between the pixel corresponding to that element in the intermediate image and the pixel corresponding to that element in the target image.

[0009] The step of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes:

[0010] The elements of the optical flow transformation matrix are traversed, and the target position information of the pixel is determined based on the position offset of the traversed element and the current position information of the pixel corresponding to the element in the intermediate image.

[0011] Obtain the current pixel value in the intermediate image that corresponds to the current position information and the target pixel value that corresponds to the target position information;

[0012] The current pixel value is replaced with the target pixel value to obtain the target image.

[0013] Furthermore, the affine transformation information is a matrix with a first predetermined size.

[0014] The process of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes:

[0015] For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and,

[0016] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0017] Furthermore, the perspective transformation information is a matrix with a second predetermined size.

[0018] The process of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes:

[0019] For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and,

[0020] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0021] Furthermore, the generative adversarial network also includes a discriminator; the training method of the generative adversarial neural network is as follows:

[0022] Obtain the original image samples and the corresponding result image samples;

[0023] The original image sample is input into the generator to obtain intermediate image samples and second pixel transformation information;

[0024] The intermediate image sample is pixel-transformed according to the second pixel transformation information to obtain the generated image;

[0025] The generator and the discriminator are trained iteratively and alternately based on the generated image, the original image sample, and the result image sample.

[0026] Further, the generator and the discriminator are trained iteratively and alternately based on the generated image, the original image samples, and the resulting image samples, including:

[0027] The generated image and the original image sample are combined into a negative sample pair, and the resulting image sample and the original image sample are combined into a positive sample pair;

[0028] The positive sample pairs are input into the discriminator to obtain a first discrimination result; the negative sample pairs are input into the discriminator to obtain a second discrimination result.

[0029] The first loss function is determined based on the first discrimination result and the second discrimination result;

[0030] A second loss function is determined based on the generated image and the resulting image samples;

[0031] The first loss function and the second loss function are linearly superimposed to obtain the target loss function; and

[0032] The generator and the discriminator are trained iteratively and alternately based on the target loss function.

[0033] Secondly, embodiments of this disclosure also provide an image processing apparatus, comprising:

[0034] The first pixel transformation information acquisition module is used to input the original image into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information;

[0035] The pixel transformation module is used to perform pixel transformation on the intermediate image according to the first pixel transformation information to obtain the target image.

[0036] Furthermore, the first pixel transformation information includes optical flow transformation information, affine transformation information, and / or perspective transformation information.

[0037] Furthermore, the optical flow transformation information is represented by an optical flow transformation matrix, where each element of the optical flow transformation matrix represents the positional offset between the pixel corresponding to that element in the intermediate image and the pixel corresponding to that element in the target image.

[0038] The pixel transformation module is further configured to:

[0039] The elements of the optical flow transformation matrix are traversed, and the target position information of the pixel is determined based on the position offset of the traversed element and the current position information of the pixel corresponding to the element in the intermediate image.

[0040] Obtain the current pixel value in the intermediate image that corresponds to the current position information and the target pixel value that corresponds to the target position information;

[0041] The current pixel value is replaced with the target pixel value to obtain the target image.

[0042] Furthermore, the affine transformation information is a matrix with a first predetermined size.

[0043] The pixel transformation module is further configured to:

[0044] For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and,

[0045] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0046] Furthermore, the perspective transformation information is a matrix with a second predetermined size.

[0047] The pixel transformation module is further configured to:

[0048] For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and,

[0049] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0050] Furthermore, the generative adversarial network further includes a discriminator; and also includes a generative adversarial neural network training module, used for:

[0051] Obtain the original image samples and the corresponding result image samples;

[0052] The original image sample is input into the generator to obtain intermediate image samples and second pixel transformation information;

[0053] The intermediate image sample is pixel-transformed according to the second pixel transformation information to obtain the generated image;

[0054] The generator and the discriminator are trained iteratively and alternately based on the generated image, the original image sample, and the result image sample.

[0055] Furthermore, the generative adversarial neural network training module is also used for:

[0056] The generated image and the original image sample are combined into a negative sample pair, and the resulting image sample and the original image sample are combined into a positive sample pair;

[0057] The positive sample pairs are input into the discriminator to obtain a first discrimination result; the negative sample pairs are input into the discriminator to obtain a second discrimination result.

[0058] The first loss function is determined based on the first discrimination result and the second discrimination result;

[0059] A second loss function is determined based on the generated image and the resulting image samples;

[0060] The first loss function and the second loss function are linearly superimposed to obtain the target loss function; and

[0061] The generator and the discriminator are trained iteratively and alternately based on the target loss function.

[0062] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0063] One or more processing devices;

[0064] Storage device for storing one or more programs;

[0065] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image processing method as described in the embodiments of this disclosure.

[0066] Fourthly, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the image processing method as described in embodiments of this disclosure.

[0067] This disclosure provides an image processing method, apparatus, device, and storage medium. An original image is input into a generator of a generative adversarial network (GAN) to obtain an intermediate image and first pixel transformation information. The intermediate image is then subjected to pixel transformation based on the first pixel transformation information to obtain a target image. The image processing method provided in this disclosure utilizes the first pixel transformation information output by the GAN to perform pixel transformation on the intermediate image to obtain the target image, achieving significant image deformation processing. This overcomes the ghosting problem caused by deformation, thereby improving the image deformation effect. Attached Figure Description

[0068] Figure 1 This is a flowchart of an image processing method according to an embodiment of the present disclosure;

[0069] Figure 2 This is an example diagram of optical flow transformation of an intermediate image in an embodiment of this disclosure;

[0070] Figure 3 This is an example diagram of training a generative adversarial neural network in an embodiment of this disclosure;

[0071] Figure 4 This is an example diagram of the network structure of a generator in one embodiment of this disclosure;

[0072] Figure 5 This is a schematic diagram of the structure of an image processing apparatus according to an embodiment of the present disclosure;

[0073] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation

[0074] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0075] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0076] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0077] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0078] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0079] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0080] Figure 1 This is a flowchart illustrating an image processing method provided in an embodiment of this disclosure. This embodiment is applicable to situations involving deformation processing of facial images. The method can be executed by an image processing device, which can consist of hardware and / or software and is generally integrated into a device with image processing capabilities. This device can be an electronic device such as a server, mobile terminal, or server cluster. Figure 1 As shown, the method specifically includes the following steps:

[0081] S110: Input the original image into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information.

[0082] The original image can be understood as an image containing a human face that needs to be deformed. It can be captured by the user through a mobile terminal's camera, or obtained from a local database or a server-side database. The generative adversarial network (GAN) can be a trained pixel-to-pixel (pixel-to-pixel) GAN, and the generator's output is multi-channel data. In this embodiment, the generator's output includes image data and pixel transformation information. The image data is three-channel data, and the pixel transformation information is one-channel or two-channel data. The number of output channels can be adjusted according to actual needs.

[0083] The first pixel transformation information can be optical flow transformation information, affine transformation information, and / or perspective transformation information. If it is optical flow transformation information, the first pixel transformation information is 2-channel data, with each channel represented by a matrix of image size. These two channels represent the pixel's position information (X, Y). If it is affine transformation information, the first pixel transformation information is 1-channel data, which is a vector containing six elements. If it is perspective transformation information, the first pixel transformation information is 1-channel data, which is a 3x3 matrix or a vector containing nine elements. In this embodiment, the first pixel transformation information can be different transformation information, enabling different types of deformation processing for facial images, thereby improving the diversity of deformation.

[0084] S120, Perform pixel transformation on the intermediate image according to the first pixel transformation information to obtain the target image.

[0085] The first pixel transformation information includes optical flow transformation information, affine transformation information, or perspective transformation information. The pixel transformation method is different for different transformation information.

[0086] Specifically, if the first pixel transformation information is optical flow transformation information, then the optical flow transformation information is represented by an optical flow transformation matrix. Each element in the optical flow transformation matrix represents the positional offset of the pixel corresponding to that element in the intermediate image and the pixel corresponding to that element in the target image. The method for obtaining the target image by performing pixel transformation on the intermediate image based on the first pixel transformation information can be as follows: traverse the elements of the optical flow transformation matrix; determine the target position information of the pixel based on the positional offset of the traversed element and the current position information of the pixel corresponding to that element in the intermediate image; obtain the current pixel value corresponding to the current position information and the target pixel value corresponding to the target position information in the intermediate image; replace the current pixel value with the target pixel value to obtain the target image.

[0087] In the optical flow transformation matrix, each element can be represented as (Δx, Δy), representing the offset between two positional information. The method to determine the target position information of a pixel based on the positional offset of the traversed element and the current position information of the corresponding pixel in the intermediate image can be as follows: The target position information is obtained by accumulating the current position information with the positional offset; that is, by accumulating the x-coordinate of the current position with the x-coordinate offset Δx to obtain the x-coordinate of the target position, and by accumulating the y-coordinate of the current position with the y-coordinate offset Δy to obtain the y-coordinate of the target position.

[0088] Specifically, assuming the current position information is (x1, y1) and the position offset at the current position is (Δx, Δy), then the target position information is (x1+Δx, y1+Δy). The pixel value of the pixel at the current position (x1, y1) in the intermediate image is then replaced with the pixel value of the pixel at position (x1+Δx, y1+Δy) in the intermediate image. This operation is performed on each pixel in the intermediate image to obtain the target image. For example, Figure 2 This is an example diagram of optical flow transformation applied to the intermediate image in this embodiment. For example... Figure 2 As shown, Figure 2 The image on the left is the intermediate image, and the image on the right is the target image. The corner of the mouth in the left image, after optical flow transformation, becomes the face in the right image. In this embodiment, by performing pixel transformation on the intermediate image using optical flow transformation information, the clarity of the target image can be improved.

[0089] Optionally, the affine transformation information is a matrix with a first predetermined size; the process of performing pixel transformation on the intermediate image according to the first pixel transformation information to obtain the target image may be: for each pixel in the intermediate image, multiplying the current position information of the pixel by the first pixel transformation information to obtain the target position information of the pixel; and transferring the pixel value of the pixel to the position corresponding to the target position information to obtain the target image.

[0090] Wherein, if the first pixel transformation information is affine transformation information, then the first predetermined size is 3*3. The affine transformation information can be represented as... Perspective transformation information can be represented as As can be seen from the above, for the affine transformation matrix, the third row is a known quantity, so the affine transformation information output by the generator is a vector containing six elements; for the perspective transformation matrix, each element is an unknown quantity, so the perspective transformation information output by the generator is a vector containing nine elements, or a 3*3 matrix.

[0091] In this embodiment, for each pixel in the intermediate image, the current pixel's position information is denoted as (x, y), and the target position information is denoted as (x1, y1). Assuming the first pixel transformation information is affine transformation information, performing pixel transformation on the intermediate image based on the first pixel transformation information can be expressed as: The current position information of the pixel is left-multiplied by the affine transformation information to obtain the target position information of the pixel. After obtaining the target position information, the pixel value of the pixel is transferred to the position corresponding to the target position information. This operation is performed on each pixel in the intermediate image to achieve an affine transformation of each pixel and obtain the target image. In this embodiment, pixel transformation of the intermediate image using affine transformation information can improve the clarity of the target image.

[0092] Optionally, the perspective transformation information is a matrix with a second predetermined size. The method for obtaining the target image by performing pixel transformation on the intermediate image based on the first pixel transformation information can be: for each pixel in the intermediate image, left-multiplying the current position information of the pixel by the first pixel transformation information to obtain the target position information of the pixel; and transferring the pixel value of the pixel to the position corresponding to the target position information to obtain the target image.

[0093] Wherein, the first pixel transformation information is perspective transformation information, and the second predetermined size is 3*3. Pixel transformation of the intermediate image based on the first pixel transformation information can be expressed as: The current position information of the pixel is multiplied by the perspective transformation information to obtain the target position information of the pixel. After obtaining the target position information, the pixel value of the pixel is transferred to the position corresponding to the target position information. This operation is performed on each pixel in the intermediate image to achieve perspective transformation for each pixel and obtain the target image. In this embodiment, performing pixel transformation on the intermediate image using perspective transformation information can improve the clarity of the target image.

[0094] Optionally, the generative adversarial network also includes a discriminator; the training method of the generative adversarial neural network is as follows: obtain the original image sample and the corresponding result image sample; input the original image sample into the generator to obtain the intermediate image sample and the second pixel transformation information; perform pixel transformation on the intermediate image sample according to the second pixel transformation information to obtain the generated image; and perform alternating iterative training on the generator and the discriminator based on the generated image, the original image sample and the result image sample.

[0095] The original image sample can be an image containing a human face that has not undergone deformation processing. The resulting image sample can be understood as a high-quality image corresponding to the original image sample after deformation processing; that is, the resulting image sample is an image obtained by deforming the original image sample. In this embodiment, the method of performing pixel transformation on the intermediate image sample according to the second pixel transformation information is the same as the method of performing pixel transformation on the intermediate image according to the first pixel transformation information in the above embodiment, and will not be repeated here.

[0096] Specifically, alternating iterative training of the generator and discriminator can be understood as follows: first, train the discriminator once; then, train the generator once based on the trained discriminator; then, train the discriminator once based on the trained generator, and so on, until the training completion condition is met. In this embodiment, alternating iterative training of the generator and discriminator based on the generated image, original image samples, and result image samples can improve the accuracy of the generator in generating intermediate images and pixel transformation information.

[0097] In this embodiment, the process of alternating iterative training of the generator and discriminator based on the generated image, original image samples, and result image samples can be as follows: forming negative sample pairs with the generated image and original image samples, and forming positive sample pairs with the result image samples and original image samples; inputting the positive sample pairs into the discriminator to obtain a first discrimination result; inputting the negative sample pairs into the discriminator to obtain a second discrimination result; determining a first loss function based on the first and second discrimination results; determining a second loss function based on the generated image and result image samples; linearly superimposing the first and second loss functions to obtain a target loss function; and alternating iterative training of the generator and discriminator based on the target loss function.

[0098] The first and second discrimination results can be values ​​between 0 and 1, used to characterize the matching degree between sample pairs. For positive sample pairs, the true discrimination result is 0, and for negative sample pairs, the true discrimination result is 1. Specifically, the method for determining the first loss function based on the first and second discrimination results can be as follows: calculate the first difference between the first discrimination result and the true discrimination result corresponding to the positive sample pair; calculate the second difference between the second discrimination result and the true discrimination result corresponding to the negative sample pair; and sum the logarithms of the first and second differences to obtain the first loss function.

[0099] The second loss function can be determined by the difference between the generated image and the resulting image samples. Specifically, all original image samples are input into the generative adversarial network (GAN) to obtain the target loss function, which is then backpropagated to adjust the discriminator's parameters. Based on the adjusted discriminator, all original image samples are input into the GAN again to obtain the target loss function, which is then backpropagated to adjust the generator's parameters. This process is repeated iteratively to train the generator and discriminator until the training termination condition is met. For example, Figure 3 This is an example diagram of training a generative adversarial neural network in this embodiment, such as... Figure 3 As shown, the original image sample is input into the generator G to obtain intermediate image samples and second pixel transformation information. Then, the intermediate image sample and second pixel transformation information are input into the pixel transformation module to obtain a generated image. Next, the generated image and the original image sample are paired and input into the discriminator D to obtain a second discrimination result. The original image sample and the resulting image sample are then paired and input into the discriminator D to obtain a first discrimination result. A first loss function is determined based on the first and second discrimination results. A second loss function is determined based on the generated image and the resulting image sample. The first and second loss functions are linearly superimposed to obtain the target loss function. Finally, the generator and discriminator are trained iteratively and alternately based on the target loss function. In this embodiment, the generator and discriminator are trained iteratively and alternately based on the target loss function to address the discrepancy between the generated image and the resulting image sample, thereby improving the generator's accuracy.

[0100] Optionally, the generator includes multiple network layers and at least one pixel transformation module; the pixel transformation module is positioned between two network layers; the forward adjacent network layer of the pixel transformation module outputs a feature map and third pixel transformation information; the pixel transformation module performs pixel transformation on the feature map according to the third pixel transformation information, outputting a transformed feature map; the transformed feature map is input to the backward adjacent network layer of the pixel transformation module. For example, Figure 4 This is an example diagram of the network structure of a generator in this embodiment, such as... Figure 4As shown, the generator contains four network layers. A pixel transformation module is located between network layer 1 and network layer 2, and another between network layer 3 and network layer 4. The first pixel transformation module performs pixel transformation on the feature map output by network layer 1 based on the third pixel transformation information, and inputs the transformed feature map into network layer 2. The second pixel transformation module performs pixel transformation on the feature map output by network layer 3 based on the third pixel transformation information, and inputs the transformed feature map into network layer 4. In this embodiment, embedding the pixel transformation modules between the network layers of the generator allows for deformation processing of facial images within the neural network, reducing workload.

[0101] The technical solution of this disclosure involves inputting the original image into the generator of a generative adversarial network (GAN) to obtain an intermediate image and first pixel transformation information; then, based on the first pixel transformation information, performing pixel transformation on the intermediate image to obtain a target image. The image processing method provided by this disclosure utilizes the first pixel transformation information output by the GAN to perform pixel transformation on the intermediate image to obtain the target image, achieving significant image deformation processing. This overcomes the ghosting problem caused by deformation, thereby improving the image deformation effect.

[0102] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this disclosure. Figure 5 As shown, the device includes:

[0103] The first pixel transformation information acquisition module 210 is used to input the original image into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information.

[0104] Pixel transformation module 220 is used to perform pixel transformation on the intermediate image according to the first pixel transformation information to obtain the target image.

[0105] Optionally, the first pixel transformation information includes optical flow transformation information, affine transformation information, and / or perspective transformation information.

[0106] Optionally, the optical flow transformation information is represented by an optical flow transformation matrix, where each element represents the positional offset between the corresponding pixel in the intermediate image and the corresponding pixel in the target image.

[0107] The pixel transformation module 220 is also used for:

[0108] Traverse the elements of the optical flow transformation matrix, and determine the target position information of the pixel based on the position offset of the traversed element and the current position information of the pixel corresponding to the element in the intermediate image.

[0109] Obtain the current pixel value in the intermediate image that corresponds to the current position information and the target pixel value that corresponds to the target position information;

[0110] Replace the current pixel value with the target pixel value to obtain the target image.

[0111] Optionally, the affine transformation information is a matrix with a first predetermined size.

[0112] The pixel transformation module 220 is also used for:

[0113] For each pixel in the intermediate image, the current position information of that pixel is left-multiplied by the transformation information of the first pixel to obtain the target position information of that pixel; and,

[0114] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0115] Optionally, the perspective transformation information is a matrix with a second predetermined size.

[0116] The pixel transformation module 220 is also used for:

[0117] For each pixel in the intermediate image, the current position information of that pixel is left-multiplied by the transformation information of the first pixel to obtain the target position information of that pixel; and,

[0118] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0119] Optionally, the generative adversarial network also includes a discriminator; and further includes: a generative adversarial neural network training module, used for:

[0120] Obtain the original image samples and the corresponding result image samples;

[0121] The original image sample is input into the generator to obtain the intermediate image sample and the second pixel transformation information;

[0122] The intermediate image samples are pixel-transformed based on the second pixel transformation information to obtain the generated image;

[0123] The generator and discriminator are trained iteratively, alternating between the generated image, the original image samples, and the resulting image samples.

[0124] Furthermore, the generative adversarial neural network training module is also used for:

[0125] The generated image and the original image samples are combined into negative sample pairs, and the resulting image samples and the original image samples are combined into positive sample pairs.

[0126] Positive samples are input into the discriminator to obtain the first discrimination result; negative samples are input into the discriminator to obtain the second discrimination result.

[0127] The first loss function is determined based on the first and second discrimination results;

[0128] The second loss function is determined based on the generated image and the resulting image samples;

[0129] The target loss function is obtained by linearly superimposing the first and second loss functions; and

[0130] The generator and discriminator are trained iteratively and alternately based on the target loss function.

[0131] The above-described apparatus can execute the methods provided in all the foregoing embodiments of this disclosure, and has the corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in this embodiment can be found in the methods provided in all the foregoing embodiments of this disclosure.

[0132] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device 300 suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs, desktop computers, or various forms of servers, such as standalone servers or server clusters. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0133] like Figure 6 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a memory device 305 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0134] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0135] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing a method of word recommendation. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a storage device 305, or installed from a ROM 302. When the computer program is executed by a processing device 301, it performs the functions defined above in the methods of embodiments of this disclosure.

[0136] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0137] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0138] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0139] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: input the original image into the generator of the generative adversarial network to obtain an intermediate image and first pixel transformation information; and perform pixel transformation on the intermediate image according to the first pixel transformation information to obtain a target image.

[0140] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0142] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0145] According to one or more embodiments of the present disclosure, an image processing method is disclosed, including:

[0146] The original image is input into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information;

[0147] The intermediate image is pixel-transformed based on the first pixel transformation information to obtain the target image.

[0148] Furthermore, the first pixel transformation information includes optical flow transformation information, affine transformation information, and / or perspective transformation information.

[0149] Furthermore, the optical flow transformation information is represented by an optical flow transformation matrix, where each element of the optical flow transformation matrix represents the positional offset between the pixel corresponding to that element in the intermediate image and the pixel corresponding to that element in the target image.

[0150] The step of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes:

[0151] The elements of the optical flow transformation matrix are traversed, and the target position information of the pixel is determined based on the position offset of the traversed element and the current position information of the pixel corresponding to the element in the intermediate image.

[0152] Obtain the current pixel value in the intermediate image that corresponds to the current position information and the target pixel value that corresponds to the target position information;

[0153] The current pixel value is replaced with the target pixel value to obtain the target image.

[0154] Furthermore, the affine transformation information is a matrix with a first predetermined size.

[0155] The process of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes:

[0156] For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and,

[0157] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0158] Furthermore, the perspective transformation information is a matrix with a second predetermined size.

[0159] The process of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes:

[0160] For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and,

[0161] The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

[0162] Furthermore, the generative adversarial network also includes a discriminator; the training method of the generative adversarial neural network is as follows:

[0163] Obtain the original image samples and the corresponding result image samples;

[0164] The original image sample is input into the generator to obtain intermediate image samples and second pixel transformation information;

[0165] The intermediate image sample is pixel-transformed according to the second pixel transformation information to obtain the generated image;

[0166] The generator and the discriminator are trained iteratively and alternately based on the generated image, the original image sample, and the result image sample.

[0167] Further, the generator and the discriminator are trained iteratively and alternately based on the generated image, the original image samples, and the resulting image samples, including:

[0168] The generated image and the original image sample are combined into a negative sample pair, and the resulting image sample and the original image sample are combined into a positive sample pair;

[0169] The positive sample pairs are input into the discriminator to obtain a first discrimination result; the negative sample pairs are input into the discriminator to obtain a second discrimination result.

[0170] The first loss function is determined based on the first discrimination result and the second discrimination result;

[0171] A second loss function is determined based on the generated image and the resulting image samples;

[0172] The first loss function and the second loss function are linearly superimposed to obtain the target loss function; and

[0173] The generator and the discriminator are trained iteratively and alternately based on the target loss function.

[0174] Note that the above description is merely a preferred embodiment and the technical principles employed in this disclosure. Those skilled in the art will understand that this disclosure is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this disclosure. Therefore, although this disclosure has been described in detail through the above embodiments, it is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this disclosure, and the scope of this disclosure is determined by the scope of the appended claims.

Claims

1. An image processing method, characterized in that, include: The original image is input into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information; The intermediate image is pixel-transformed according to the first pixel transformation information to obtain the target image; The generator outputs multi-channel data, including image data and pixel transformation information. The generator includes a network layer and a pixel transformation module; the pixel transformation module is disposed between two network layers; the forward adjacent network layer of the pixel transformation module outputs a feature map and third pixel transformation information; the pixel transformation module is used to perform pixel transformation on the feature map according to the third pixel transformation information and output the transformed feature map; the transformed feature map is input to the backward adjacent network layer of the pixel transformation module.

2. The method according to claim 1, characterized in that, The first pixel transformation information includes optical flow transformation information, affine transformation information, and / or perspective transformation information.

3. The method according to claim 2, characterized in that, The optical flow transformation information is represented by an optical flow transformation matrix, where each element represents the positional offset between the corresponding pixel in the intermediate image and the corresponding pixel in the target image. The step of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes: The elements of the optical flow transformation matrix are traversed, and the target position information of the pixel is determined based on the position offset of the traversed element and the current position information of the pixel corresponding to the element in the intermediate image. Obtain the current pixel value in the intermediate image that corresponds to the current position information and the target pixel value that corresponds to the target position information; The current pixel value is replaced with the target pixel value to obtain the target image.

4. The method according to claim 2, characterized in that, The affine transformation information is a matrix with a first predetermined size. The process of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes: For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and, The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

5. The method according to claim 2, characterized in that, The perspective transformation information is a matrix with a second predetermined size. The process of performing pixel transformation on the intermediate image based on the first pixel transformation information to obtain the target image includes: For each pixel in the intermediate image, the current position information of the pixel is left-multiplied by the first pixel transformation information to obtain the target position information of the pixel; and, The pixel value of the pixel is transferred to the position corresponding to the target location information to obtain the target image.

6. The method according to claim 1, characterized in that, The generative adversarial network further includes a discriminator; the training method of the generative adversarial neural network is as follows: Obtain the original image samples and the corresponding result image samples; The original image sample is input into the generator to obtain intermediate image samples and second pixel transformation information; The intermediate image sample is pixel-transformed according to the second pixel transformation information to obtain the generated image; The generator and the discriminator are trained iteratively and alternately based on the generated image, the original image sample, and the result image sample.

7. The method according to claim 6, characterized in that, The generator and the discriminator are trained iteratively and alternately based on the generated image, the original image samples, and the resulting image samples, including: The generated image and the original image sample are combined into a negative sample pair, and the resulting image sample and the original image sample are combined into a positive sample pair; The positive sample pairs are input into the discriminator to obtain a first discrimination result; the negative sample pairs are input into the discriminator to obtain a second discrimination result. The first loss function is determined based on the first discrimination result and the second discrimination result; A second loss function is determined based on the generated image and the resulting image samples; The first loss function and the second loss function are linearly superimposed to obtain the target loss function; and The generator and the discriminator are trained iteratively and alternately based on the target loss function.

8. An image processing apparatus, characterized in that, include: The first pixel transformation information acquisition module is used to input the original image into the generator of the generative adversarial network to obtain the intermediate image and the first pixel transformation information; Pixel transformation module, used to perform pixel transformation on the intermediate image according to the first pixel transformation information to obtain the target image; The generator outputs multi-channel data, including image data and pixel transformation information. The generator includes a network layer and a pixel transformation module; the pixel transformation module is disposed between two network layers; the forward adjacent network layer of the pixel transformation module outputs a feature map and third pixel transformation information; the pixel transformation module is used to perform pixel transformation on the feature map according to the third pixel transformation information and output the transformed feature map; the transformed feature map is input to the backward adjacent network layer of the pixel transformation module.

9. An electronic device, characterized in that, The electronic device includes: One or more processing devices; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image processing method as described in any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the image processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable medium

    CN111402159A

  • Generation method of age transformation face image and generative adversarial network model

    CN112883756A

  • Image processing method and device, electronic equipment and storage medium

    CN113706369A