Image Processing Method, Apparatus, Device, and Storage Medium
By generating and adversarial networks to obtain facial defect information and deficit processing on facial images, the problem of missing details of facial features and skin texture in the prior art is solved, and the authenticity of the image is improved.
Patent Information
- Application Number
- CN202210108249.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-01-28
AI Technical Summary
The existing facial image removal method uses skin grinding technology to cause the loss of facial features and skin texture details, forming a fake face effect.
Generative adversarial network (GAN) is used to obtain facial defect information, and remove defects on the facial images to be processed based on this information to avoid global processing.
It effectively avoids the missing details of facial features and skin texture, and improves the authenticity of facial images after deficiencies.
Smart Images

Figure CN114494071B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art
[0002] Existing methods for removing defects from facial images mainly use skin smoothing techniques. Traditional skin smoothing algorithms are composed of various high-pass filtering algorithms and image processing algorithms. This method usually performs global processing on facial images, resulting in the loss of details of facial features and skin textures after skin smoothing, thus forming an obvious fake face effect. Summary of the Invention
[0003] Embodiments of the present disclosure provide an image processing method, apparatus, device, and storage medium, which can achieve defect removal processing of facial images, avoid the loss of details of facial features and skin textures, and thus improve the authenticity of facial images after defect removal.
[0004] In a first aspect, embodiments of the present disclosure provide an image processing method, including:
[0005] Obtain a facial image to be processed;
[0006] Input the facial image to be processed into a generator of a set generative adversarial network to obtain facial defect information;
[0007] Perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image.
[0008] In a second aspect, embodiments of the present disclosure further provide an image processing apparatus, including:
[0009] An image to be processed acquisition module, configured to obtain a facial image to be processed;
[0010] A facial defect information acquisition module, configured to input the facial image to be processed into a generator of a set generative adversarial network to obtain facial defect information;
[0011] A target facial image acquisition module, configured to perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image.
[0012] In a third aspect, embodiments of the present disclosure further provide an electronic device, where the electronic device includes:
[0013] One or more processing devices;
[0014] A storage device, configured to store one or more programs;
[0015] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image processing method as described in the embodiments of the present disclosure.
[0016] In a fourth aspect, embodiments of the present disclosure further provide a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, it implements the image processing method as described in the embodiments of the present disclosure.
[0017] Embodiments of the present disclosure disclose an image processing method, apparatus, device, and storage medium. Obtain a facial image to be processed; input the facial image to be processed into the generator of a set generative adversarial network to obtain facial defect information; perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image. The image processing method provided by the embodiments of the present disclosure performs defect removal processing on the facial image to be processed based on the facial defect information obtained by the generative adversarial network, without performing global processing on the facial image to be processed, avoiding the loss of details of facial features and skin texture, thereby improving the authenticity of the facial image after defect removal. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of an image processing method in the embodiments of the present disclosure;
[0019] Figure 2 is a schematic diagram of the principle of defect removal processing for a facial image in the embodiments of the present disclosure;
[0020] Figure 3 is a schematic structural diagram of an image processing apparatus in the embodiments of the present disclosure;
[0021] Figure 4 is a schematic structural diagram of an electronic device in the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0023] It should be understood that the steps described in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0024] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions of other terms will be given in the following description.
[0025] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0026] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0027] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0028] Figure 1 The following is a flowchart of an image processing method provided for an embodiment of this disclosure. This embodiment is applicable to the situation of removing defects from facial images. This method can be executed by an image processing device, which can be composed of hardware and / or software and is generally integrated in an electronic device with image processing functions. This device can be an electronic device such as a server, a mobile terminal or a server cluster. As Figure 1 shown, the method specifically includes the following steps:
[0029] S110, obtain a facial image to be processed.
[0030] Among them, the facial image to be processed can be an image with defects on the face. In this embodiment, the facial image to be processed can be an image intercepted from the original image, with a set size and face alignment, which can reduce the spatial distribution differences of the facial image to be processed and reduce the learning difficulty of the generative adversarial network. Face alignment can be understood as the line connecting the two eyes being parallel to the horizontal line.
[0031] Optionally, the way to obtain the facial image to be processed can be: perform facial key point recognition on the original image to obtain initial key point information; determine affine transformation information according to the initial key point information and standard key point information; process the original image according to the affine transformation information to obtain the facial image to be processed.
[0032] Among them, the standard key-point information is the key-point information corresponding to the standard facial image. The affine transformation information is characterized by a matrix of a first set size, and the first set size can be 256*256. The key-point information can be the position information of the key points, which is the key-point information of a set number (such as 106) of key points in the human face, and is the point information used to identify the facial contour and the positions of the facial features. The affine transformation information can characterize information such as image flipping, rotation, shearing, and translation. In this embodiment, the affine transformation information characterizes information such as flipping, rotation, shearing, and translation of the to-be-processed facial image to the standard facial image.
[0033] In this embodiment, both the initial key-point information and the characterized key-point information can be represented by matrices. Determining the affine transformation information according to the initial key-point information and the standard key-point information can be transformed into a process of solving the matrix corresponding to the affine transformation information. Assume that A represents the matrix corresponding to the initial key-point information, B represents the matrix corresponding to the standard key-point information, and X is the matrix corresponding to the affine transformation information. Then the following relationship exists among the three: B = XA, so X = B / A. -1 . Among them, the matrices corresponding to the standard key-point information and the matrix corresponding to the affine transformation information have the same size. After obtaining the affine transformation information, the operations of intercepting and aligning the original image can be achieved in one step according to the affine transformation information, which can improve the efficiency of obtaining the to-be-processed facial image.
[0034] Optionally, the method for obtaining the to-be-processed facial image by processing the original image according to the affine transformation information can be: calling a set image transformation function; inputting the affine transformation information and the original image into the set image transformation function to obtain the to-be-processed facial image.
[0035] Among them, the set image transformation function can be a function for implementing operations such as image flipping, rotation, shearing, and translation from an open-source database (opencv) or a local database (mobilecv), for example, it can be the warpAffine function. Specifically, inputting the affine transformation information and the original image into the set image transformation function can achieve the operations of intercepting and aligning the original image. There is no need for developers to rewrite program code, which can reduce the workload of technical personnel.
[0036] Optionally, another method for obtaining the to-be-processed facial image can be: intercepting the facial area of the original image to obtain an initial facial image; performing an alignment operation on the initial facial image; scaling the aligned initial facial image to an image of a set size to obtain the to-be-processed facial image.
[0037] It should be noted that in the translation of the formula in , there may be some inaccuracies in the understanding of the mathematical relationship. The correct formula should be \(X = A^{-1}B\). Here, it is translated according to the content you provided, but please pay attention to the possible error in the original text.The process of cropping the facial region from the original image can be as follows: perform face recognition on the original image, and crop the region where the recognized face is located from the original image to obtain an initial facial image. Performing an alignment operation on the initial facial image can be understood as: rotating the initial facial image so that the line connecting the two eyes of the rotated facial image is parallel to the horizontal line. Scaling the aligned initial facial image can be understood as reducing or enlarging the image size. In this embodiment, the operations of cropping, aligning, and scaling are sequentially performed on the original image, so that the obtained facial image to be processed can be recognized by the generator of the generative adversarial network, thereby improving the recognition accuracy of the network.
[0038] Specifically, the way to obtain the facial image to be processed by performing an alignment operation on the scaled initial facial image can be: obtain the angle between the line connecting the two eyes and the horizontal line in the initial facial image; rotate the original facial image based on the angle so that the line connecting the two eyes is parallel to the horizontal line to obtain the facial image to be processed.
[0039] Among them, rotating the original facial image based on the angle can be understood as rotating the original facial image clockwise or counterclockwise by the angle corresponding to the angle between the line connecting the two eyes and the horizontal line so that the line connecting the two eyes is parallel to the horizontal line. In this embodiment, the original facial image is aligned by applying an affine transformation to ensure that the facial image to be processed can be better processed for defects by the generator of the generative adversarial network.
[0040] S120, input the facial image to be processed into the generator of the set generative adversarial network to obtain facial defect information.
[0041] Among them, the generator of the set generative adversarial network has the functions of defect detection and defect removal. The output of the network is a four-channel image of RGBA or an image with only the A channel. The A-channel image carries the position information of the defect and the pixel transformation information of the defect. The pixel transformation information of the defect can be understood as the weight information of the pixel transformation before and after the defect removal.
[0042] S130, perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image.
[0043] Among them, the facial defect information can be represented by a matrix carrying the position information of the defect and the pixel transformation information of the defect.
[0044] Specifically, the way to perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image can be: fuse the color channel information of the facial image to be processed with the facial defect information respectively to obtain an intermediate facial image; perform an inverse transformation on the matrix corresponding to the affine transformation information to obtain the inverse affine transformation information; process the intermediate facial image according to the inverse affine transformation information to obtain a target facial image.
[0045] Among them, the color channels include a red (R) channel, a green (G) channel, and a blue (B) channel. In this embodiment, the RGB three-channel information in the facial image to be processed is respectively fused with the facial defect information to obtain the fused RGB three-channel information, and the fused RGB three-channel information constitutes an intermediate facial image.
[0046] In this embodiment, since the intermediate facial image is an image after an affine transformation, it is necessary to multiply the matrix corresponding to the inverse affine transformation information by the intermediate facial image to obtain the target facial image, and finally paste the target facial image back into the original image to achieve the defect removal process of the facial image.
[0047] Specifically, the way to fuse the color channel information of the facial image to be processed with the facial defect information to obtain the intermediate facial image can be: multiplying the matrix corresponding to the color channel information of the facial image to be processed by the matrix corresponding to the facial defect information respectively to obtain the target facial image.
[0048] In this embodiment, since the A channel is a relatively sparse matrix, the scaling of the image has little influence on it. After fusion, most of the pixels on the original human face are well preserved, and the changed pixels are only the defect areas such as pimples. This method solves the problem of reduced clarity often brought by using a generative adversarial network for this task, enabling the completion of high-definition portrait tasks even with a small resolution.
[0049] Optionally, after obtaining the facial image to be processed, the following steps are further included: scaling the facial image to be processed to a second set size.
[0050] Among them, the second set size can be understood as the image size that the generative adversarial neural network can recognize.
[0051] Correspondingly, before performing defect removal processing on the facial image to be processed according to the facial defect information, it also includes: scaling the facial defect information to a first set size.
[0052] In this embodiment, scaling the facial defect information output by the generator to the first set size is beneficial to protecting the clarity of the original human face image.
[0053] Optionally, it is set that the generative adversarial network further includes a discriminator. In this embodiment, the training method of the generative adversarial network is as follows: Obtain a first facial image sample with defects and a corresponding second facial image sample without defects; input the first facial image sample into the generator to obtain a facial defect information sample; perform defect removal processing on the first facial image sample according to the facial defect information sample to obtain a third facial image; form a negative sample pair with the third facial image and the first facial image, and form a positive sample pair with the second facial image and the first facial image; input the negative sample pair into the discriminator to obtain a first discrimination result; input the positive sample pair into the discriminator to obtain a second discrimination result; alternately iteratively train the generator and the discriminator based on the first discrimination result and the second discrimination result.
[0054] Among them, the first facial image sample can be an image obtained by intercepting, scaling, and aligning the original image. The second facial image sample can be an image obtained by performing defect removal processing on the first facial image sample using the set retouching software. Alternate iterative training can be understood as: First, train the discriminator once, then train the generator once based on the trained discriminator, and then train the discriminator once based on the trained generator, and so on, until the training completion condition is met.
[0055] Among them, the method of performing defect removal processing on the first facial image sample according to the facial defect information sample refers to the above embodiment and will not be elaborated here.
[0056] Specifically, the method of alternately iteratively training the generator and the discriminator based on the first discrimination result and the second discrimination result can be: Determine a first loss function according to the first discrimination result, generate a second loss function according to the second discrimination result, linearly superimpose the first loss function and the second loss function to obtain a target loss function, and alternately iteratively train the generator and the discriminator based on the target loss function, which can improve the accuracy of the generative adversarial network.
[0057] Exemplarily, Figure 2 is the schematic diagram of defect removal processing for facial images in this embodiment. As Figure 2 shown, first perform key point recognition on the original image, and determine the affine transformation matrix based on the recognized key point information and the standard key point information; perform interception and alignment operations on the original image based on the affine transformation matrix to obtain a facial image to be processed; perform scaling processing on the facial image to be processed, and input the scaled image into the generative adversarial network (GAN) to output an RGBA image; scale the A channel in the RGBA image, fuse the scaled A channel with the image to be processed, and multiply the fused image by the inverse affine transformation matrix to obtain a target facial image; finally, paste the target facial image back into the original image to obtain the result image.
[0058] In the technical solution of the embodiment of the present disclosure, a facial image to be processed is obtained; the facial image to be processed is input into the generator of a set generative adversarial network to obtain facial defect information; and the facial image to be processed is processed to remove defects according to the facial defect information to obtain a target facial image. The image processing method provided by the embodiment of the present disclosure processes the facial image to be processed to remove defects based on the facial defect information obtained by the generative adversarial network, without globally processing the facial image to be processed, avoiding the loss of details of facial features and skin texture, thereby improving the authenticity of the facial image after defect removal.
[0059] Figure 3 It is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As Figure 3 shown, the device includes:
[0060] A facial image to be processed acquisition module 210, configured to acquire a facial image to be processed;
[0061] A facial defect information acquisition module 220, configured to input the facial image to be processed into the generator of a set generative adversarial network to obtain facial defect information;
[0062] A target facial image acquisition module 230, configured to process the facial image to be processed to remove defects according to the facial defect information to obtain a target facial image.
[0063] Optionally, the facial image to be processed acquisition module 210 is further configured to:
[0064] Perform facial key point recognition on the original image to obtain initial key point information;
[0065] Determine affine transformation information according to the initial key point information and standard key point information; wherein, the standard key point information is the key point information corresponding to a standard facial image; the affine transformation information is represented by a matrix of a first set size;
[0066] Process the original image according to the affine transformation information to obtain a facial image to be processed.
[0067] Optionally, the facial image to be processed acquisition module 210 is further configured to:
[0068] Call a set image transformation function;
[0069] Input the affine transformation information and the original image into the set image transformation function to perform cropping and alignment operations on the original image to obtain a facial image to be processed.
[0070] Optionally, the target facial image acquisition module 230 is further configured to:
[0071] Fuse the color channel information of the facial image to be processed with the facial defect information respectively to obtain an intermediate facial image;
[0072] Perform an inverse transformation on the matrix corresponding to the affine transformation information to obtain the inverse affine transformation information;
[0073] Process the intermediate facial image according to the inverse affine transformation information to obtain the target facial image.
[0074] Optionally, the target facial defect information includes defect position information and defect pixel transformation information; the target facial image acquisition module 230 is further configured to:
[0075] Multiply the matrices corresponding to the color channel information of the facial image to be processed by the matrix corresponding to the facial defect information respectively to obtain the target facial image.
[0076] Optionally, it further includes: a scaling module, configured to:
[0077] Scale the facial image to be processed to a second set size;
[0078] Scale the facial defect information to a first set size.
[0079] Optionally, the set generative adversarial network further includes a discriminator; it further includes: a training module of the set generative adversarial network, configured to:
[0080] Obtain a first facial image sample with defects and a corresponding second facial image sample without defects;
[0081] Input the first facial image sample into the generator to obtain a facial defect information sample;
[0082] Perform defect removal processing on the first facial image sample according to the facial defect information sample to obtain a third facial image;
[0083] Form a negative sample pair with the third facial image and the first facial image, and form a positive sample pair with the second facial image and the first facial image;
[0084] Input the negative sample pair into the discriminator to obtain a first discrimination result; input the positive sample pair into the discriminator to obtain a second discrimination result;
[0085] Perform alternating iterative training on the generator and the discriminator based on the first discrimination result and the second discrimination result.
[0086] The above device can execute the methods provided by all the foregoing embodiments of the present disclosure, and has corresponding functional modules and beneficial effects for executing the above methods. For technical details not described in detail in this embodiment, reference may be made to the methods provided by all the foregoing embodiments of the present disclosure.
[0087] Next, refer to Figure 4, which shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., or various forms of servers, such as independent servers or server clusters. Figure 4 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0088] As Figure 4 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to the program stored in the read-only storage device (ROM) 302 or the program loaded from the storage device 305 into the random access storage device (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0089] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.
[0090] Specifically, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the word recommendation method. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 309, or installed from the storage device 305, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of the embodiments of the present disclosure are executed.
[0091] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0092] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0093] The above computer-readable medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device.
[0094] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a facial image to be processed; input the facial image to be processed into a generator of a set generative adversarial network to obtain facial defect information; and perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image.
[0095] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0097] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0098] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0100] According to one or more embodiments of the embodiments of the present disclosure, the embodiments of the present disclosure disclose an image processing method, including:
[0101] Obtain a facial image to be processed;
[0102] Input the facial image to be processed into the generator of a set generative adversarial network to obtain facial defect information;
[0103] Perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image.
[0104] Further, obtaining a facial image to be processed includes:
[0105] Perform facial key point recognition on the original image to obtain initial key point information;
[0106] Determine affine transformation information according to the initial key point information and standard key point information; wherein, the standard key point information is the key point information corresponding to a standard facial image; the affine transformation information is characterized by a matrix of a first set size;
[0107] Process the original image according to the affine transformation information to obtain a facial image to be processed.
[0108] Further, processing the original image according to the affine transformation information to obtain a facial image to be processed, including:
[0109] Invoking a set image transformation function;
[0110] Inputting the affine transformation information and the original image into the set image transformation function to perform cropping and alignment operations on the original image, thereby obtaining a facial image to be processed.
[0111] Further, performing blemish removal processing on the facial image to be processed according to the facial blemish information to obtain a target facial image, including:
[0112] Fusing the color channel information of the facial image to be processed with the facial blemish information respectively to obtain an intermediate facial image;
[0113] Performing an inverse transformation on the matrix corresponding to the affine transformation information to obtain inverse affine transformation information;
[0114] Processing the intermediate facial image according to the inverse affine transformation information to obtain a target facial image.
[0115] Further, the target facial blemish information includes blemish position information and blemish pixel transformation information; fusing the color channel information of the facial image to be processed with the facial blemish information respectively to obtain an intermediate facial image, including:
[0116] Multiplying the matrices corresponding to the color channel information of the facial image to be processed with the matrices corresponding to the facial blemish information respectively to obtain a target facial image.
[0117] Further, after obtaining the facial image to be processed, it further includes:
[0118] Scaling the facial image to be processed to a second set size;
[0119] Before performing blemish removal processing on the facial image to be processed according to the facial blemish information, it further includes:
[0120] Scaling the facial blemish information to a first set size.
[0121] Further, the set generative adversarial network further includes a discriminator; the training method of the set generative adversarial network is:
[0122] Obtaining a first facial image sample with blemishes and a corresponding second facial image sample without blemishes;
[0123] Inputting the first facial image sample into the generator to obtain a facial blemish information sample;
[0124] Perform defect removal processing on the first facial image sample according to the facial defect information sample to obtain a third facial image;
[0125] Form a negative sample pair with the third facial image and the first facial image, and form a positive sample pair with the second facial image and the first facial image;
[0126] Input the negative sample pair into the discriminator to obtain a first discrimination result; input the positive sample pair into the discriminator to obtain a second discrimination result;
[0127] Perform alternating iterative training on the generator and the discriminator based on the first discrimination result and the second discrimination result.
[0128] Note that the above is only a preferred embodiment of the present disclosure and the applied technical principles. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present disclosure. Therefore, although the present disclosure has been described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments. Without departing from the concept of the present disclosure, more other equivalent embodiments can be included, and the scope of the present disclosure is determined by the scope of the appended claims.
Claims
1. An image processing method, characterized in that, Including: Obtain a facial image to be processed; Input the facial image to be processed into the generator of a set generative adversarial network to obtain facial defect information; wherein, the facial defect information is a four-channel RGBA image or an image with only the A channel, and the A channel image carries the position information of the defect and the defect pixel transformation information, and the defect pixel transformation information includes the weight information of the pixel transformation before and after defect removal; Perform defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image.
2. The method according to claim 1, characterized in that, Obtaining a facial image to be processed includes: Perform facial key point recognition on the original image to obtain initial key point information; Determine affine transformation information according to the initial key point information and standard key point information; wherein, the standard key point information is the key point information corresponding to a standard facial image; the affine transformation information is represented by a matrix of a first set size; Process the original image according to the affine transformation information to obtain a facial image to be processed.
3. The method according to claim 2, characterized in that, Processing the original image according to the affine transformation information to obtain a facial image to be processed includes: Call a set image transformation function; Input the affine transformation information and the original image into the set image transformation function to perform cropping and alignment operations on the original image to obtain a facial image to be processed.
4. The method according to claim 2, characterized in that, Performing defect removal processing on the facial image to be processed according to the facial defect information to obtain a target facial image includes: Fuse the color channel information of the facial image to be processed with the facial defect information respectively to obtain an intermediate facial image; Perform an inverse transformation on the matrix corresponding to the affine transformation information to obtain inverse affine transformation information; Process the intermediate facial image according to the inverse affine transformation information to obtain a target facial image.
5. The method according to claim 4, characterized in that, The facial defect information includes defect position information and defect pixel transformation information; fusing the color channel information of the facial image to be processed with the facial defect information respectively to obtain an intermediate facial image includes: Multiply the matrices corresponding to the color channel information of the facial image to be processed with the matrix corresponding to the facial defect information respectively to obtain a target facial image.
6. The method according to claim 2, characterized in that, After obtaining the facial image to be processed, it further includes: Scale the facial image to be processed to a second set size; Before performing defect removal processing on the facial image to be processed according to the facial defect information, it further includes: Scale the facial defect information to a first set size.
7. The method according to claim 1, characterized in that, The set generative adversarial network further includes a discriminator; the training method of the set generative adversarial network is: Obtain a first facial image sample with defects and a corresponding second facial image sample without defects; Input the first facial image sample into the generator to obtain a facial defect information sample; Perform defect removal processing on the first facial image sample according to the facial defect information sample to obtain a third facial image; Form a negative sample pair with the third facial image and the first facial image, and form a positive sample pair with the second facial image and the first facial image; Input the negative sample pair into the discriminator to obtain a first discrimination result; Input the positive sample pair into the discriminator to obtain a second discrimination result; Based on the first discrimination result and the second discrimination result, alternately and iteratively train the generator and the discriminator.
8. An image processing apparatus, characterized in that, It includes: A to-be-processed image acquisition module, configured to acquire a to-be-processed facial image; A facial defect information acquisition module, configured to input the to-be-processed facial image into a generator of a set generative adversarial network to obtain facial defect information; wherein, the facial defect information is a four-channel image of RGBA or an image with only the A channel, and the A-channel image carries the position information of the defect and the defect pixel transformation information, and the defect pixel transformation information includes the weight information of the pixel transformation before and after the defect removal; A target facial image acquisition module, configured to perform defect removal processing on the to-be-processed facial image according to the facial defect information to obtain a target facial image.
9. An electronic device, characterized in that, The electronic device includes: One or more processing devices; A storage device, configured to store one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image processing method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the image processing method according to any one of claims 1-7.
Citation Information
Patent Citations
An appearance defect detection method based on deep convolutional generative adversarial network sample generation
CN109598287A
Neural network training method and device and identification method and device
CN109784255A
Image processing method and electronic equipment
CN111553854A